Volume 27 - Issue 1

Research Article Biomedical Science and Research Biomedical Science and Research CC by Creative Commons, CC-BY

How to Know the Inter and Intra Genotype Variability of the H Gene in the Canine Distemper Virus?

*Corresponding author: Eva Kuennemann, MVS Pharma GmbH, Leinfelder Str. 62, Leinfelden Echterdingen, Germany.

Received: May 12, 2025; Published: May 15, 2025

DOI: 10.34297/AJBSR.2025.26.003515

Abstract

Currently, the Canine Distemper Virus (CDV) is one of the most common and observable pathogens within the area of canine clinical medicine, where in addition to affecting other wild species, it can cause a variety of clinical signs, since it covers various systems. Within individuals, manifesting itself in a respiratory, digestive, cutaneous, nervous or ocular form, it should be noted that it is common due to its clinical signs, this condition is also known as canine distemper disease.

This virus has the ability to affect a wide range of hosts, including species of carnivorous mammals, which gives it great capacity to generate reservoirs in the species and thus allow it to establish its permanence over time, along with the ability to mutate and generate new genotypes.

This aspect is of great importance, since it prevents the correct control of this pathogenic agent through vaccines that can be prepared based on viral hemagglutinin together with diagnostic tests such as the reverse transcription Polymerase Chain Reaction test (RT- PCR), whose H gene is considered the most variable within the virus genome.

Due to these reasons, determining the current variability of the CDV H gene would allow an approach to what is described in the veterinary clinic: “, in relation to this could lead to lower effectiveness of both vaccines and detection methods used in both preventive and internal medicine.

Despite the extensive vaccination that exists today to prevent serious clinical conditions caused by the virus, failures have been recorded in the immunity granted by the application of the vaccines, where it is suspected that there have been interactions of the strains used in the manufacture of vaccines with wild strains that generate a reservoir, this is due to the high percentage of variability that the virus presents.

Thus, in this document we propose to identify and use all the sequences of the genotypes currently identified for the H gene stored in Genbank®, the genetic sequence database of the NIH (USA) to date, based on the tree phylogenetic study published by Ke et al, in 2015, in order to cover all these published genomic sequences, and then identify and group them according to their established genotype. Then we will proceed with its alignment and comparison using the Clustal Omega computer program or another similar one, thus obtaining the nucleotide identity matrix or the set of genomic sequences that compose it, which would allow us by difference to know the percentage of variability of the gene. H of VDC in a general way and, in the same way, know the variability that exists between and within genotypes

Thus, a more in-depth and detailed discussion and comparative analysis could be carried out regarding the variability and number of genotypes that currently exist worldwide regarding the H gene in CDV and what this could impact on animal species in danger of extinction

Background

The Virus

Canine Distemper Virus (CDV) is known to cause one of the most common multisystem viral diseases in domestic canines called Canine Distemper (CD), with wide worldwide distribution [1] that produces a multisystem clinical infection with high rates of morbidity and mortality that mainly affects young individuals, where despite the use of preventive vaccines, the percentage of registered cases has been increasing in recent years [2,3].

This virus corresponds to a Morbillivirus of the Paramyxoviridae family [4], which is a highly pathogenic agent that is characterized by having an envelope and a size of 150 to 300 nm in diameter [1] containing a molecule of negative sense single-stranded RNA which is associated with a nucleoprotein, which is present in serum, whole blood and cerebrospinal fluid, being a target ideal for detection by molecular detection methods, such as reverse transcription- PCR (RT-PCR) [5], where its genome has 6 genes whose function is to encode the six structural proteins of the virion, which correspond to N, P, M, F, H and L, from which the H gene is identified as the one with the greatest genetic variability, which has served as the basis for various studies to characterize strains and genotypes of the virus [4,6].

Viral Pathogenesis

The virus is released mainly through the oronasal route, although it can be found in other types of secretions in the same way, where its main routes of entry into the body are the ocular, respiratory and oral routes, through direct contact or inhalation of the virus transported in the air or in droplets, where upon reaching mucosal tissue it establishes the first interaction with the host and its immune system through the early infection of local lymphocytes and CD150+ mononuclear cells, where the virus deploys a series of mechanisms that allow neutralize and evade the innate and adaptive antiviral immune response [7].

From this point, it can use the cells of the immune system as a transport vehicle to the regional lymph nodes and carry out replication in subpopulations of lymphocytes between the first and third day post infection, establishment of primary viremia associated with leukocytes, massive replication in lymphoid organs with selective depletion of the Th1 subpopulation and the establishment of the systemic condition on the seventh day post infection [7,8]. Therefore, the initial replication organ is the lymphoid tissue of the upper respiratory tract, subsequently dispersing throughout the organism in the mononuclear cells of the bloodstream directly to the respiratory, digestive and nervous systems, where depending on the immune status of the host it can produce symptoms. of pneumonia, gastroenteritis, skin disorders and a condition of the central nervous system [9].

Symptoms and Clinical Signs

The typical course of the disease is characterized by replication in lymphoid tissue that causes severe immunosuppression that per sists over several weeks, where a variety of signs can be observed that includes fever, respiratory and enteric signs, which in cases of patients with a higher degree of immunosuppression may progress to neurological signs. Similarly, DCV infection can cause eye diseases, skin lesions, dental defects and abortions [8,10].

The degree of the clinical picture and the tissues involved varies depending on the age of the animal, its predisposition to the disease, the strain of the virus, which increasingly present a greater record of mutations along with the appearance of new strains and the immunological status of the animal patient [11,12].

It affects individuals of all ages, although it is more common to find in puppies between 3 and 6 months, which is related to the decrease in maternal immunity. Polysystemic and acute disease occurs in individuals with deficient immune responses, while individuals with an adequate immune response may not be clinically affected [12].

In the case of a low immune response, the virus can reach epithelial tissue and the central nervous system, where it can invade astrocytes, microglia, oligodendrocytes, neurons, ependymal cells and cells of the choroid plexus, which produces lesions characteristic of demyelination when they are oligodendrocytes. However, astrocytes are the cell population that is mainly affected [12].

Neurological signs typically include partial or complete progressive tetraparesis, vestibular signs, seizures, and dementia. Mainly, the most common signs that can be observed in canines diagnosed with distemper encephalitis are myoclonus of the temporal muscles, the anterior limbs, and mouth movements [12,13].

Diagnosis and Detection

There are various methodologies for the detection of CDV, where histological detection stands out, where inclusion bodies can be found in cells of the oral and conjunctival mucosa and in the analysis of cerebrospinal fluid [14], serological techniques such as immunohistochemistry assays and ELISA. (Enzyme-Linked Immunosorbent Assay) and molecular techniques such as the RT-PCR assay, since the clinical symptoms associated with DCV infection are similar to those produced by other viral agents, which is why this no specificity makes it difficult to reach the final diagnosis using this tool [15].

One of the most used tests is ELISA, which detects serum antibodies IgM (against the nucleocapsid proteins, N and P of the VDC) or IgG (against the envelope antigens, H and F of the VDC). However, the detection of these antibodies is non-specific, yielding false negatives and positives, in addition to not specifying whether they correspond to maternal, vaccine or infection antibodies. Furthermore, this test depends on the stage of infection in which the patient is, and may be absent in the initial and final stages of the infection, giving negative results [16].

The H Gene

The CDV genome spans approximately 15,700 nucleotides and is composed of 6 genes that encode viral proteins: N (nucleoprotein); P (phosphoprotein), M (Matrix), F (Fusion), H (hemagglutinin) and L (Large Polymerase) [4].

The H gene encodes the viral hemagglutinin, located in the viral envelope and the target used for the generation of antibodies by vaccines against DCV and has implications in both the tropism and cytopathogenicity of the virus [7,17].

In 2015, a phylogenetic tree was constructed taking sequences from the H gene and the existence of at least 14 VDC genotypes was proposed, from which the intrinsic variability of the H gene is derived [18].

The H protein is a constituent of the envelope glycoprotein spikes in the virion and initiates entry into the host cell by binding to cellular receptors such as signaling lymphocyte activation molecule (SLAM, CD150) or PVRL4 [18], where it should be noted that it was discovered that lymphocyte infection is mediated by the viral hemagglutinin or protein H that is located, as mentioned, in the viral envelope together with the other integral protein called fusion (F), which, being a glycoprotein of the lipid envelope, recognizes and mediates the binding of the virus to these lymphocyte receptors CD150/SLAM (Signaling Lymphocyte Activation Molecule) in the cell membrane [17]. Therefore, it has been determined that the H protein is the main agent of fusogenicity or fusion of the virus with the host and is also the main determinant of viral tropism, along with its contribution to the growth and tropism of the VDC [17].

Therefore, it has been concluded that the H protein has an essential role in cellular tropism, where its ability to have a high variability or antigenic and sequence variation can affect the levels of virulence that the pathogenic agent may have, along with the host range that it can affect and the VDC neutralization epitopes [18].

Phylogenetic Tree and Use of MEGA Software

Viral genetic variability as well as evolution over time is information valuable for, as an example, the development of long-term effective vaccines and the monitoring its effectiveness in the future [19]. Phylogenetic trees are used to visualize evolutionary relationships between species. The high availability of genetic material information allows us to build trees phylogenetics in thousands of species and has had important contributions in taxonomy, epidemiology or virology [20].

The MEGA software (acronym from Molecular Evolutionary Genetic Analysis - Analysis of Molecular Evolutionary Genetics) allows estimating evolutionary distances, reconstruct phylogenetic trees and calculate basic statistical quantities from data molecular [21]. MEGA version 11 adds many methods and tools to keep pace with the growing needs of researchers [22].

In this context, through this work, it is proposed to use the sequences of the H gene existing in the official database of the NIH (USA) - called Genbank® - to carry out a comparative analysis of the known nucleotide sequences of the gene. H at a global level with respect to the 14 genotypes published [18], since as previously described, the H gene has the greatest genetic variability within the viral genome, which is why it is commonly used for molecular typing of existing CDV strains.

As a result, a nucleotide identity matrix or set of genomic sequences that make up the H gene (0-100%) could be established, from which, by difference, the percentage or current genomic variability of the H gene could be established from VDC to through comparative analysis facilitated by the use of computer programs inter and intra genotypes, where the difference and variability present within and between them can be established in order to characterize them and carry out a comparative analysis regarding the difference that may exist. and categorize them within the same H gene.

To determine the membership of the sequences to be used, it will be done through the MEGA program when entering the sequence data to build the new phylogenetic tree.

Materials and Methods

This work can be carried out in any Virology or Animal Microbiology laboratory in the third world, which has qualified personnel capable of selecting the nucleotide sequences of the H gene published in Genbank® and identifying the different genotypes of the selected sequences and finally determining the variability genetic within and between genotypes of the H gene with respect to the updated information of the new phylogenetic tree.

Activity 1

Obtain The Pool of Nucleotide Sequences to Analyze:

To carry out this work, an ASUS computer, Ryzen 7 fourth generation at 3.8 Ghz, will be used, together with the Microsoft Office Word computer program for the taking and collection of data and genomic sequences, together with the Google Chrome search engine.

As a comparison criterion and database contemplated for this report, the 14 genotypes described in the last updated phylogenetic tree regarding the CDV H gene that has been published [18] called “Phylodynamic analysis of the canine distemper virus hemagglutinin gene”, where the access codes or loci that are published in the phylogenetic tree will be extracted and subsequently introduced into the search engine of the website called Genbank®.

As mentioned above, the obtaining and analysis of data contemplates using the existing information in the Genbank® database, where the access codes of the 14 genotypes will be used as keywords, which are subdivided in 14 genotypes:

1 and 2. Europe-1/South America-1: (33 recorded sequences); 3. South America-2: (10 recorded sequences); 4. South America-3: (3 recorded sequences); 5. European wildlife: (9 recorded sequences); 6. Asia-4: (2 recorded sequences); 7. America-2: (15 recorded sequences); 8. Rockborn-like: (5 recorded sequences); 9. Africa-2: (6 recorded sequences), 10. Asia-1: (64 recorded sequences); 11. Artic: (13 recorded sequences); 12. Africa-1: (5 recorded sequences); 13. Asia-2: (11 recorded sequences); and 14. America-1: (12 recorded sequences).

After obtaining the previous sequences, these will be integrated into the MEGA program to obtain again the phylogenetic tree made [18] and verify if there are changes or if it remains the same to obtain a point of comparison for the collection of updated data.

Following in the same way the line of obtaining data through Genbank®, the keywords H gene, hemagglutinin gene or similar will be used to obtain a broader search of the published gene sequences where, when entered into the search engine, approximately 3200 results that represent the existing sequences in the database until the date of this year 2024, in order to obtain the new sequences that have been published to obtain updated data and the sequences not counted in the previous work in order to obtain a more representative quantity, and in the same way update the information present up to this last tree.

In this work we suggest to obtain a maximum number of 2000 sequences from the Genbank® database in order to ensure that the numbers we obtain when obtaining the percentages are representative and more accurately communicate the information we want to obtain.

Likewise, once these sequences are obtained, they will be placed again in the Mega program to give way to a new phylogenetic tree updated to the current date to compare it with the previous one published and use it as a point of comparison of the different established and unknown genotypes that were found. may have added to date.

All new sequences that were not considered for the previous study will be ordered and assigned using the Mega computer program to identify the different genotypes established for Gene H.

It should be taken into account that all sequences published in the database present an access code or locus to identify their genotype within those already known, while those that are not found within the 2015 phylogenetic tree will be considered new sequences.

For the sequences that will be used within the maximum of 2000 to which we aspire, the temporality and geography of the published sequences will be used as search criteria, in order to not always obtain the same samples without some variation and thus be able to ensure that the variability data are representative and come from different years and locations, providing greater dynamism in the sequences.

In this way, the previously known sequences will be used and the new ones published since 2016 will be added to obtain updated results that give greater representativeness to the values that are sought to be obtained and that provide a greater point of comparison and discussion regarding the criteria mentioned above that will mainly include identified genotypes, temporality and geographical location.

Regarding the chain length criteria for the genomic sequences to be included the total size of the H gene that appears in Genbank® will be considered and, in those in which the complete genome appears, the nt that correspond to the gene will be considered. H.

Once the nucleotide sequences have been identified after entering the codes in the search engine and the keywords, they will be selected in FASTA format and will be ordered one by one in groups with respect to their corresponding genotype in the Microsoft Office Word computer program to generate the pool of sequences. to perform the analysis later.

Activity 2

Obtain The Percentages of Variability of the H Gene Within and Between Genotypes:

The sequences obtained will be subsequently entered into the Clustal Omega computer program through Internet access through the Google Chrome search engine, where they will be placed with respect to the percentage of variability that is to be generated, depending on whether the aim is to obtain the intra-genotype figure or between genotypes and subsequently in general, in order to obtain the alignment of the sequences.

Within the Clustal Omega computer program website, you choose the option to analyze in RNA format, then copy the sequences selected to analyze, and then choose the Submit option and let the program work and generate the identity matrix. corresponding nucleotide or the set of genomic sequences that make up the H genotypes, this will be located in the “result files” section within the Clustal Omega page where we will also have the percentage of identity nucleotide (PIN) among these.

The percentage of variability of the H gene can be obtained using the expression (100-PIN), where after generating the results they will be compared to each other to determine the degree of similarity of nucleotide base pairs, and thus by difference between these values begin to calculate the percentage of variability that exists by calculating the difference in percentages, which will be done for each of the 14 determined genotypes and their sequences registered in the phylogenetic tree for the H gene along with the possible new sequences that can be incorporated and thus being able to make the comparison between and within genotypes and also obtain a general value that will allow for comparative analysis and discussion regarding the values that will be obtained.

Discussion

By following this methodology that involves biotools such as Clustal Omega and the MEGA program, it will be possible to understand the variability of the most variable gene in the genome of the Canine Distemper Virus, a virus that has crossed the animal species barrier, affecting animals that are currently in danger of extinction.

Conclusion

The fantastic idea of Kary Mullis [23] in conjunction with the total or partial sequencing of genomes has allowed us to advance in great steps and that medicine is divided into before and after PCR is not trivial.

Acknowledgements

The authors thank Dr. Aron Mosnaim of Wolf Found, Illinois, USA, for his constant support in the development of science in countries like ours.

Conflict of Interest

None.

References

Sign up for Newsletter

Sign up for our newsletter to receive the latest updates. We respect your privacy and will never share your email address with anyone else.