Publications

Subjects: AgingAnalytical methodCell biologyComparative genomicsDisease studyEpigenomicsEvolutionary genomicsFunctional genomicsGenetic variationGenomic sequencingPseudogeneReview / PerspectiveSoftware / Pipeline / DatabaseStructural RNAsSystems biology

Years


Aging (18)

Evidence for Improved DNA Repair in the Long-Lived Bowhead Whale.
Firsanov D, Zacher M, Tian X, Sformo TL, Zhao Y, Tombline G, Lu JY, Zheng Z, Perelli L, Gurreri E, Zhang L, Guo J, Korotkov A, Volobaev V, Biashad SA, Zhang Z, Heid J, Maslov AY, Sun S, Wu Z, Gigas J, Hillpot EC, Martinez JC, Lee M, Williams A, Gilman A, Hamilton N, Strelkova E, Haseljic E, Patel A, Straight ME, Miller N, Ablaeva J, Tam LM, Couderc C, Hoopmann MR, Moritz RL, Fujii S, Pelletier A, Hayman DJ, Liu H, Cai Y, Leung AKL, Zhang Z, Nelson CB, Abegglen LM, Schiffman JD, Gladyshev VN, Maley CC, Modesti M, Genovese G, Simons MJP, Vijg J, Seluanov A, Gorbunova V (2025) Nature 648(8094):717-725.  JOURNAL   PUBMED   REPRINT     AGING · CELL BIOLOGY · COMPARATIVE GENOMICS
ABSTRACT: At more than 200 years, the maximum lifespan of the bowhead whale exceeds that of all other mammals. The bowhead is also the second-largest animal on Earth1, reaching over 80,000 kg. Despite its very large number of cells and long lifespan, the bowhead is not highly cancer-prone, an incongruity termed Peto's paradox2. Here, to understand the mechanisms that underlie the cancer resistance of the bowhead whale, we examined the number of oncogenic hits required for malignant transformation of whale primary fibroblasts. Unexpectedly, bowhead whale fibroblasts required fewer oncogenic hits to undergo malignant transformation than human fibroblasts. However, bowhead whale cells exhibited enhanced DNA double-strand break repair capacity and fidelity, and lower mutation rates than cells of other mammals. We found the cold-inducible RNA-binding protein CIRBP to be highly expressed in bowhead fibroblasts and tissues. Bowhead whale CIRBP enhanced both non-homologous end joining and homologous recombination repair in human cells, reduced micronuclei formation, promoted DNA end protection, and stimulated end joining in vitro. CIRBP overexpression in Drosophila extended lifespan and improved resistance to irradiation. These findings provide evidence supporting the hypothesis that, rather than relying on additional tumour suppressor genes to prevent oncogenesis3-5, the bowhead whale maintains genome integrity through enhanced DNA repair. This strategy, which does not eliminate damaged cells but faithfully repairs them, may be contributing to the exceptional longevity and low cancer incidence in the bowhead whale.

Identification of Functional Rare Coding Variants in IGF-1 Gene in Humans with Exceptional Longevity.
Ali A, Zhang ZD, Gao T, Aleksic S, Gavathiotis E, Barzilai N, Milman S (2025) Sci Rep 15(1):10199.  JOURNAL   PUBMED   REPRINT     AGING · CELL BIOLOGY · COMPARATIVE GENOMICS
ABSTRACT: Diminished signaling via insulin/insulin-like growth factor-1 (IGF-1) axis is associated with longevity in different model organisms. IGF-1 gene is highly conserved across species, with only few evolutionary changes identified in it. Despite its potential role in regulating lifespan, no coding variants in IGF-1 have been reported in human longevity cohorts to date. This study investigated the whole exome sequencing data from 2,108 individuals in a cohort of Ashkenazi Jewish centenarians, their offspring, and controls without familial longevity to identify functional IGF-1 coding variants. We identified two likely functional coding variants IGF-1:p.Ile91Leu and IGF-1:p.Ala118Thr in our longevity cohort. Notably, a centenarian specific novel variant IGF-1:p.Ile91Leu was located at the binding interface of IGF-1-IGF-1R, whereas IGF-1:p.Ala118Thr was significantly associated with lower circulating levels of IGF-1. We performed extended all-atom molecular dynamics simulations to evaluate the impact of Ile91Leu on stability, binding dynamics and energetics of IGF-1 bound to IGF-1R. The IGF-1:p.Ile91Leu formed less stable interactions with IGF-1R's critical binding pocket residues and demonstrated lower binding affinity at the extracellular binding site compared to wild-type IGF-1. Our findings suggest that IGF-1:p.Ile91Leu and IGF-1:p.Ala118Thr variants attenuate IGF-1R activity by impairing IGF-1 binding and diminishing the circulatory levels of IGF-1, respectively. Consequently, diminished IGF-1 signaling resulting from these variants may contribute to exceptional longevity in humans.

Genetic Variants Associated with Age-Related Episodic Memory Decline Implicate Distinct Memory Pathologies.
Ali A, Milman S, Weiss EF, Gao T, Napolioni V, Barzilai N, Zhang ZD, Lin JR (2025) Alzheimers Dement 21(1):e14379.  JOURNAL   PUBMED   REPRINT     AGING · DISEASE STUDY · GENETIC VARIATION
ABSTRACT: BACKGROUND: Approximately 40% of people aged ≥ 65 experience memory loss, particularly in episodic memory. Identifying the genetic basis of episodic memory decline is crucial for uncovering its underlying causes. METHODS: We investigated common and rare genetic variants associated with episodic memory decline in 742 (632 for rare variants) Ashkenazi Jewish individuals (mean age 75) from the LonGenity study. All-atom molecular dynamics simulations were performed to uncover mechanistic insights underlying rare variants associated with episodic memory decline. RESULTS: In addition to the common polygenic risk of Alzheimer's disease, we identified and replicated rare variant associations in ITSN1 and CRHR2. Structural analyses revealed distinct memory pathologies mediated by interfacial rare coding variants such as impaired receptor activation of corticotropin releasing hormone and dysregulated L-serine synthesis. DISCUSSION: Our study uncovers novel risk loci for episodic memory decline. The identified underlying mechanisms point toward heterogenous memory pathologies mediated by rare coding variants. HIGHLIGHTS: We demonstrated the contribution of the common polygenic risk of Alzheimer's disease to episodic memory decline. We discovered and replicated two risk genes associated with episodic memory decline implicated by rare variants, were discovered and replicated. We demonstrated molecular mechanisms and potential novel memory pathologies underlying interfacial rare coding variants. Molecular dynamics simulations were performed to understand the downstream effects of risk rare coding variants.

Polygenic Prediction of Human Longevity on the Supposition of Pervasive Pleiotropy.
Jabalameli MR, Lin JR, Zhang Q, Wang Z, Mitra J, Nguyen N, Gao T, Khusidman M, Sathyan S, Atzmon G, Milman S, Vijg J, Barzilai N, Zhang ZD (2024) Sci Rep 14(1):19981.  JOURNAL   PUBMED   REPRINT     AGING · GENETIC VARIATION
ABSTRACT: The highly polygenic nature of human longevity renders pleiotropy an indispensable feature of its genetic architecture. Leveraging the genetic correlation between aging-related traits (ARTs), we aimed to model the additive variance in lifespan as a function of the cumulative liability from pleiotropic segregating variants. We tracked allele frequency changes as a function of viability across different age bins and prioritized 34 variants with an immediate implication on lipid metabolism, body mass index (BMI), and cognitive performance, among other traits, revealed by PheWAS analysis in the UK Biobank. Given the highly complex and non-linear interactions between the genetic determinants of longevity, we reasoned that a composite polygenic score would approximate a substantial portion of the variance in lifespan and developed the integrated longevity genetic scores (iLGSs) for distinguishing exceptional survival. We showed that coefficients derived from our ensemble model could potentially reveal an interesting pattern of genomic pleiotropy specific to lifespan. We assessed the predictive performance of our model for distinguishing the enrichment of exceptional longevity among long-lived individuals in two replication cohorts (the Scripps Wellderly cohort and the Medical Genome Reference Bank (MRGB)) and showed that the median lifespan in the highest decile of our composite prognostic index is up to 4.8 years longer. Finally, using the proteomic correlates of iLGS, we identified protein markers associated with exceptional longevity irrespective of chronological age and prioritized drugs with repurposing potentials for gerotherapeutics. Together, our approach demonstrates a promising framework for polygenic modeling of additive liability conferred by ARTs in defining exceptional longevity and assisting the identification of individuals at a higher risk of mortality for targeted lifestyle modifications earlier in life. Furthermore, the proteomic signature associated with iLGS highlights the functional pathway upstream of the PI3K-Akt that can be effectively targeted to slow down aging and extend lifespan.

Frailty Resilience Score: A Novel Measure of Frailty Resilience Associated with Protection from Frailty and Survival.
Milman S, Lerman B, Ayers E, Zhang Z, Sathyan S, Levine M, Ye K, Gao T, Higgins-Chen A, Barzilai N, Verghese J (2023) J Gerontol A Biol Sci Med Sci 78(10):1771-1777.  JOURNAL   PUBMED   REPRINT     AGING
ABSTRACT: Frailty is characterized by increased vulnerability to disability and high risk for mortality in older adults. Identification of factors that contribute to frailty resilience is an important step in the development of effective therapies that protect against frailty. First, a reliable quantification of frailty resilience is needed. We developed a novel measure of frailty resilience, the Frailty Resilience Score (FRS), that integrates frailty genetic risk, age, and sex. Application of FRS to the LonGenity cohort (n = 467, mean age 74.4) demonstrated its validity compared to phenotypic frailty and its utility as a reliable predictor of overall survival. In a multivariable-adjusted analysis, 1-standard deviation increase in FRS predicted a 38% reduction in the hazard of mortality, independent of baseline frailty (p < .001). Additionally, FRS was used to identify a proteomic profile of frailty resilience. FRS was shown to be a reliable measure of frailty resilience that can be applied to biological studies of resilience.

◸ NEWS AND VIEWS 
Unravelling Genetic Components of Longevity.
Jabalameli MR, Zhang ZD (2022) Nat Aging 2(1):5-6.  JOURNAL   PUBMED   REPRINT     AGING · DISEASE STUDY · GENETIC VARIATION · REVIEW / PERSPECTIVE
ABSTRACT: Many aging-related traits share a common genetic component. How to disentangle it from the trait-specific effects has remained largely unexplored. A new study in Nature Aging uses an analysis framework for isolating the shared genetic component in genome-wide association studies of aging-related traits and identifies genomic loci that contribute to aging.

Genomic Expansion of Aldh1a1 Protects Beavers Against High Metabolic Aldehydes from Lipid Oxidation.
Zhang Q, Tombline G, Ablaeva J, Zhang L, Zhou X, Smith Z, Zhao Y, Xiaoli AM, Wang Z, Lin JR, Jabalameli MR, Mitra J, Nguyen N, Vijg J, Seluanov A, Gladyshev VN, Gorbunova V, Zhang ZD (2021) Cell Rep 37(6):109965.  JOURNAL   PUBMED   REPRINT     AGING · COMPARATIVE GENOMICS · EVOLUTIONARY GENOMICS
ABSTRACT: The North American beaver is an exceptionally long-lived and cancer-resistant rodent species. Here, we report the evolutionary changes in its gene coding sequences, copy numbers, and expression. We identify changes that likely increase its ability to detoxify aldehydes, enhance tumor suppression and DNA repair, and alter lipid metabolism, potentially contributing to its longevity and cancer resistance. Hpgd, a tumor suppressor gene, is uniquely duplicated in beavers among rodents, and several genes associated with tumor suppression and longevity are under positive selection in beavers. Lipid metabolism genes show positive selection signals, changes in copy numbers, or altered gene expression in beavers. Aldh1a1, encoding an enzyme for aldehydes detoxification, is particularly notable due to its massive expansion in beavers, which enhances their cellular resistance to ethanol and capacity to metabolize diverse aldehyde substrates from lipid oxidation and their woody diet. We hypothesize that the amplification of Aldh1a1 may contribute to the longevity of beavers.

Rare Genetic Coding Variants Associated with Human Longevity and Protection Against Age-Related Diseases.
Lin JR, Sin-Chan P, Napolioni V, Torres GG, Mitra J, Zhang Q, Jabalameli MR, Wang Z, Nguyen N, Gao T, Regeneron Genetics Center, Laudes M, Görg S, Franke A, Nebel A, Greicius MD, Atzmon G, Ye K, Gorbunova V, Ladiges WC, Shuldiner AR, Niedernhofer LJ, Robbins PD, Milman S, Suh Y, Vijg J, Barzilai N, Zhang ZD (2021) Nat Aging 1(9):783-794.  JOURNAL   PUBMED   REPRINT   WEBSITE     AGING · DISEASE STUDY · GENETIC VARIATION
ABSTRACT: Extreme longevity in humans has a strong genetic component, but whether this involves genetic variation in the same longevity pathways as found in model organisms is unclear. Using whole-exome sequences of a large cohort of Ashkenazi Jewish centenarians to examine enrichment for rare coding variants, we found most longevity-associated rare coding variants converge upon conserved insulin/insulin-like growth factor 1 signaling and AMP-activating protein kinase signaling pathways. Centenarians have a number of pathogenic rare coding variants similar to control individuals, suggesting that rare variants detected in the conserved longevity pathways are protective against age-related pathology. Indeed, we detected a pro-longevity effect of rare coding variants in the Wnt signaling pathway on individuals harboring the known common risk allele APOE4. The genetic component of extreme human longevity constitutes, at least in part, rare coding variants in pathways that protect against aging, including those that control longevity in model organisms.

Genetics of Extreme Human Longevity to Guide Drug Discovery for Healthy Ageing.
Zhang ZD, Milman S, Lin JR, Wierbowski S, Yu H, Barzilai N, Gorbunova V, Ladiges WC, Niedernhofer LJ, Suh Y, Robbins PD, Vijg J (2020) Nat Metab 2(8):663-672.  JOURNAL   PUBMED   REPRINT     AGING · DISEASE STUDY · GENETIC VARIATION · REVIEW / PERSPECTIVE
ABSTRACT: Ageing is the greatest risk factor for most common chronic human diseases, and it therefore is a logical target for developing interventions to prevent, mitigate or reverse multiple age-related morbidities. Over the past two decades, genetic and pharmacologic interventions targeting conserved pathways of growth and metabolism have consistently led to substantial extension of the lifespan and healthspan in model organisms as diverse as nematodes, flies and mice. Recent genetic analysis of long-lived individuals is revealing common and rare variants enriched in these same conserved pathways that significantly correlate with longevity. In this Perspective, we summarize recent insights into the genetics of extreme human longevity and propose the use of this rare phenotype to identify genetic variants as molecular targets for gaining insight into the physiology of healthy ageing and the development of new therapies to extend the human healthspan.

Beaver and Naked Mole Rat Genomes Reveal Common Paths to Longevity.
Zhou X, Dou Q, Fan G, Zhang Q, Sanderford M, Kaya A, Johnson J, Karlsson EK, Tian X, Mikhalchenko A, Kumar S, Seluanov A, Zhang ZD, Gorbunova V, Liu X, Gladyshev VN (2020) Cell Rep 32(4):107949.  JOURNAL   PUBMED   REPRINT     AGING · COMPARATIVE GENOMICS · EVOLUTIONARY GENOMICS
ABSTRACT: Long-lived rodents have become an attractive model for the studies on aging. To understand evolutionary paths to long life, we prepare chromosome-level genome assemblies of the two longest-lived rodents, Canadian beaver (Castor canadensis) and naked mole rat (NMR, Heterocephalus glaber), which were scaffolded with in vitro proximity ligation and chromosome conformation capture data and complemented with long-read sequencing. Our comparative genomic analyses reveal that amino acid substitutions at "disease-causing" sites are widespread in the rodent genomes and that identical substitutions in long-lived rodents are associated with common adaptive phenotypes, e.g., enhanced resistance to DNA damage and cellular stress. By employing a newly developed substitution model and likelihood ratio test, we find that energy and fatty acid metabolism pathways are enriched for signals of positive selection in both long-lived rodents. Thus, the high-quality genome resource of long-lived rodents can assist in the discovery of genetic factors that control longevity and adaptive evolution.

Inducible Aging in Hydra Oligactis Implicates Sexual Reproduction, Loss of Stem Cells, and Genome Maintenance as Major Pathways.
Sun S, White RR, Fischer KE, Zhang Z, Austad SN, Vijg J (2020) Geroscience 42(4):1119-1132.  JOURNAL   PUBMED   REPRINT     AGING · CELL BIOLOGY · COMPARATIVE GENOMICS
ABSTRACT: Freshwater polyps of the genus Hydra do not age. However, temperature stress induces aging and a shift from reproduction by asexual budding to sexual gamete production in a cold-sensitive (CS) strain of H. oligactis. We sequenced the transcriptome of a male CS strain before and after this life history shift and compared changes in gene expression relative to those seen in a cold-resistant (CR) strain that does not undergo a life history shift in response to altered temperature. We found that the switch from non-aging asexual reproduction to aging and sexual reproduction involves upregulation of genes not only involved in gametogenesis but also genes involved in cellular senescence, apoptosis, and DNA repair accompanied by a downregulation of genes involved in stem cell maintenance. These results suggest that aging is a byproduct of sexual reproduction-associated cellular reprogramming and underscore the power of these H. oligactis strains to identify intrinsic mechanisms of aging.

SIRT6 Is Responsible for More Efficient DNA Double-Strand Break Repair in Long-Lived Species.
Tian X, Firsanov D, Zhang Z, Cheng Y, Luo L, Tombline G, Tan R, Simon M, Henderson S, Steffan J, Goldfarb A, Tam J, Zheng K, Cornwell A, Johnson A, Yang JN, Mao Z, Manta B, Dang W, Zhang Z, Vijg J, Wolfe A, Moody K, Kennedy BK, Bohmann D, Gladyshev VN, Seluanov A, Gorbunova V (2019) Cell 177(3):622-638.e22.  JOURNAL   PUBMED   REPRINT     AGING · CELL BIOLOGY · COMPARATIVE GENOMICS
ABSTRACT: DNA repair has been hypothesized to be a longevity determinant, but the evidence for it is based largely on accelerated aging phenotypes of DNA repair mutants. Here, using a panel of 18 rodent species with diverse lifespans, we show that more robust DNA double-strand break (DSB) repair, but not nucleotide excision repair (NER), coevolves with longevity. Evolution of NER, unlike DSB, is shaped primarily by sunlight exposure. We further show that the capacity of the SIRT6 protein to promote DSB repair accounts for a major part of the variation in DSB repair efficacy between short- and long-lived species. We dissected the molecular differences between a weak (mouse) and a strong (beaver) SIRT6 protein and identified five amino acid residues that are fully responsible for their differential activities. Our findings demonstrate that DSB repair and SIRT6 have been optimized during the evolution of longevity, which provides new targets for anti-aging interventions.

Translation Fidelity Coevolves with Longevity.
Ke Z, Mallik P, Johnson AB, Luna F, Nevo E, Zhang ZD, Gladyshev VN, Seluanov A, Gorbunova V (2017) Aging Cell 16(5):988-993.  JOURNAL   PUBMED   REPRINT     AGING · CELL BIOLOGY
ABSTRACT: Whether errors in protein synthesis play a role in aging has been a subject of intense debate. It has been suggested that rare mistakes in protein synthesis in young organisms may result in errors in the protein synthesis machinery, eventually leading to an increasing cascade of errors as organisms age. Studies that followed generally failed to identify a dramatic increase in translation errors with aging. However, whether translation fidelity plays a role in aging remained an open question. To address this issue, we examined the relationship between translation fidelity and maximum lifespan across 17 rodent species with diverse lifespans. To measure translation fidelity, we utilized sensitive luciferase-based reporter constructs with mutations in an amino acid residue critical to luciferase activity, wherein misincorporation of amino acids at this mutated codon re-activated the luciferase. The frequency of amino acid misincorporation at the first and second codon positions showed strong negative correlation with maximum lifespan. This correlation remained significant after phylogenetic correction, indicating that translation fidelity coevolves with longevity. These results give new life to the role of protein synthesis errors in aging: Although the error rate may not significantly change with age, the basal rate of translation errors is important in defining lifespan across mammals.

Cell Culture-Based Profiling Across Mammals Reveals DNA Repair and Metabolism as Determinants of Species Longevity.
Ma S, Upneja A, Galecki A, Tsai YM, Burant CF, Raskind S, Zhang Q, Zhang ZD, Seluanov A, Gorbunova V, Clish CB, Miller RA, Gladyshev VN (2016) Elife 5:e19130.  JOURNAL   PUBMED   REPRINT     AGING · COMPARATIVE GENOMICS
ABSTRACT: Mammalian lifespan differs by >100 fold, but the mechanisms associated with such longevity differences are not understood. Here, we conducted a study on primary skin fibroblasts isolated from 16 species of mammals and maintained under identical cell culture conditions. We developed a pipeline for obtaining species-specific ortholog sequences, profiled gene expression by RNA-seq and small molecules by metabolite profiling, and identified genes and metabolites correlating with species longevity. Cells from longer lived species up-regulated genes involved in DNA repair and glucose metabolism, down-regulated proteolysis and protein transport, and showed high levels of amino acids but low levels of lysophosphatidylcholine and lysophosphatidylethanolamine. The amino acid patterns were recapitulated by further analyses of primate and bird fibroblasts. The study suggests that fibroblast profiling captures differences in longevity across mammals at the level of global gene expression and metabolite levels and reveals pathways that define these differences.

Systems-Level Analysis of Human Aging Genes Shed New Light on Mechanisms of Aging.
Zhang Q, Nogales-Cadenas R, Lin JR, Zhang W, Cai Y, Vijg J, Zhang ZD (2016) Hum Mol Genet 25(14):2934-2947.  JOURNAL   PUBMED   REPRINT     AGING · SYSTEMS BIOLOGY
ABSTRACT: Although studies over the last decades have firmly connected a number of genes and molecular pathways to aging, the aging process as a whole still remains poorly understood. To gain novel insights into the mechanisms underlying aging, instead of considering aging genes individually, we studied their characteristics at the systems level in the context of biological networks. We calculated a comprehensive set of network characteristics for human aging-related genes from the GenAge database. By comparing them with other functional groups of genes, we identified a robust group of aging-specific network characteristics. To find the structural basis and the molecular mechanisms underlying this aging-related network specificity, we also analyzed protein domain interactions and gene expression patterns across different tissues. Our study revealed that aging genes not only tend to be network hubs, playing important roles in communication among different functional modules or pathways, but also are more likely to physically interact and be co-expressed with essential genes. The high expression of aging genes across a large number of tissue types also points to a high level of connectivity among aging genes. Unexpectedly, contrary to the depletion of interactions among hub genes in biological networks, we observed close interactions among aging hubs, which renders the aging subnetworks vulnerable to random attacks and thus may contribute to the aging process. Comparison across species reveals the evolution process of the aging subnetwork. As the organisms become more complex, the complexity of its aging mechanisms increases and their aging hub genes are more functionally connected.

DNA Repair in Species with Extreme Lifespan Differences.
MacRae SL, Croken MM, Calder RB, Aliper A, Milholland B, White RR, Zhavoronkov A, Gladyshev VN, Seluanov A, Gorbunova V, Zhang ZD, Vijg J (2015) Aging (Albany NY) 7(12):1171-84.  JOURNAL   PUBMED   REPRINT     AGING · COMPARATIVE GENOMICS · EVOLUTIONARY GENOMICS
ABSTRACT: Differences in DNA repair capacity have been hypothesized to underlie the great range of maximum lifespans among mammals. However, measurements of individual DNA repair activities in cells and animals have not substantiated such a relationship because utilization of repair pathways among animals--depending on habitats, anatomical characteristics, and life styles--varies greatly between mammalian species. Recent advances in high-throughput genomics, in combination with increased knowledge of the genetic pathways involved in genome maintenance, now enable a comprehensive comparison of DNA repair transcriptomes in animal species with extreme lifespan differences. Here we compare transcriptomes of liver, an organ with high oxidative metabolism and abundant spontaneous DNA damage, from humans, naked mole rats, and mice, with maximum lifespans of ~120, 30, and 3 years, respectively, with a focus on genes involved in DNA repair. The results show that the longer-lived species, human and naked mole rat, share higher expression of DNA repair genes, including core genes in several DNA repair pathways. A more systematic approach of signaling pathway analysis indicates statistically significant upregulation of several DNA repair signaling pathways in human and naked mole rat compared with mouse. The results of this present work indicate, for the first time, that DNA repair is upregulated in a major metabolic organ in long-lived humans and naked mole rats compared with short-lived mice. These results strongly suggest that DNA repair can be considered a genuine longevity assurance system.

Comparative Analysis of Genome Maintenance Genes in Naked Mole Rat, Mouse, and Human.
MacRae SL, Zhang Q, Lemetre C, Seim I, Calder RB, Hoeijmakers J, Suh Y, Gladyshev VN, Seluanov A, Gorbunova V, Vijg J, Zhang ZD (2015) Aging Cell 14(2):288-91.  JOURNAL   PUBMED   REPRINT     AGING · COMPARATIVE GENOMICS · EVOLUTIONARY GENOMICS
ABSTRACT: Genome maintenance (GM) is an essential defense system against aging and cancer, as both are characterized by increased genome instability. Here, we compared the copy number variation and mutation rate of 518 GM-associated genes in the naked mole rat (NMR), mouse, and human genomes. GM genes appeared to be strongly conserved, with copy number variation in only four genes. Interestingly, we found NMR to have a higher copy number of CEBPG, a regulator of DNA repair, and TINF2, a protector of telomere integrity. NMR, as well as human, was also found to have a lower rate of germline nucleotide substitution than the mouse. Together, the data suggest that the long-lived NMR, as well as human, has more robust GM than mouse and identifies new targets for the analysis of the exceptional longevity of the NMR.

Comparative Genetics of Longevity and Cancer: Insights from Long-Lived Rodents.
Gorbunova V, Seluanov A, Zhang Z, Gladyshev VN, Vijg J (2014) Nat Rev Genet 15(8):531-40.  JOURNAL   PUBMED   REPRINT     AGING · REVIEW / PERSPECTIVE
ABSTRACT: Mammals have evolved a remarkable diversity of ageing rates. Within the single order of Rodentia, maximum lifespans range from 4 years in mice to 32 years in naked mole rats. Cancer rates also differ substantially between cancer-prone mice and almost cancer-proof naked mole rats and blind mole rats. Recent progress in rodent comparative biology, together with the emergence of whole-genome sequence information, has opened opportunities for the discovery of genetic factors that control longevity and cancer susceptibility.

Analytical method (19)

Direct Probabilistic Quantification of Mosaic Loss of Chromosome Y from Sequencing Data
Lin J-R, Chang Y-C, Maslov AY, Song Y, Gao T, Shan J, Bennett D, Milman S, Barzilai N, Vijg J, Montagna C, Zhang ZD (2026) bioRxiv 2026.06.26.734767.  JOURNAL     ANALYTICAL METHOD · SOFTWARE / PIPELINE / DATABASE
ABSTRACT: Loss of chromosome Y (LOY) is the most common aneuploidy in aging men and is increasingly recognized as a marker of aging and genomic instability. Because LOY occurs in mosaic form, its degree reflects the fraction of cells lacking the Y chromosome. Existing SNP-array- and sequencing-based methods rely largely on single genomic features and indirect transformations to estimate this fraction. We developed BaySeq-Y, a Bayesian method that directly estimates LOY mosaicism from sequencing data using VCF files with read depth (DP) and allelic depth (AD). Within a rigorous Bayesian framework, BaySeq-Y integrates complementary LOY-associated genomic features, including decreased read depth and allelic imbalance, and can additionally leverage haplotype phasing to improve precision. In simulations and fluorescence in situ hybridization validation (FISH), BaySeq-Y provided accurate estimates and outperformed existing methods. Applications to ROSMAP and GTEx supported its biological relevance through transcriptomic validation, demonstrating its utility for quantifying LOY across diverse sequencing datasets.

Bayesian Estimation of Mosaic Loss of Chromosome Y from Bulk RNA Sequencing Data
Lin J-R, Zhang ZD (2026) bioRxiv 2026.05.20.726153.  JOURNAL     ANALYTICAL METHOD · SOFTWARE / PIPELINE / DATABASE
ABSTRACT: Mosaic loss of chromosome Y (LOY) is a common age-associated somatic alteration in men and is typically measured from DNA-based assays. Many cohorts, however, contain bulk RNA-seq data without matched DNA-based LOY measurements. We developed a Bayesian framework to estimate the fraction of cells with LOY from male bulk RNA-seq by modeling reduced Y-linked gene expression relative to expected expression after adjustment for age, expression covariates, and autosomal/X-linked control genes. In 377 male GTEx samples, individual Y-linked genes showed negative correlations with separately obtained DNA-based LOY measurements, supporting a shared Y-expression depletion signal. The primary fast empirical Bayes estimator achieved a Pearson correlation of 0.678 with measured LOY, a mean absolute error of 1.79%, a root mean squared error of 3.72%, and 95.2% empirical coverage of measured LOY. Performance was strongest for identifying large LOY events, with an AUC of 0.964 for measured LOY greater than 20%, while fine ranking among low-LOY samples remained uncertain. A mixture/PCA hierarchical Bayesian sensitivity model provided similar validation performance and interpretable posterior quantities but did not improve point estimation. Leave-one-Y-gene-out and prior-sensitivity analyses showed that the signal was distributed across multiple Y-linked transcripts and that prior shrinkage affected calibration. In an external whole-blood RNA-seq dataset without measured LOY, estimated LOY showed a modest age-related increase, but ex vivo immune stimulation shifted RNA-derived LOY estimates and reduced multiple Y-linked transcripts, indicating transcriptional confounding. These results show that bulk RNA-seq contains usable information about LOY, especially for larger events, but RNA-derived LOY should be interpreted as a probabilistic transcriptome-based estimate rather than a direct substitute for DNA-based mosaicism measurement.

Rare Coding Variants as Risk Modifiers of the 22q11.2 Deletion Implicate Postnatal Cortical Development in Syndromic Schizophrenia.
Lin JR, Zhao Y, Jabalameli MR, Nguyen N, Mitra J, International 22q11.DS Brain and Behavior Consortium, Swillen A, Vorstman JAS, Chow EWC, van den Bree M, Emanuel BS, Vermeesch JR, Owen MJ, Williams NM, Bassett AS, McDonald-McGinn DM, Gur RE, Bearden CE, Morrow BE, Lachman HM, Zhang ZD (2023) Mol Psychiatry 28(5):2071-2080.  JOURNAL   PUBMED   REPRINT   WEBSITE     ANALYTICAL METHOD · DISEASE STUDY · GENETIC VARIATION
ABSTRACT: 22q11.2 deletion is one of the strongest known genetic risk factors for schizophrenia. Recent whole-genome sequencing of schizophrenia cases and controls with this deletion provided an unprecedented opportunity to identify risk modifying genetic variants and investigate their contribution to the pathogenesis of schizophrenia in 22q11.2 deletion syndrome. Here, we apply a novel analytic framework that integrates gene network and phenotype data to investigate the aggregate effects of rare coding variants and identified modifier genes in this etiologically homogenous cohort (223 schizophrenia cases and 233 controls of European descent). Our analyses revealed significant additive genetic components of rare nonsynonymous variants in 110 modifier genes (adjusted P = 9.4E-04) that overall accounted for 4.6% of the variance in schizophrenia status in this cohort, of which 4.0% was independent of the common polygenic risk for schizophrenia. The modifier genes affected by rare coding variants were enriched with genes involved in synaptic function and developmental disorders. Spatiotemporal transcriptomic analyses identified an enrichment of coexpression between modifier and 22q11.2 genes in cortical brain regions from late infancy to young adulthood. Corresponding gene coexpression modules are enriched with brain-specific protein-protein interactions of SLC25A1, COMT, and PI4KA in the 22q11.2 deletion region. Overall, our study highlights the contribution of rare coding variants to the SCZ risk. They not only complement common variants in disease genetics but also pinpoint brain regions and developmental stages critical to the etiology of syndromic schizophrenia.

Protocol for Gene Annotation, Prediction, and Validation of Genomic Gene Expansion.
Zhang Q, Zhang ZD (2022) STAR Protoc 3(4):101692.  JOURNAL   PUBMED   REPRINT     ANALYTICAL METHOD · COMPARATIVE GENOMICS · EVOLUTIONARY GENOMICS · SOFTWARE / PIPELINE / DATABASE
ABSTRACT: Although gene expansion plays an important role in evolution, its identification remains a challenge due to potential errors in genome assembly and annotation. Here, we describe a detailed step-by-step protocol for gene annotation, prediction of genomic gene expansion, and its computational and experimental validation. Finally, we also detail steps to discover functionality of each copy of replicated genes. For complete details on the use and execution of this protocol, please refer to Zhang et al. (2021).

Deep Post-Gwas Analysis Identifies Potential Risk Genes and Risk Variants for Alzheimer's Disease, Providing New Insights into Its Disease Mechanisms.
Wang Z, Zhang Q, Lin JR, Jabalameli MR, Mitra J, Nguyen N, Zhang ZD (2021) Sci Rep 11(1):20511.  JOURNAL   PUBMED   REPRINT     ANALYTICAL METHOD · DISEASE STUDY · GENETIC VARIATION
ABSTRACT: Alzheimer's disease (AD) is a genetically complex, multifactorial neurodegenerative disease. It affects more than 45 million people worldwide and currently remains untreatable. Although genome-wide association studies (GWAS) have identified many AD-associated common variants, only about 25 genes are currently known to affect the risk of developing AD, despite its highly polygenic nature. Moreover, the risk variants underlying GWAS AD-association signals remain unknown. Here, we describe a deep post-GWAS analysis of AD-associated variants, using an integrated computational framework for predicting both disease genes and their risk variants. We identified 342 putative AD risk genes in 203 risk regions spanning 502 AD-associated common variants. 246 AD risk genes have not been identified as AD risk genes by previous GWAS collected in GWAS catalogs, and 115 of 342 AD risk genes are outside the risk regions, likely under the regulation of transcriptional regulatory elements contained therein. Even more significantly, for 109 AD risk genes, we predicted 150 risk variants, of both coding and regulatory (in promoters or enhancers) types, and 85 (57%) of them are supported by functional annotation. In-depth functional analyses showed that AD risk genes were overrepresented in AD-related pathways or GO terms-e.g., the complement and coagulation cascade and phosphorylation and activation of immune response-and their expression was relatively enriched in microglia, endothelia, and pericytes of the human brain. We found nine AD risk genes-e.g., IL1RAP, PMAIP1, LAMTOR4-as predictors for the prognosis of AD survival and genes such as ARL6IP5 with altered network connectivity between AD patients and normal individuals involved in AD progression. Our findings open new strategies for developing therapeutics targeting AD risk genes or risk variants to influence AD pathogenesis.

PGA: Post-Gwas Analysis for Disease Gene Identification.
Lin JR, Jaroslawicz D, Cai Y, Zhang Q, Wang Z, Zhang ZD (2018) Bioinformatics 34(10):1786-1788.  JOURNAL   PUBMED   REPRINT   WEBSITE     ANALYTICAL METHOD · DISEASE STUDY · GENETIC VARIATION
ABSTRACT: SUMMARY: Although the genome-wide association study (GWAS) is a powerful method to identify disease-associated variants, it does not directly address the biological mechanisms underlying such genetic association signals. Here, we present PGA, a Perl- and Java-based program for post-GWAS analysis that predicts likely disease genes given a list of GWAS-reported variants. Designed with a command line interface, PGA incorporates genomic and eQTL data in identifying disease gene candidates and uses gene network and ontology data to score them based upon the strength of their relationship to the disease in question. AVAILABILITY AND IMPLEMENTATION: http://zdzlab.einstein.yu.edu/1/pga.html. CONTACT: zhengdong.zhang@einstein.yu.edu. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.

Integrated Rare Variant-Based Risk Gene Prioritization in Disease Case-Control Sequencing Studies.
Lin JR, Zhang Q, Cai Y, Morrow BE, Zhang ZD (2017) PLoS Genet 13(12):e1007142.  JOURNAL   PUBMED   REPRINT   WEBSITE     ANALYTICAL METHOD · DISEASE STUDY · GENETIC VARIATION
ABSTRACT: Rare variants of major effect play an important role in human complex diseases and can be discovered by sequencing-based genome-wide association studies. Here, we introduce an integrated approach that combines the rare variant association test with gene network and phenotype information to identify risk genes implicated by rare variants for human complex diseases. Our data integration method follows a 'discovery-driven' strategy without relying on prior knowledge about the disease and thus maintains the unbiased character of genome-wide association studies. Simulations reveal that our method can outperform a widely-used rare variant association test method by 2 to 3 times. In a case study of a small disease cohort, we uncovered putative risk genes and the corresponding rare variants that may act as genetic modifiers of congenital heart disease in 22q11.2 deletion syndrome patients. These variants were missed by a conventional approach that relied on the rare variant association test alone.

◸ JOURNAL HIGHLIGHTED ARTICLE 
Integrated Post-Gwas Analysis Sheds New Light on the Disease Mechanisms of Schizophrenia.
Lin JR, Cai Y, Zhang Q, Zhang W, Nogales-Cadenas R, Zhang ZD (2016) Genetics 204(4):1587-1600.  JOURNAL   PUBMED   REPRINT   WEBSITE     ANALYTICAL METHOD · DISEASE STUDY
ABSTRACT: Schizophrenia is a severe mental disorder with a large genetic component. Recent genome-wide association studies (GWAS) have identified many schizophrenia-associated common variants. For most of the reported associations, however, the underlying biological mechanisms are not clear. The critical first step for their elucidation is to identify the most likely disease genes as the source of the association signals. Here, we describe a general computational framework of post-GWAS analysis for complex disease gene prioritization. We identify 132 putative schizophrenia risk genes in 76 risk regions spanning 120 schizophrenia-associated common variants, 78 of which have not been recognized as schizophrenia disease genes by previous GWAS. Even more significantly, 29 of them are outside the risk regions, likely under regulation of transcriptional regulatory elements contained therein. These putative schizophrenia risk genes are transcriptionally active in both brain and the immune system, and highly enriched among cellular pathways, consistent with leading pathophysiological hypotheses about the pathogenesis of schizophrenia. With their involvement in distinct biological processes, these putative schizophrenia risk genes, with different association strengths, show distinctive temporal expression patterns, and play specific biological roles during brain development.

Prioritization of Schizophrenia Risk Genes by a Network-Regularized Logistic Regression Method.
Zhang W, Lin JR, Nogales-Cadenas R, Zhang Q, Cai Y, Zhang ZD (2016) Lecture Notes in Bioinformatics 9565:434-445.  JOURNAL   REPRINT     ANALYTICAL METHOD · DISEASE STUDY
ABSTRACT: Schizophrenia (SCZ) is a severe mental disorder with a large genetic component. While recent large-scale microarray- and sequencing-based genome wide association studies have made significant progress toward finding SCZ risk variants and genes of subtle effect, the interactions among them were not considered in those studies. Using a protein-protein interaction network both in our regression model and to generate a SCZ gene subnetwork, we developed an analytical framework with Logit-Lapnet, the graphical Laplacian-regularized logistic regression, for whole exome sequencing (WES) data analysis to detect SCZ gene subnetworks. Using simulated data from sequencing-based association study, we compared the performances of Logit-Lapnet with other logistic regression (LR)-based models. We use Logit-Lapnet to prioritize genes according to their coefficients and select top-ranked genes as seeds to generate the gene sub-network that is associated to SCZ. The comparison demonstrated not only the applicability but also better performance of Logit-Lapnet to score disease risk genes using sequencing-based association data. We applied our method to SCZ whole exome sequencing data and selected top-ranked risk genes, the majority of which are either known SCZ genes or genes potentially associated with SCZ. We then used the seed genes to construct SCZ gene subnetworks. This result demonstrates that by rank-ing gene according to their disease contributions our method scores and thus prioritiz-es disease risk genes for further investigation. An implementation of our approach in MATLAB is freely available for download at:
http://zdzlab.einstein.yu.edu/1/publications/LapNet-MATLAB.zip.

SubNet: A Java Application for Subnetwork Extraction.
Lemetre C, Zhang Q, Zhang ZD (2013) Bioinformatics 29(19):2509-11.  JOURNAL   PUBMED   REPRINT   WEBSITE     ANALYTICAL METHOD · SOFTWARE / PIPELINE / DATABASE · SYSTEMS BIOLOGY
ABSTRACT: SUMMARY: The extraction of targeted subnetworks is a powerful way to identify functional modules and pathways within complex networks. Here, we present SubNet, a Java-based stand-alone program for extracting subnetworks, given a basal network and a set of selected nodes. Designed with a graphical user-friendly interface, SubNet combines four different extraction methods, which offer the possibility to interrogate a biological network according to the question investigated. Of note, we developed a method based on the highly successful Google PageRank algorithm to extract the subnetwork using the node centrality metric, to which possible node weights of the selected genes can be incorporated. AVAILABILITY: http://www.zdzlab.org/1/subnet.html

Identification of Genomic Indels and Structural Variations Using Split Reads.
Zhang ZD, Du J, Lam H, Abyzov A, Urban AE, Snyder M, Gerstein M (2011) BMC Genomics 12:375.  JOURNAL   PUBMED   REPRINT     ANALYTICAL METHOD · SOFTWARE / PIPELINE / DATABASE
ABSTRACT: BACKGROUND: Recent studies have demonstrated the genetic significance of insertions, deletions, and other more complex structural variants (SVs) in the human population. With the development of the next-generation sequencing technologies, high-throughput surveys of SVs on the whole-genome level have become possible. Here we present split-read identification, calibrated (SRiC), a sequence-based method for SV detection. RESULTS: We start by mapping each read to the reference genome in standard fashion using gapped alignment. Then to identify SVs, we score each of the many initial mappings with an assessment strategy designed to take into account both sequencing and alignment errors (e.g. scoring more highly events gapped in the center of a read). All current SV calling methods have multilevel biases in their identifications due to both experimental and computational limitations (e.g. calling more deletions than insertions). A key aspect of our approach is that we calibrate all our calls against synthetic data sets generated from simulations of high-throughput sequencing (with realistic error models). This allows us to calculate sensitivity and the positive predictive value under different parameter-value scenarios and for different classes of events (e.g. long deletions vs. short insertions). We run our calculations on representative data from the 1000 Genomes Project. Coupling the observed numbers of events on chromosome 1 with the calibrations gleaned from the simulations (for different length events) allows us to construct a relatively unbiased estimate for the total number of SVs in the human genome across a wide range of length scales. We estimate in particular that an individual genome contains ~670,000 indels/SVs. CONCLUSIONS: Compared with the existing read-depth and read-pair approaches for SV identification, our method can pinpoint the exact breakpoints of SV events, reveal the actual sequence content of insertions, and cover the whole size spectrum for deletions. Moreover, with the advent of the third-generation sequencing technologies that produce longer reads, we expect our method to be even more useful.

ACT: Aggregation and Correlation Toolbox for Analyses of Genome Tracks.
Jee J, Rozowsky J, Yip KY, Lochovsky L, Bjornson R, Zhong G, Zhang Z, Fu Y, Wang J, Weng Z, Gerstein M (2011) Bioinformatics 27(8):1152-4.  JOURNAL   PUBMED   REPRINT   WEBSITE     ANALYTICAL METHOD · SOFTWARE / PIPELINE / DATABASE
ABSTRACT: UNLABELLED: We have implemented aggregation and correlation toolbox (ACT), an efficient, multifaceted toolbox for analyzing continuous signal and discrete region tracks from high-throughput genomic experiments, such as RNA-seq or ChIP-chip signal profiles from the ENCODE and modENCODE projects, or lists of single nucleotide polymorphisms from the 1000 genomes project. It is able to generate aggregate profiles of a given track around a set of specified anchor points, such as transcription start sites. It is also able to correlate related tracks and analyze them for saturation--i.e. how much of a certain feature is covered with each new succeeding experiment. The ACT site contains downloadable code in a variety of formats, interactive web servers (for use on small quantities of data), example datasets, documentation and a gallery of outputs. Here, we explain the components of the toolbox in more detail and apply them in various contexts. AVAILABILITY: ACT is available at http://act.gersteinlab.org CONTACT: pi@gersteinlab.org.

Detection of Copy Number Variation from Array Intensity and Sequencing Read Depth Using a Stepwise Bayesian Model.
Zhang ZD, Gerstein MB (2010) BMC Bioinformatics 11:539.  JOURNAL   PUBMED   REPRINT     ANALYTICAL METHOD · SOFTWARE / PIPELINE / DATABASE
ABSTRACT: BACKGROUND: Copy number variants (CNVs) have been demonstrated to occur at a high frequency and are now widely believed to make a significant contribution to the phenotypic variation in human populations. Array-based comparative genomic hybridization (array-CGH) and newly developed read-depth approach through ultrahigh throughput genomic sequencing both provide rapid, robust, and comprehensive methods to identify CNVs on a whole-genome scale. RESULTS: We developed a Bayesian statistical analysis algorithm for the detection of CNVs from both types of genomic data. The algorithm can analyze such data obtained from PCR-based bacterial artificial chromosome arrays, high-density oligonucleotide arrays, and more recently developed high-throughput DNA sequencing. Treating parameters--e.g., the number of CNVs, the position of each CNV, and the data noise level--that define the underlying data generating process as random variables, our approach derives the posterior distribution of the genomic CNV structure given the observed data. Sampling from the posterior distribution using a Markov chain Monte Carlo method, we get not only best estimates for these unknown parameters but also Bayesian credible intervals for the estimates. We illustrate the characteristics of our algorithm by applying it to both synthetic and experimental data sets in comparison to other segmentation algorithms. CONCLUSIONS: In particular, the synthetic data comparison shows that our method is more sensitive than other approaches at low false positive rates. Furthermore, given its Bayesian origin, our method can also be seen as a technique to refine CNVs identified by fast point-estimate methods and also as a framework to integrate array-CGH and sequencing data with other CNV-related biological knowledge, all through informative priors.

Integrating Sequencing Technologies in Personal Genomics: Optimal Low Cost Reconstruction of Structural Variants.
Du J, Bjornson RD, Zhang ZD, Kong Y, Snyder M, Gerstein MB (2009) PLoS Comput Biol 5(7):e1000432.  JOURNAL   PUBMED   REPRINT   WEBSITE     ANALYTICAL METHOD
ABSTRACT: The goal of human genome re-sequencing is obtaining an accurate assembly of an individual's genome. Recently, there has been great excitement in the development of many technologies for this (e.g. medium and short read sequencing from companies such as 454 and SOLiD, and high-density oligo-arrays from Affymetrix and NimbelGen), with even more expected to appear. The costs and sensitivities of these technologies differ considerably from each other. As an important goal of personal genomics is to reduce the cost of re-sequencing to an affordable point, it is worthwhile to consider optimally integrating technologies. Here, we build a simulation toolbox that will help us optimally combine different technologies for genome re-sequencing, especially in reconstructing large structural variants (SVs). SV reconstruction is considered the most challenging step in human genome re-sequencing. (It is sometimes even harder than de novo assembly of small genomes because of the duplications and repetitive sequences in the human genome.) To this end, we formulate canonical problems that are representative of issues in reconstruction and are of small enough scale to be computationally tractable and simulatable. Using semi-realistic simulations, we show how we can combine different technologies to optimally solve the assembly at low cost. With mapability maps, our simulations efficiently handle the inhomogeneous repeat-containing structure of the human genome and the computational complexity of practical assembly algorithms. They quantitatively show how combining different read lengths is more cost-effective than using one length, how an optimal mixed sequencing strategy for reconstructing large novel SVs usually also gives accurate detection of SNPs/indels, how paired-end reads can improve reconstruction efficiency, and how adding in arrays is more efficient than just sequencing for disentangling some complex SVs. Our strategy should facilitate the sequencing of human genomes at maximum accuracy and low cost.

PEMer: A Computational Framework with Simulation-Based Error Models for Inferring Genomic Structural Variants from Massive Paired-End Sequencing Data.
Korbel JO, Abyzov A, Mu XJ, Carriero N, Cayting P, Zhang Z, Snyder M, Gerstein MB (2009) Genome Biol 10(2):R23.  JOURNAL   PUBMED   REPRINT   WEBSITE     ANALYTICAL METHOD · SOFTWARE / PIPELINE / DATABASE
ABSTRACT: Personal-genomics endeavors, such as the 1000 Genomes project, are generating maps of genomic structural variants by analyzing ends of massively sequenced genome fragments. To process these we developed Paired-End Mapper (PEMer; http://sv.gersteinlab.org/pemer). This comprises an analysis pipeline, compatible with several next-generation sequencing platforms; simulation-based error models, yielding confidence-values for each structural variant; and a back-end database. The simulations demonstrated high structural variant reconstruction efficiency for PEMer's coverage-adjusted multi-cutoff scoring-strategy and showed its relative insensitivity to base-calling errors.

PeakSeq Enables Systematic Scoring of ChIP-seq Experiments Relative to Controls.
Rozowsky J, Euskirchen G, Auerbach RK, Zhang ZD, Gibson T, Bjornson R, Carriero N, Snyder M, Gerstein MB (2009) Nat Biotechnol 27(1):66-75.  JOURNAL   PUBMED   REPRINT   WEBSITE     ANALYTICAL METHOD · SOFTWARE / PIPELINE / DATABASE
ABSTRACT: Chromatin immunoprecipitation (ChIP) followed by tag sequencing (ChIP-seq) using high-throughput next-generation instrumentation is fast, replacing chromatin immunoprecipitation followed by genome tiling array analysis (ChIP-chip) as the preferred approach for mapping of sites of transcription-factor binding and chromatin modification. Using two deeply sequenced data sets for human RNA polymerase II and STAT1, each with matching input-DNA controls, we describe a general scoring approach to address unique challenges in ChIP-seq data analysis. Our approach is based on the observation that sites of potential binding are strongly correlated with signal peaks in the control, likely revealing features of open chromatin. We develop a two-pass strategy called PeakSeq to compensate for this. A two-pass strategy compensates for signal caused by open chromatin, as revealed by inclusion of the controls. The first pass identifies putative binding sites and compensates for genomic variation in the 'mappability' of sequences. The second pass filters out sites not significantly enriched compared to the normalized control, computing precise enrichments and significances. Our scoring procedure enables us to optimize experimental design by estimating the depth of sequencing required for a desired level of coverage and demonstrating that more than two replicates provides only a marginal gain in information.

Modeling ChIP Sequencing in Silico with Applications.
Zhang ZD, Rozowsky J, Snyder M, Chang J, Gerstein M (2008) PLoS Comput Biol 4(8):e1000158.  JOURNAL   PUBMED   REPRINT   WEBSITE     ANALYTICAL METHOD
ABSTRACT: ChIP sequencing (ChIP-seq) is a new method for genomewide mapping of protein binding sites on DNA. It has generated much excitement in functional genomics. To score data and determine adequate sequencing depth, both the genomic background and the binding sites must be properly modeled. To develop a computational foundation to tackle these issues, we first performed a study to characterize the observed statistical nature of this new type of high-throughput data. By linking sequence tags into clusters, we show that there are two components to the distribution of tag counts observed in a number of recent experiments: an initial power-law distribution and a subsequent long right tail. Then we develop in silico ChIP-seq, a computational method to simulate the experimental outcome by placing tags onto the genome according to particular assumed distributions for the actual binding sites and for the background genomic sequence. In contrast to current assumptions, our results show that both the background and the binding sites need to have a markedly nonuniform distribution in order to correctly model the observed ChIP-seq data, with, for instance, the background tag counts modeled by a gamma distribution. On the basis of these results, we extend an existing scoring approach by using a more realistic genomic-background model. This enables us to identify transcription-factor binding sites in ChIP-seq data in a statistically rigorous fashion.

Statistical Analysis of the Genomic Distribution and Correlation of Regulatory Elements in the ENCODE Regions.
Zhang ZD, Paccanaro A, Fu Y, Weissman S, Weng Z, Chang J, Snyder M, Gerstein MB (2007) Genome Res 17(6):787-97.  JOURNAL   PUBMED   REPRINT   POSTER   WEBSITE     ANALYTICAL METHOD · FUNCTIONAL GENOMICS
ABSTRACT: The comprehensive inventory of functional elements in 44 human genomic regions carried out by the ENCODE Project Consortium enables for the first time a global analysis of the genomic distribution of transcriptional regulatory elements. In this study we developed an intuitive and yet powerful approach to analyze the distribution of regulatory elements found in many different ChIP-chip experiments on a 10 approximately 100-kb scale. First, we focus on the overall chromosomal distribution of regulatory elements in the ENCODE regions and show that it is highly nonuniform. We demonstrate, in fact, that regulatory elements are associated with the location of known genes. Further examination on a local, single-gene scale shows an enrichment of regulatory elements near both transcription start and end sites. Our results indicate that overall these elements are clustered into regulatory rich "islands" and poor "deserts." Next, we examine how consistent the nonuniform distribution is between different transcription factors. We perform on all the factors a multivariate analysis in the framework of a biplot, which enhances biological signals in the experiments. This groups transcription factors into sequence-specific and sequence-nonspecific clusters. Moreover, with experimental variation carefully controlled, detailed correlations show that the distribution of sites was generally reproducible for a specific factor between different laboratories and microarray platforms. Data sets associated with histone modifications have particularly strong correlations. Finally, we show how the correlations between factors change when only regulatory elements far from the transcription start sites are considered.

A Supervised Hidden Markov Model Framework for Efficiently Segmenting Tiling Array Data in Transcriptional and chIP-chip Experiments: Systematically Incorporating Validated Biological Knowledge.
Du J, Rozowsky JS, Korbel JO, Zhang ZD, Royce TE, Schultz MH, Snyder M, Gerstein M (2006) Bioinformatics 22(24):3016-24.  JOURNAL   PUBMED   REPRINT   WEBSITE     ANALYTICAL METHOD
ABSTRACT: MOTIVATION: Large-scale tiling array experiments are becoming increasingly common in genomics. In particular, the ENCODE project requires the consistent segmentation of many different tiling array datasets into 'active regions' (e.g. finding transfrags from transcriptional data and putative binding sites from ChIP-chip experiments). Previously, such segmentation was done in an unsupervised fashion mainly based on characteristics of the signal distribution in the tiling array data itself. Here we propose a supervised framework for doing this. It has the advantage of explicitly incorporating validated biological knowledge into the model and allowing for formal training and testing. METHODOLOGY: In particular, we use a hidden Markov model (HMM) framework, which is capable of explicitly modeling the dependency between neighboring probes and whose extended version (the generalized HMM) also allows explicit description of state duration density. We introduce a formal definition of the tiling-array analysis problem, and explain how we can use this to describe sampling small genomic regions for experimental validation to build up a gold-standard set for training and testing. We then describe various ideal and practical sampling strategies (e.g. maximizing signal entropy within a selected region versus using gene annotation or known promoters as positives for transcription or ChIP-chip data, respectively). RESULTS: For the practical sampling and training strategies, we show how the size and noise in the validated training data affects the performance of an HMM applied to the ENCODE transcriptional and ChIP-chip experiments. In particular, we show that the HMM framework is able to efficiently process tiling array data as well as or better than previous approaches. For the idealized sampling strategies, we show how we can assess their performance in a simulation framework and how a maximum entropy approach, which samples sub-regions with very different signal intensities, gives the maximally performing gold-standard. This latter result has strong implications for the optimum way medium-scale validation experiments should be carried out to verify the results of the genome-scale tiling array experiments.

Cell biology (9)

Evidence for Improved DNA Repair in the Long-Lived Bowhead Whale.
Firsanov D, Zacher M, Tian X, Sformo TL, Zhao Y, Tombline G, Lu JY, Zheng Z, Perelli L, Gurreri E, Zhang L, Guo J, Korotkov A, Volobaev V, Biashad SA, Zhang Z, Heid J, Maslov AY, Sun S, Wu Z, Gigas J, Hillpot EC, Martinez JC, Lee M, Williams A, Gilman A, Hamilton N, Strelkova E, Haseljic E, Patel A, Straight ME, Miller N, Ablaeva J, Tam LM, Couderc C, Hoopmann MR, Moritz RL, Fujii S, Pelletier A, Hayman DJ, Liu H, Cai Y, Leung AKL, Zhang Z, Nelson CB, Abegglen LM, Schiffman JD, Gladyshev VN, Maley CC, Modesti M, Genovese G, Simons MJP, Vijg J, Seluanov A, Gorbunova V (2025) Nature 648(8094):717-725.  JOURNAL   PUBMED   REPRINT     AGING · CELL BIOLOGY · COMPARATIVE GENOMICS
ABSTRACT: At more than 200 years, the maximum lifespan of the bowhead whale exceeds that of all other mammals. The bowhead is also the second-largest animal on Earth1, reaching over 80,000 kg. Despite its very large number of cells and long lifespan, the bowhead is not highly cancer-prone, an incongruity termed Peto's paradox2. Here, to understand the mechanisms that underlie the cancer resistance of the bowhead whale, we examined the number of oncogenic hits required for malignant transformation of whale primary fibroblasts. Unexpectedly, bowhead whale fibroblasts required fewer oncogenic hits to undergo malignant transformation than human fibroblasts. However, bowhead whale cells exhibited enhanced DNA double-strand break repair capacity and fidelity, and lower mutation rates than cells of other mammals. We found the cold-inducible RNA-binding protein CIRBP to be highly expressed in bowhead fibroblasts and tissues. Bowhead whale CIRBP enhanced both non-homologous end joining and homologous recombination repair in human cells, reduced micronuclei formation, promoted DNA end protection, and stimulated end joining in vitro. CIRBP overexpression in Drosophila extended lifespan and improved resistance to irradiation. These findings provide evidence supporting the hypothesis that, rather than relying on additional tumour suppressor genes to prevent oncogenesis3-5, the bowhead whale maintains genome integrity through enhanced DNA repair. This strategy, which does not eliminate damaged cells but faithfully repairs them, may be contributing to the exceptional longevity and low cancer incidence in the bowhead whale.

Identification of Functional Rare Coding Variants in IGF-1 Gene in Humans with Exceptional Longevity.
Ali A, Zhang ZD, Gao T, Aleksic S, Gavathiotis E, Barzilai N, Milman S (2025) Sci Rep 15(1):10199.  JOURNAL   PUBMED   REPRINT     AGING · CELL BIOLOGY · COMPARATIVE GENOMICS
ABSTRACT: Diminished signaling via insulin/insulin-like growth factor-1 (IGF-1) axis is associated with longevity in different model organisms. IGF-1 gene is highly conserved across species, with only few evolutionary changes identified in it. Despite its potential role in regulating lifespan, no coding variants in IGF-1 have been reported in human longevity cohorts to date. This study investigated the whole exome sequencing data from 2,108 individuals in a cohort of Ashkenazi Jewish centenarians, their offspring, and controls without familial longevity to identify functional IGF-1 coding variants. We identified two likely functional coding variants IGF-1:p.Ile91Leu and IGF-1:p.Ala118Thr in our longevity cohort. Notably, a centenarian specific novel variant IGF-1:p.Ile91Leu was located at the binding interface of IGF-1-IGF-1R, whereas IGF-1:p.Ala118Thr was significantly associated with lower circulating levels of IGF-1. We performed extended all-atom molecular dynamics simulations to evaluate the impact of Ile91Leu on stability, binding dynamics and energetics of IGF-1 bound to IGF-1R. The IGF-1:p.Ile91Leu formed less stable interactions with IGF-1R's critical binding pocket residues and demonstrated lower binding affinity at the extracellular binding site compared to wild-type IGF-1. Our findings suggest that IGF-1:p.Ile91Leu and IGF-1:p.Ala118Thr variants attenuate IGF-1R activity by impairing IGF-1 binding and diminishing the circulatory levels of IGF-1, respectively. Consequently, diminished IGF-1 signaling resulting from these variants may contribute to exceptional longevity in humans.

Inducible Aging in Hydra Oligactis Implicates Sexual Reproduction, Loss of Stem Cells, and Genome Maintenance as Major Pathways.
Sun S, White RR, Fischer KE, Zhang Z, Austad SN, Vijg J (2020) Geroscience 42(4):1119-1132.  JOURNAL   PUBMED   REPRINT     AGING · CELL BIOLOGY · COMPARATIVE GENOMICS
ABSTRACT: Freshwater polyps of the genus Hydra do not age. However, temperature stress induces aging and a shift from reproduction by asexual budding to sexual gamete production in a cold-sensitive (CS) strain of H. oligactis. We sequenced the transcriptome of a male CS strain before and after this life history shift and compared changes in gene expression relative to those seen in a cold-resistant (CR) strain that does not undergo a life history shift in response to altered temperature. We found that the switch from non-aging asexual reproduction to aging and sexual reproduction involves upregulation of genes not only involved in gametogenesis but also genes involved in cellular senescence, apoptosis, and DNA repair accompanied by a downregulation of genes involved in stem cell maintenance. These results suggest that aging is a byproduct of sexual reproduction-associated cellular reprogramming and underscore the power of these H. oligactis strains to identify intrinsic mechanisms of aging.

The Nutritional Environment Determines Which and How Intestinal Stem Cells Contribute to Homeostasis and Tumorigenesis.
Li W, Zimmerman SE, Peregrina K, Houston M, Mayoral J, Zhang J, Maqbool S, Zhang Z, Cai Y, Ye K, Augenlicht LH (2019) Carcinogenesis 40(8):937-946.  JOURNAL   PUBMED   REPRINT     CELL BIOLOGY
ABSTRACT: Sporadic colon cancer accounts for approximately 80% of colorectal cancer (CRC) with high incidence in Western societies strongly linked to long-term dietary patterns. A unique mouse model for sporadic CRC results from feeding a purified rodent Western-style diet (NWD1) recapitulating intake for the mouse of common nutrient risk factors each at its level consumed in higher risk Western populations. This causes sporadic large and small intestinal tumors in wild-type mice at an incidence and frequency similar to that in humans. NWD1 perturbs intestinal cell maturation and Wnt signaling throughout villi and colonic crypts and decreases mouse Lgr5hi intestinal stem cell contribution to homeostasis and tumor development. Here we establish that NWD1 transcriptionally reprograms Lgr5hi cells, and that nutrients are interactive in reprogramming. Furthermore, the DNA mismatch repair pathway is elevated in Lgr5hi cells by lower vitamin D3 and/or calcium in NWD1, paralleled by reduced accumulation of relevant somatic mutations detected by single-cell exome sequencing. In compensation, NWD1 also reprograms Bmi1+ cells to function and persist as stem-like cells in mucosal homeostasis and tumor development. The data establish the key role of the nutrient environment in defining the contribution of two different stem cell populations to both mucosal homeostasis and tumorigenesis. This raises important questions regarding impact of variable human diets on which and how stem cell populations function in the human mucosa and give rise to tumors. Moreover, major differences reported in turnover of human and mouse crypt base stem cells may be linked to their very different nutrient exposures.

SIRT6 Is Responsible for More Efficient DNA Double-Strand Break Repair in Long-Lived Species.
Tian X, Firsanov D, Zhang Z, Cheng Y, Luo L, Tombline G, Tan R, Simon M, Henderson S, Steffan J, Goldfarb A, Tam J, Zheng K, Cornwell A, Johnson A, Yang JN, Mao Z, Manta B, Dang W, Zhang Z, Vijg J, Wolfe A, Moody K, Kennedy BK, Bohmann D, Gladyshev VN, Seluanov A, Gorbunova V (2019) Cell 177(3):622-638.e22.  JOURNAL   PUBMED   REPRINT     AGING · CELL BIOLOGY · COMPARATIVE GENOMICS
ABSTRACT: DNA repair has been hypothesized to be a longevity determinant, but the evidence for it is based largely on accelerated aging phenotypes of DNA repair mutants. Here, using a panel of 18 rodent species with diverse lifespans, we show that more robust DNA double-strand break (DSB) repair, but not nucleotide excision repair (NER), coevolves with longevity. Evolution of NER, unlike DSB, is shaped primarily by sunlight exposure. We further show that the capacity of the SIRT6 protein to promote DSB repair accounts for a major part of the variation in DSB repair efficacy between short- and long-lived species. We dissected the molecular differences between a weak (mouse) and a strong (beaver) SIRT6 protein and identified five amino acid residues that are fully responsible for their differential activities. Our findings demonstrate that DSB repair and SIRT6 have been optimized during the evolution of longevity, which provides new targets for anti-aging interventions.

Global, Integrated Analysis of Methylomes and Transcriptomes from Laser Capture Microdissected Bronchial and Alveolar Cells in Human Lung.
Dong X, Shi M, Lee M, Toro R, Gravina S, Han W, Yasuda S, Wang T, Zhang Z, Vijg J, Suh Y, Spivack SD (2018) Epigenetics 13(3):264-274.  JOURNAL   PUBMED   REPRINT     CELL BIOLOGY · EPIGENOMICS
ABSTRACT: Gene regulatory analysis of highly diverse human tissues in vivo is essentially constrained by the challenge of performing genome-wide, integrated epigenetic and transcriptomic analysis in small selected groups of specific cell types. Here we performed genome-wide bisulfite sequencing and RNA-seq from the same small groups of bronchial and alveolar cells isolated by laser capture microdissection from flash-frozen lung tissue of 12 donors and their peripheral blood T cells. Methylation and transcriptome patterns differed between alveolar and bronchial cells, while each of these epithelia showed more differences from mesodermally-derived T cells. Differentially methylated regions (DMRs) between alveolar and bronchial cells tended to locate at regulatory regions affecting promoters of 4,350 genes. A large number of pathways enriched for these DMRs including GTPase signal transduction, cell death, and skeletal muscle. Similar patterns of transcriptome differences were observed: 4,108 differentially expressed genes (DEGs) enriched in GTPase signal transduction, inflammation, cilium assembly, and others. Prioritizing using DMR-DEG regulatory network, we highlighted genes, e.g., ETS1, PPARG, and RXRG, at prominent alveolar vs. bronchial cell discriminant nodes. Our results show that multi-omic analysis of small, highly specific cells is feasible and yields unique physiologic loci distinguishing human lung cell types in situ.

Translation Fidelity Coevolves with Longevity.
Ke Z, Mallik P, Johnson AB, Luna F, Nevo E, Zhang ZD, Gladyshev VN, Seluanov A, Gorbunova V (2017) Aging Cell 16(5):988-993.  JOURNAL   PUBMED   REPRINT     AGING · CELL BIOLOGY
ABSTRACT: Whether errors in protein synthesis play a role in aging has been a subject of intense debate. It has been suggested that rare mistakes in protein synthesis in young organisms may result in errors in the protein synthesis machinery, eventually leading to an increasing cascade of errors as organisms age. Studies that followed generally failed to identify a dramatic increase in translation errors with aging. However, whether translation fidelity plays a role in aging remained an open question. To address this issue, we examined the relationship between translation fidelity and maximum lifespan across 17 rodent species with diverse lifespans. To measure translation fidelity, we utilized sensitive luciferase-based reporter constructs with mutations in an amino acid residue critical to luciferase activity, wherein misincorporation of amino acids at this mutated codon re-activated the luciferase. The frequency of amino acid misincorporation at the first and second codon positions showed strong negative correlation with maximum lifespan. This correlation remained significant after phylogenetic correction, indicating that translation fidelity coevolves with longevity. These results give new life to the role of protein synthesis errors in aging: Although the error rate may not significantly change with age, the basal rate of translation errors is important in defining lifespan across mammals.

Cyclin C Regulates Adipogenesis by Stimulating Transcriptional Activity of CCAAT/enhancer-binding Protein Α.
Song Z, Xiaoli AM, Zhang Q, Zhang Y, Yang EST, Wang S, Chang R, Zhang ZD, Yang G, Strich R, Pessin JE, Yang F (2017) J Biol Chem 292(21):8918-8932.  JOURNAL   PUBMED   REPRINT     CELL BIOLOGY
ABSTRACT: Brown adipose tissue is important for maintaining energy homeostasis and adaptive thermogenesis in rodents and humans. As disorders arising from dysregulated energy metabolism, such as obesity and metabolic diseases, have increased, so has interest in the molecular mechanisms of adipocyte biology. Using a functional screen, we identified cyclin C (CycC), a conserved subunit of the Mediator complex, as a novel regulator for brown adipocyte formation. siRNA-mediated CycC knockdown (KD) in brown preadipocytes impaired the early transcriptional program of differentiation, and genetic KO of CycC completely blocked the differentiation process. RNA sequencing analyses of CycC-KD revealed a critical role of CycC in activating genes co-regulated by peroxisome proliferator activated receptor γ (PPARγ) and CCAAT/enhancer-binding protein α (C/EBPα). Overexpression of PPARγ2 or addition of the PPARγ ligand rosiglitazone rescued the defects in CycC-KO brown preadipocytes and efficiently activated the PPARγ-responsive promoters in both WT and CycC-KO cells, suggesting that CycC is not essential for PPARγ transcriptional activity. In contrast, CycC-KO significantly reduced C/EBPα-dependent gene expression. Unlike for PPARγ, overexpression of C/EBPα could not induce C/EBPα target gene expression in CycC-KO cells or rescue the CycC-KO defects in brown adipogenesis, suggesting that CycC is essential for C/EBPα-mediated gene activation. CycC physically interacted with C/EBPα, and this interaction was required for C/EBPα transactivation domain activity. Consistent with the role of C/EBPα in white adipogenesis, CycC-KD also inhibited differentiation of 3T3-L1 cells into white adipocytes. Together, these data indicate that CycC activates adipogenesis in part by stimulating the transcriptional activity of C/EBPα.

INK4 Locus of the Tumor-Resistant Rodent, the Naked Mole Rat, Expresses a Functional p15/p16 Hybrid Isoform.
Tian X, Azpurua J, Ke Z, Augereau A, Zhang ZD, Vijg J, Gladyshev VN, Gorbunova V, Seluanov A (2015) Proc Natl Acad Sci U S A 112(4):1053-8.  JOURNAL   PUBMED   REPRINT     CELL BIOLOGY
ABSTRACT: The naked mole rat (Heterocephalus glaber) is a long-lived and tumor-resistant rodent. Tumor resistance in the naked mole rat is mediated by the extracellular matrix component hyaluronan of very high molecular weight (HMW-HA). HMW-HA triggers hypersensitivity of naked mole rat cells to contact inhibition, which is associated with induction of the INK4 (inhibitors of cyclin dependent kinase 4) locus leading to cell-cycle arrest. The INK4a/b locus is among the most frequently mutated in human cancer. This locus encodes three distinct tumor suppressors: p15(INK4b), p16(INK4a), and ARF (alternate reading frame). Although p15(INK4b) has its own ORF, p16(INK4a) and ARF share common second and third exons with alternative reading frames. Here, we show that, in the naked mole rat, the INK4a/b locus encodes an additional product that consists of p15(INK4b) exon 1 joined to p16(INK4a) exons 2 and 3. We have named this isoform pALT(INK4a/b) (for alternative splicing). We show that pALT(INK4a/b) is present in both cultured cells and naked mole rat tissues but is absent in human and mouse cells. Additionally, we demonstrate that pALT(INK4a/b) expression is induced during early contact inhibition and upon a variety of stresses such as UV, gamma irradiation-induced senescence, loss of substrate attachment, and expression of oncogenes. When overexpressed in naked mole rat or human cells, pALT(INK4a/b) has stronger ability to induce cell-cycle arrest than either p15(INK4b) or p16(INK4a). We hypothesize that the presence of the fourth product, pALT(INK4a/b) of the INK4a/b locus in the naked mole rat, contributes to the increased resistance to tumorigenesis of this species.

Comparative genomics (14)

Evidence for Improved DNA Repair in the Long-Lived Bowhead Whale.
Firsanov D, Zacher M, Tian X, Sformo TL, Zhao Y, Tombline G, Lu JY, Zheng Z, Perelli L, Gurreri E, Zhang L, Guo J, Korotkov A, Volobaev V, Biashad SA, Zhang Z, Heid J, Maslov AY, Sun S, Wu Z, Gigas J, Hillpot EC, Martinez JC, Lee M, Williams A, Gilman A, Hamilton N, Strelkova E, Haseljic E, Patel A, Straight ME, Miller N, Ablaeva J, Tam LM, Couderc C, Hoopmann MR, Moritz RL, Fujii S, Pelletier A, Hayman DJ, Liu H, Cai Y, Leung AKL, Zhang Z, Nelson CB, Abegglen LM, Schiffman JD, Gladyshev VN, Maley CC, Modesti M, Genovese G, Simons MJP, Vijg J, Seluanov A, Gorbunova V (2025) Nature 648(8094):717-725.  JOURNAL   PUBMED   REPRINT     AGING · CELL BIOLOGY · COMPARATIVE GENOMICS
ABSTRACT: At more than 200 years, the maximum lifespan of the bowhead whale exceeds that of all other mammals. The bowhead is also the second-largest animal on Earth1, reaching over 80,000 kg. Despite its very large number of cells and long lifespan, the bowhead is not highly cancer-prone, an incongruity termed Peto's paradox2. Here, to understand the mechanisms that underlie the cancer resistance of the bowhead whale, we examined the number of oncogenic hits required for malignant transformation of whale primary fibroblasts. Unexpectedly, bowhead whale fibroblasts required fewer oncogenic hits to undergo malignant transformation than human fibroblasts. However, bowhead whale cells exhibited enhanced DNA double-strand break repair capacity and fidelity, and lower mutation rates than cells of other mammals. We found the cold-inducible RNA-binding protein CIRBP to be highly expressed in bowhead fibroblasts and tissues. Bowhead whale CIRBP enhanced both non-homologous end joining and homologous recombination repair in human cells, reduced micronuclei formation, promoted DNA end protection, and stimulated end joining in vitro. CIRBP overexpression in Drosophila extended lifespan and improved resistance to irradiation. These findings provide evidence supporting the hypothesis that, rather than relying on additional tumour suppressor genes to prevent oncogenesis3-5, the bowhead whale maintains genome integrity through enhanced DNA repair. This strategy, which does not eliminate damaged cells but faithfully repairs them, may be contributing to the exceptional longevity and low cancer incidence in the bowhead whale.

Identification of Functional Rare Coding Variants in IGF-1 Gene in Humans with Exceptional Longevity.
Ali A, Zhang ZD, Gao T, Aleksic S, Gavathiotis E, Barzilai N, Milman S (2025) Sci Rep 15(1):10199.  JOURNAL   PUBMED   REPRINT     AGING · CELL BIOLOGY · COMPARATIVE GENOMICS
ABSTRACT: Diminished signaling via insulin/insulin-like growth factor-1 (IGF-1) axis is associated with longevity in different model organisms. IGF-1 gene is highly conserved across species, with only few evolutionary changes identified in it. Despite its potential role in regulating lifespan, no coding variants in IGF-1 have been reported in human longevity cohorts to date. This study investigated the whole exome sequencing data from 2,108 individuals in a cohort of Ashkenazi Jewish centenarians, their offspring, and controls without familial longevity to identify functional IGF-1 coding variants. We identified two likely functional coding variants IGF-1:p.Ile91Leu and IGF-1:p.Ala118Thr in our longevity cohort. Notably, a centenarian specific novel variant IGF-1:p.Ile91Leu was located at the binding interface of IGF-1-IGF-1R, whereas IGF-1:p.Ala118Thr was significantly associated with lower circulating levels of IGF-1. We performed extended all-atom molecular dynamics simulations to evaluate the impact of Ile91Leu on stability, binding dynamics and energetics of IGF-1 bound to IGF-1R. The IGF-1:p.Ile91Leu formed less stable interactions with IGF-1R's critical binding pocket residues and demonstrated lower binding affinity at the extracellular binding site compared to wild-type IGF-1. Our findings suggest that IGF-1:p.Ile91Leu and IGF-1:p.Ala118Thr variants attenuate IGF-1R activity by impairing IGF-1 binding and diminishing the circulatory levels of IGF-1, respectively. Consequently, diminished IGF-1 signaling resulting from these variants may contribute to exceptional longevity in humans.

Protocol for Gene Annotation, Prediction, and Validation of Genomic Gene Expansion.
Zhang Q, Zhang ZD (2022) STAR Protoc 3(4):101692.  JOURNAL   PUBMED   REPRINT     ANALYTICAL METHOD · COMPARATIVE GENOMICS · EVOLUTIONARY GENOMICS · SOFTWARE / PIPELINE / DATABASE
ABSTRACT: Although gene expansion plays an important role in evolution, its identification remains a challenge due to potential errors in genome assembly and annotation. Here, we describe a detailed step-by-step protocol for gene annotation, prediction of genomic gene expansion, and its computational and experimental validation. Finally, we also detail steps to discover functionality of each copy of replicated genes. For complete details on the use and execution of this protocol, please refer to Zhang et al. (2021).

Genomic Expansion of Aldh1a1 Protects Beavers Against High Metabolic Aldehydes from Lipid Oxidation.
Zhang Q, Tombline G, Ablaeva J, Zhang L, Zhou X, Smith Z, Zhao Y, Xiaoli AM, Wang Z, Lin JR, Jabalameli MR, Mitra J, Nguyen N, Vijg J, Seluanov A, Gladyshev VN, Gorbunova V, Zhang ZD (2021) Cell Rep 37(6):109965.  JOURNAL   PUBMED   REPRINT     AGING · COMPARATIVE GENOMICS · EVOLUTIONARY GENOMICS
ABSTRACT: The North American beaver is an exceptionally long-lived and cancer-resistant rodent species. Here, we report the evolutionary changes in its gene coding sequences, copy numbers, and expression. We identify changes that likely increase its ability to detoxify aldehydes, enhance tumor suppression and DNA repair, and alter lipid metabolism, potentially contributing to its longevity and cancer resistance. Hpgd, a tumor suppressor gene, is uniquely duplicated in beavers among rodents, and several genes associated with tumor suppression and longevity are under positive selection in beavers. Lipid metabolism genes show positive selection signals, changes in copy numbers, or altered gene expression in beavers. Aldh1a1, encoding an enzyme for aldehydes detoxification, is particularly notable due to its massive expansion in beavers, which enhances their cellular resistance to ethanol and capacity to metabolize diverse aldehyde substrates from lipid oxidation and their woody diet. We hypothesize that the amplification of Aldh1a1 may contribute to the longevity of beavers.

Beaver and Naked Mole Rat Genomes Reveal Common Paths to Longevity.
Zhou X, Dou Q, Fan G, Zhang Q, Sanderford M, Kaya A, Johnson J, Karlsson EK, Tian X, Mikhalchenko A, Kumar S, Seluanov A, Zhang ZD, Gorbunova V, Liu X, Gladyshev VN (2020) Cell Rep 32(4):107949.  JOURNAL   PUBMED   REPRINT     AGING · COMPARATIVE GENOMICS · EVOLUTIONARY GENOMICS
ABSTRACT: Long-lived rodents have become an attractive model for the studies on aging. To understand evolutionary paths to long life, we prepare chromosome-level genome assemblies of the two longest-lived rodents, Canadian beaver (Castor canadensis) and naked mole rat (NMR, Heterocephalus glaber), which were scaffolded with in vitro proximity ligation and chromosome conformation capture data and complemented with long-read sequencing. Our comparative genomic analyses reveal that amino acid substitutions at "disease-causing" sites are widespread in the rodent genomes and that identical substitutions in long-lived rodents are associated with common adaptive phenotypes, e.g., enhanced resistance to DNA damage and cellular stress. By employing a newly developed substitution model and likelihood ratio test, we find that energy and fatty acid metabolism pathways are enriched for signals of positive selection in both long-lived rodents. Thus, the high-quality genome resource of long-lived rodents can assist in the discovery of genetic factors that control longevity and adaptive evolution.

Inducible Aging in Hydra Oligactis Implicates Sexual Reproduction, Loss of Stem Cells, and Genome Maintenance as Major Pathways.
Sun S, White RR, Fischer KE, Zhang Z, Austad SN, Vijg J (2020) Geroscience 42(4):1119-1132.  JOURNAL   PUBMED   REPRINT     AGING · CELL BIOLOGY · COMPARATIVE GENOMICS
ABSTRACT: Freshwater polyps of the genus Hydra do not age. However, temperature stress induces aging and a shift from reproduction by asexual budding to sexual gamete production in a cold-sensitive (CS) strain of H. oligactis. We sequenced the transcriptome of a male CS strain before and after this life history shift and compared changes in gene expression relative to those seen in a cold-resistant (CR) strain that does not undergo a life history shift in response to altered temperature. We found that the switch from non-aging asexual reproduction to aging and sexual reproduction involves upregulation of genes not only involved in gametogenesis but also genes involved in cellular senescence, apoptosis, and DNA repair accompanied by a downregulation of genes involved in stem cell maintenance. These results suggest that aging is a byproduct of sexual reproduction-associated cellular reprogramming and underscore the power of these H. oligactis strains to identify intrinsic mechanisms of aging.

SIRT6 Is Responsible for More Efficient DNA Double-Strand Break Repair in Long-Lived Species.
Tian X, Firsanov D, Zhang Z, Cheng Y, Luo L, Tombline G, Tan R, Simon M, Henderson S, Steffan J, Goldfarb A, Tam J, Zheng K, Cornwell A, Johnson A, Yang JN, Mao Z, Manta B, Dang W, Zhang Z, Vijg J, Wolfe A, Moody K, Kennedy BK, Bohmann D, Gladyshev VN, Seluanov A, Gorbunova V (2019) Cell 177(3):622-638.e22.  JOURNAL   PUBMED   REPRINT     AGING · CELL BIOLOGY · COMPARATIVE GENOMICS
ABSTRACT: DNA repair has been hypothesized to be a longevity determinant, but the evidence for it is based largely on accelerated aging phenotypes of DNA repair mutants. Here, using a panel of 18 rodent species with diverse lifespans, we show that more robust DNA double-strand break (DSB) repair, but not nucleotide excision repair (NER), coevolves with longevity. Evolution of NER, unlike DSB, is shaped primarily by sunlight exposure. We further show that the capacity of the SIRT6 protein to promote DSB repair accounts for a major part of the variation in DSB repair efficacy between short- and long-lived species. We dissected the molecular differences between a weak (mouse) and a strong (beaver) SIRT6 protein and identified five amino acid residues that are fully responsible for their differential activities. Our findings demonstrate that DSB repair and SIRT6 have been optimized during the evolution of longevity, which provides new targets for anti-aging interventions.

Cell Culture-Based Profiling Across Mammals Reveals DNA Repair and Metabolism as Determinants of Species Longevity.
Ma S, Upneja A, Galecki A, Tsai YM, Burant CF, Raskind S, Zhang Q, Zhang ZD, Seluanov A, Gorbunova V, Clish CB, Miller RA, Gladyshev VN (2016) Elife 5:e19130.  JOURNAL   PUBMED   REPRINT     AGING · COMPARATIVE GENOMICS
ABSTRACT: Mammalian lifespan differs by >100 fold, but the mechanisms associated with such longevity differences are not understood. Here, we conducted a study on primary skin fibroblasts isolated from 16 species of mammals and maintained under identical cell culture conditions. We developed a pipeline for obtaining species-specific ortholog sequences, profiled gene expression by RNA-seq and small molecules by metabolite profiling, and identified genes and metabolites correlating with species longevity. Cells from longer lived species up-regulated genes involved in DNA repair and glucose metabolism, down-regulated proteolysis and protein transport, and showed high levels of amino acids but low levels of lysophosphatidylcholine and lysophosphatidylethanolamine. The amino acid patterns were recapitulated by further analyses of primate and bird fibroblasts. The study suggests that fibroblast profiling captures differences in longevity across mammals at the level of global gene expression and metabolite levels and reveals pathways that define these differences.

DNA Repair in Species with Extreme Lifespan Differences.
MacRae SL, Croken MM, Calder RB, Aliper A, Milholland B, White RR, Zhavoronkov A, Gladyshev VN, Seluanov A, Gorbunova V, Zhang ZD, Vijg J (2015) Aging (Albany NY) 7(12):1171-84.  JOURNAL   PUBMED   REPRINT     AGING · COMPARATIVE GENOMICS · EVOLUTIONARY GENOMICS
ABSTRACT: Differences in DNA repair capacity have been hypothesized to underlie the great range of maximum lifespans among mammals. However, measurements of individual DNA repair activities in cells and animals have not substantiated such a relationship because utilization of repair pathways among animals--depending on habitats, anatomical characteristics, and life styles--varies greatly between mammalian species. Recent advances in high-throughput genomics, in combination with increased knowledge of the genetic pathways involved in genome maintenance, now enable a comprehensive comparison of DNA repair transcriptomes in animal species with extreme lifespan differences. Here we compare transcriptomes of liver, an organ with high oxidative metabolism and abundant spontaneous DNA damage, from humans, naked mole rats, and mice, with maximum lifespans of ~120, 30, and 3 years, respectively, with a focus on genes involved in DNA repair. The results show that the longer-lived species, human and naked mole rat, share higher expression of DNA repair genes, including core genes in several DNA repair pathways. A more systematic approach of signaling pathway analysis indicates statistically significant upregulation of several DNA repair signaling pathways in human and naked mole rat compared with mouse. The results of this present work indicate, for the first time, that DNA repair is upregulated in a major metabolic organ in long-lived humans and naked mole rats compared with short-lived mice. These results strongly suggest that DNA repair can be considered a genuine longevity assurance system.

Comparative Analysis of Genome Maintenance Genes in Naked Mole Rat, Mouse, and Human.
MacRae SL, Zhang Q, Lemetre C, Seim I, Calder RB, Hoeijmakers J, Suh Y, Gladyshev VN, Seluanov A, Gorbunova V, Vijg J, Zhang ZD (2015) Aging Cell 14(2):288-91.  JOURNAL   PUBMED   REPRINT     AGING · COMPARATIVE GENOMICS · EVOLUTIONARY GENOMICS
ABSTRACT: Genome maintenance (GM) is an essential defense system against aging and cancer, as both are characterized by increased genome instability. Here, we compared the copy number variation and mutation rate of 518 GM-associated genes in the naked mole rat (NMR), mouse, and human genomes. GM genes appeared to be strongly conserved, with copy number variation in only four genes. Interestingly, we found NMR to have a higher copy number of CEBPG, a regulator of DNA repair, and TINF2, a protector of telomere integrity. NMR, as well as human, was also found to have a lower rate of germline nucleotide substitution than the mouse. Together, the data suggest that the long-lived NMR, as well as human, has more robust GM than mouse and identifies new targets for the analysis of the exceptional longevity of the NMR.

Naked Mole-Rat Has Increased Translational Fidelity Compared with the Mouse, as Well as a Unique 28S Ribosomal RNA Cleavage.
Azpurua J, Ke Z, Chen IX, Zhang Q, Ermolenko DN, Zhang ZD, Gorbunova V, Seluanov A (2013) Proc Natl Acad Sci U S A 110(43):17350-5.  JOURNAL   PUBMED   REPRINT     COMPARATIVE GENOMICS · EVOLUTIONARY GENOMICS
ABSTRACT: The naked mole-rat (Heterocephalus glaber) is a subterranean eusocial rodent with a markedly long lifespan and resistance to tumorigenesis. Multiple data implicate modulation of protein translation in longevity. Here we report that 28S ribosomal RNA (rRNA) of the naked mole-rat is processed into two smaller fragments of unequal size. The two breakpoints are located in the 28S rRNA divergent region 6 and excise a fragment of 263 nt. The excised fragment is unique to the naked mole-rat rRNA and does not show homology to other genomic regions. Because this hidden break site could alter ribosome structure, we investigated whether translation rate and amino acid incorporation fidelity were altered. We report that naked mole-rat fibroblasts have significantly increased translational fidelity despite having comparable translation rates with mouse fibroblasts. Although we cannot directly test whether the unique 28S rRNA structure contributes to the increased fidelity of translation, we speculate that it may change the folding or dynamics of the large ribosomal subunit, altering the rate of GTP hydrolysis and/or interaction of the large subunit with tRNA during accommodation, thus affecting the fidelity of protein synthesis. In summary, our results show that naked mole-rat cells produce fewer aberrant proteins, supporting the hypothesis that the more stable proteome of the naked mole-rat contributes to its longevity.

Identification and Analysis of Unitary Pseudogenes: Historic and Contemporary Gene Losses in Humans and Other Primates.
Zhang ZD, Frankish A, Hunt T, Harrow J, Gerstein M (2010) Genome Biol 11(3):R26.  JOURNAL   PUBMED   REPRINT     COMPARATIVE GENOMICS · EVOLUTIONARY GENOMICS · PSEUDOGENE · SOFTWARE / PIPELINE / DATABASE
ABSTRACT: BACKGROUND: Unitary pseudogenes are a class of unprocessed pseudogenes without functioning counterparts in the genome. They constitute only a small fraction of annotated pseudogenes in the human genome. However, as they represent distinct functional losses over time, they shed light on the unique features of humans in primate evolution. RESULTS: We have developed a pipeline to detect human unitary pseudogenes through analyzing the global inventory of orthologs between the human genome and its mammalian relatives. We focus on gene losses along the human lineage after the divergence from rodents about 75 million years ago. In total, we identify 76 unitary pseudogenes, including previously annotated ones, and many novel ones. By comparing each of these to its functioning ortholog in other mammals, we can approximately date the creation of each unitary pseudogene (that is, the gene 'death date') and show that for our group of 76, the functional genes appear to be disabled at a fairly uniform rate throughout primate evolution - not all at once, correlated, for instance, with the 'Alu burst'. Furthermore, we identify 11 unitary pseudogenes that are polymorphic - that is, they have both nonfunctional and functional alleles currently segregating in the human population. Comparing them with their orthologs in other primates, we find that two of them are in fact pseudogenes in non-human primates, suggesting that they represent cases of a gene being resurrected in the human lineage. CONCLUSIONS: This analysis of unitary pseudogenes provides insights into the evolutionary constraints faced by different organisms and the timescales of functional gene loss in humans.

Rapid Evolution by Positive Darwinian Selection in T-Cell Antigen CD4 in Primates.
Zhang ZD, Weinstock G, Gerstein M (2008) J Mol Evol 66(5):446-56.  JOURNAL   PUBMED   REPRINT     COMPARATIVE GENOMICS · EVOLUTIONARY GENOMICS
ABSTRACT: CD4, an integral membrane glycoprotein, plays a critical role in the immune response and in the life cycle of simian and human immunodeficiency virus (SIV and HIV). Pairwise comparisons of orthologous human and mouse genes show that CD4 is evolving much faster than the majority of mammalian genes. The acceleration is too great to be attributed to a simple relaxation of the action of purifying selection alone. Here we show that the selective pressure acting on CD4 is highly variable between regions in the protein and identify codon sites under strong positive selection. We reconstruct the coding sequences for ancestral primate CD4s and model tertiary structures of all ancestral and extant sequences. Structural mapping of positively selected sites shows they distribute on the surface of the D1 domain of CD4, where the exogenous SIV gp120 protein binds. Moreover, structural models of the ancestral sequences show substantially larger variation in the interfacial electrostatic charge on CD4 and in the surface complementary between CD4 and gp120 in CD4 lineages from primates with natural SIV infections than those without. Thus, positive selection on CD4 among primates may reflect forces driven by SIV infection and could provide a link between changes in sequence and structure of CD4 during evolution and the interaction with the immunodeficiency virus.

Genomic Analysis of the Nuclear Receptor Family: New Insights into Structure, Regulation, and Evolution from the Rat Genome.
Zhang Z, Burch PE, Cooney AJ, Lanz RB, Pereira FA, Wu J, Gibbs RA, Weinstock G, Wheeler DA (2004) Genome Res 14(4):580-90.  JOURNAL   PUBMED   REPRINT   POSTER     COMPARATIVE GENOMICS · FUNCTIONAL GENOMICS
ABSTRACT: Completion of the Rattus norvegicus genome sequence enabled a global inventory and analysis of the nuclear receptors (NRs) in three mammalian species. Forty-nine NR members were found in mouse, 48 in human. Forty-seven were found in the rat, with gaps at the locations expected for the other two. Pairwise comparisons of their distribution in rat, mouse, and human identified 11 syntenic NR gene blocks, including three small clusters of two or three closely related genes, each spanning 40 kb to 1700 kb. The exon structure of the ligand-binding domain suggests that exon shuffling has played a role in the evolution of this family. An invariant splice junction in all members of the NR family except LXRbeta suggests a functional role for the intron. The ligand-binding domains of PXR and CAR are among the most divergent in the family. Their higher nucleotide substitution rates may be related to the central role played by these two NRs in the metabolism of the foreign compounds and may have resulted from limited positive selection.

Disease study (22)

Deletion Size and Background Genetic Variation Shape Congenital Heart Disease Phenotypes in 3,016 Individuals with 22q11.2 Deletion Syndrome
Lin J-R, Miller D, Luong D, Nelson T, Crowley TB, Tran OT, Thiruvahindrapuram B, Hajianpour A, Campbell L, Busa T, Heine-Suner D, Garcia-Minaur S, Fernandez L, Murphy KC, Murphy D, Hawula W, Angkustsiri K, Shashi V, Schoch K, Bearden CE, Tomita Mitchell A, Mitchell ME, Carmel M, Weizman A, Michaelovsky E, Gothelf D, van den Bree MBM, Owen MJ, Vorstman JAS, Boot E, Vingerhoets C, van Amelsvoort T, Swillen A, Breckpot J, Vermeesch JR, Devriendt K, Schneider M, Eliez S, Digilio MC, Unolt M, Putotto C, Marino B, Pontillo M, Armando M, Vicari S, Repetto GM, Kates WR, Shprintzen RJ, Gur RE, Zackai EH, Goldmuntz E, Wang T, Raj S, Emanuel BS, McDonald-McGinn DM, Scherer SC, Bassett AS, Zhang ZD, Morrow BE (2026) medRxiv 2026.02.23.26346918.  JOURNAL     DISEASE STUDY · GENETIC VARIATION
ABSTRACT: Congenital heart disease (CHD) occurs in over half of individuals with 22q11.2 deletion syndrome (22q11.2DS) and the types of lesions range from mild to severe. To determine the basis of variation in cardiac phenotypes we analyzed demographic data from 3,016 unrelated individuals with 22q11.2DS from centers in the Northeast US, Canada, Europe, South America, Israel and Australia. Most individuals in this cohort had a 3 million base pair hemizygous deletion between low copy repeat, LCR22 A-D (87.2%), while some had nested deletions. We performed multivariable mixed-effects logistic regression and uncovered significant differences between CHD phenotypes and basic demographic features. Individuals with the A-D deletion had a lower risk of persistent truncus arteriosus (OR = 0.37, 95% CI 0.18-0.75) but a higher risk of septal defects (OR = 4.7, 95% CI 1.7-12.8) compared to those with the smaller A-B deletion, suggesting distinct developmental pathways sensitive to 22q11.2 gene dosage. In addition, genome-wide genetic principal components (PCs) were associated with specific CHD subtypes, including reduced risk of pulmonary stenosis or atresia with other heart lesions (PC2; OR = 0.73, 95% CI 0.61-0.87) and increased risk of abnormal origin of the subclavian arteries (PC4; OR = 2.6, 95% CI 1.4-4.9), indicating that background genetic variation modifies heart lesion-specific susceptibility. Together, these results suggest that both deletion size and background genetic variation shape the highly variable cardiac phenotypes in 22q11.2DS.

Prevalence and Spectrum of Congenital Heart Disease in Individuals with Distal Chromosome 22q11.22-23 Deletions.
Nelson TJ, McGinn DE, Crowley TB, Rockart L, Green A, Giunta V, Tran O, Miller D, Breckpot J, Swillen A, Digilio MC, Unolt M, Putotto C, Pulvirenti F, Marino B, Emanuel BS, Zackai EH, Zhang ZD, Goldmuntz E, Boot E, Bassett AS, Morrow BE, McDonald-McGinn DM (2026) Clin Genet 109(5):859-868.  JOURNAL   PUBMED   REPRINT     DISEASE STUDY · GENETIC VARIATION
ABSTRACT: This study is aimed at determining the spectrum of congenital heart disease associated with distal 22q11.22-23 deletions flanked by low copy repeats, LCR22 D-H. We analyzed cardiology findings in 128 unrelated individuals with distal LCR22 D-H deletions. A total of 62 were newly described and 66 were derived from previous reports. We found that deletions which included LCR22-D as the proximal endpoint were the most prevalent in the cohort (104/128, 81.3%). Clinically relevant congenital heart disease was identified in 48 individuals (37.5%, 95% CI 29%-46%), which is lower than the prevalence reported for typical, proximal LCR22 A-D deletions (p = 3.7E-4), especially for conotruncal defects (13/128, 10.2%; p = 7.1E-13). Mild to moderate CHD predominated, including ventricular septal defects (22/128), bicuspid aortic valve (9/128) and mild cardiomyopathy (3/128). Persistent truncus arteriosus was the most prevalent (n = 8/13) conotruncal heart defect, but other anomalies also occurred in singleton cases. These findings support the need for cardiac evaluation in all individuals with distal 22q11.22-23 deletions, increased use of clinical genetic testing in syndromic individuals with these findings, and molecular studies in model systems. The results demonstrate that reduced gene dosage of distal 22q11.21-23, particularly within the D-E region including MAPK1 and HIC2 convey risk for CHD.

The Somatic Aneuploidy Landscape of Adult Glia Reveals 16p as a Hotspot and Differentiates Mosaicism in Normal Glia from Chromosomal Instability in Glioblastoma.
Montagna C, Albert O, Sun S, Lin JR, Lee M, Chan C, Maslov A, Ellerby L, Huttner A, Zhang Z, Vijg J (2025) Res Sq :rs.3.rs-6497851.  JOURNAL   PUBMED     DISEASE STUDY
ABSTRACT: Aneuploidy, an abnormal number of chromosomes, is a hallmark of cancer and has been proposed as an initiating event in tumorigenesis. In glioblastoma (GBM), a highly aggressive brain tumor, cells almost universally display gain of chromosome 7 and loss of chromosome 10. However, it remains unclear whether these alterations arise de novo during malignant transformation or reflect pre-existing chromosomal instability in normal brain tissue. Here, we used single-nucleus whole-genome sequencing (snWGS) on 225 NeuN-negative (non-neuronal) cortical nuclei from 12 healthy individuals and 6 GBM patients, including matched tumor cores and non-tumor brain regions. In healthy brains, approximately 15% of glial nuclei harbored somatic aneuploidies, most often involving chromosome arms, with recurrent 16p alterations detected in up to 3% of nuclei from both healthy controls and GBM non-tumor tissue. These findings establish 16p is a hotspot of structural variation in adult glia. Non-tumor regions in GBM patients closely resembled healthy controls in aneuploidy burden and chromosomal instability metrics and lacked hallmark tumor alterations. In contrast, GBM tumors exhibited significantly elevated aneuploidy (~50%), enrichment for canonical chromosomal instability-driven events, and sex-specific karyotype patterns, consistent with transformation-associated chromosomal instability. Thus, aneuploidy is a recurrent but constrained feature of normal adult glia, whereas chromosome instability and GBM-defining aneuploidies emerge only during malignant transformation.

Genetic Variants Associated with Age-Related Episodic Memory Decline Implicate Distinct Memory Pathologies.
Ali A, Milman S, Weiss EF, Gao T, Napolioni V, Barzilai N, Zhang ZD, Lin JR (2025) Alzheimers Dement 21(1):e14379.  JOURNAL   PUBMED   REPRINT     AGING · DISEASE STUDY · GENETIC VARIATION
ABSTRACT: BACKGROUND: Approximately 40% of people aged ≥ 65 experience memory loss, particularly in episodic memory. Identifying the genetic basis of episodic memory decline is crucial for uncovering its underlying causes. METHODS: We investigated common and rare genetic variants associated with episodic memory decline in 742 (632 for rare variants) Ashkenazi Jewish individuals (mean age 75) from the LonGenity study. All-atom molecular dynamics simulations were performed to uncover mechanistic insights underlying rare variants associated with episodic memory decline. RESULTS: In addition to the common polygenic risk of Alzheimer's disease, we identified and replicated rare variant associations in ITSN1 and CRHR2. Structural analyses revealed distinct memory pathologies mediated by interfacial rare coding variants such as impaired receptor activation of corticotropin releasing hormone and dysregulated L-serine synthesis. DISCUSSION: Our study uncovers novel risk loci for episodic memory decline. The identified underlying mechanisms point toward heterogenous memory pathologies mediated by rare coding variants. HIGHLIGHTS: We demonstrated the contribution of the common polygenic risk of Alzheimer's disease to episodic memory decline. We discovered and replicated two risk genes associated with episodic memory decline implicated by rare variants, were discovered and replicated. We demonstrated molecular mechanisms and potential novel memory pathologies underlying interfacial rare coding variants. Molecular dynamics simulations were performed to understand the downstream effects of risk rare coding variants.

Rare Coding Variants as Risk Modifiers of the 22q11.2 Deletion Implicate Postnatal Cortical Development in Syndromic Schizophrenia.
Lin JR, Zhao Y, Jabalameli MR, Nguyen N, Mitra J, International 22q11.DS Brain and Behavior Consortium, Swillen A, Vorstman JAS, Chow EWC, van den Bree M, Emanuel BS, Vermeesch JR, Owen MJ, Williams NM, Bassett AS, McDonald-McGinn DM, Gur RE, Bearden CE, Morrow BE, Lachman HM, Zhang ZD (2023) Mol Psychiatry 28(5):2071-2080.  JOURNAL   PUBMED   REPRINT   WEBSITE     ANALYTICAL METHOD · DISEASE STUDY · GENETIC VARIATION
ABSTRACT: 22q11.2 deletion is one of the strongest known genetic risk factors for schizophrenia. Recent whole-genome sequencing of schizophrenia cases and controls with this deletion provided an unprecedented opportunity to identify risk modifying genetic variants and investigate their contribution to the pathogenesis of schizophrenia in 22q11.2 deletion syndrome. Here, we apply a novel analytic framework that integrates gene network and phenotype data to investigate the aggregate effects of rare coding variants and identified modifier genes in this etiologically homogenous cohort (223 schizophrenia cases and 233 controls of European descent). Our analyses revealed significant additive genetic components of rare nonsynonymous variants in 110 modifier genes (adjusted P = 9.4E-04) that overall accounted for 4.6% of the variance in schizophrenia status in this cohort, of which 4.0% was independent of the common polygenic risk for schizophrenia. The modifier genes affected by rare coding variants were enriched with genes involved in synaptic function and developmental disorders. Spatiotemporal transcriptomic analyses identified an enrichment of coexpression between modifier and 22q11.2 genes in cortical brain regions from late infancy to young adulthood. Corresponding gene coexpression modules are enriched with brain-specific protein-protein interactions of SLC25A1, COMT, and PI4KA in the 22q11.2 deletion region. Overall, our study highlights the contribution of rare coding variants to the SCZ risk. They not only complement common variants in disease genetics but also pinpoint brain regions and developmental stages critical to the etiology of syndromic schizophrenia.

◸ NEWS AND VIEWS 
Unravelling Genetic Components of Longevity.
Jabalameli MR, Zhang ZD (2022) Nat Aging 2(1):5-6.  JOURNAL   PUBMED   REPRINT     AGING · DISEASE STUDY · GENETIC VARIATION · REVIEW / PERSPECTIVE
ABSTRACT: Many aging-related traits share a common genetic component. How to disentangle it from the trait-specific effects has remained largely unexplored. A new study in Nature Aging uses an analysis framework for isolating the shared genetic component in genome-wide association studies of aging-related traits and identifies genomic loci that contribute to aging.

Substance Abuse and the Risk of Severe COVID-19: Mendelian Randomization Confirms the Causal Role of Opioids but Hints a Negative Causal Effect for Cannabinoids.
Jabalameli MR, Zhang ZD (2022) Front Genet 13:1070428.  JOURNAL   PUBMED   REPRINT     DISEASE STUDY · GENETIC VARIATION
ABSTRACT: Since the start of the COVID-19 global pandemic, our understanding of the underlying disease mechanism and factors associated with the disease severity has dramatically increased. A recent study investigated the relationship between substance use disorders (SUD) and the risk of severe COVID-19 in the United States and concluded that the risk of hospitalization and death due to COVID-19 is directly correlated with substance abuse, including opioid use disorder (OUD) and cannabis use disorder (CUD). While we found this analysis fascinating, we believe this observation may be biased due to comorbidities (such as hypertension, diabetes, and cardiovascular disease) confounding the direct effect of SUD on severe COVID-19 illness. To answer this question, we sought to investigate the causal relationship between substance abuse and medication-taking history (as a proxy trait for comorbidities) with the risk of COVID-19 adverse outcomes. Our Mendelian randomization analysis confirms the causal relationship between OUD and severe COVID-19 illness but suggests an inverse causal effect for cannabinoids. Considering that COVID-19 mortality is largely attributed to disturbed immune regulation, the possible modulatory impact of cannabinoids in alleviating cytokine storms merits further investigation.

Rare Genetic Coding Variants Associated with Human Longevity and Protection Against Age-Related Diseases.
Lin JR, Sin-Chan P, Napolioni V, Torres GG, Mitra J, Zhang Q, Jabalameli MR, Wang Z, Nguyen N, Gao T, Regeneron Genetics Center, Laudes M, Görg S, Franke A, Nebel A, Greicius MD, Atzmon G, Ye K, Gorbunova V, Ladiges WC, Shuldiner AR, Niedernhofer LJ, Robbins PD, Milman S, Suh Y, Vijg J, Barzilai N, Zhang ZD (2021) Nat Aging 1(9):783-794.  JOURNAL   PUBMED   REPRINT   WEBSITE     AGING · DISEASE STUDY · GENETIC VARIATION
ABSTRACT: Extreme longevity in humans has a strong genetic component, but whether this involves genetic variation in the same longevity pathways as found in model organisms is unclear. Using whole-exome sequences of a large cohort of Ashkenazi Jewish centenarians to examine enrichment for rare coding variants, we found most longevity-associated rare coding variants converge upon conserved insulin/insulin-like growth factor 1 signaling and AMP-activating protein kinase signaling pathways. Centenarians have a number of pathogenic rare coding variants similar to control individuals, suggesting that rare variants detected in the conserved longevity pathways are protective against age-related pathology. Indeed, we detected a pro-longevity effect of rare coding variants in the Wnt signaling pathway on individuals harboring the known common risk allele APOE4. The genetic component of extreme human longevity constitutes, at least in part, rare coding variants in pathways that protect against aging, including those that control longevity in model organisms.

Enhancer Release and Retargeting Activates Disease-Susceptibility Genes.
Oh S, Shao J, Mitra J, Xiong F, D'Antonio M, Wang R, Garcia-Bassets I, Ma Q, Zhu X, Lee JH, Nair SJ, Yang F, Ohgi K, Frazer KA, Zhang ZD, Li W, Rosenfeld MG (2021) Nature 595(7869):735-740.  JOURNAL   PUBMED   REPRINT     DISEASE STUDY · FUNCTIONAL GENOMICS · GENETIC VARIATION
ABSTRACT: The functional engagement between an enhancer and its target promoter ensures precise gene transcription1. Understanding the basis of promoter choice by enhancers has important implications for health and disease. Here we report that functional loss of a preferred promoter can release its partner enhancer to loop to and activate an alternative promoter (or alternative promoters) in the neighbourhood. We refer to this target-switching process as 'enhancer release and retargeting'. Genetic deletion, motif perturbation or mutation, and dCas9-mediated CTCF tethering reveal that promoter choice by an enhancer can be determined by the binding of CTCF at promoters, in a cohesin-dependent manner-consistent with a model of 'enhancer scanning' inside the contact domain. Promoter-associated CTCF shows a lower affinity than that at chromatin domain boundaries and often lacks a preferred motif orientation or a partnering CTCF at the cognate enhancer, suggesting properties distinct from boundary CTCF. Analyses of cancer mutations, data from the GTEx project and risk loci from genome-wide association studies, together with a focused CRISPR interference screen, reveal that enhancer release and retargeting represents an overlooked mechanism that underlies the activation of disease-susceptibility genes, as exemplified by a risk locus for Parkinson's disease (NUCKS1-RAB7L1) and three loci associated with cancer (CLPTM1L-TERT, ZCCHC7-PAX5 and PVT1-MYC).

Deep Post-Gwas Analysis Identifies Potential Risk Genes and Risk Variants for Alzheimer's Disease, Providing New Insights into Its Disease Mechanisms.
Wang Z, Zhang Q, Lin JR, Jabalameli MR, Mitra J, Nguyen N, Zhang ZD (2021) Sci Rep 11(1):20511.  JOURNAL   PUBMED   REPRINT     ANALYTICAL METHOD · DISEASE STUDY · GENETIC VARIATION
ABSTRACT: Alzheimer's disease (AD) is a genetically complex, multifactorial neurodegenerative disease. It affects more than 45 million people worldwide and currently remains untreatable. Although genome-wide association studies (GWAS) have identified many AD-associated common variants, only about 25 genes are currently known to affect the risk of developing AD, despite its highly polygenic nature. Moreover, the risk variants underlying GWAS AD-association signals remain unknown. Here, we describe a deep post-GWAS analysis of AD-associated variants, using an integrated computational framework for predicting both disease genes and their risk variants. We identified 342 putative AD risk genes in 203 risk regions spanning 502 AD-associated common variants. 246 AD risk genes have not been identified as AD risk genes by previous GWAS collected in GWAS catalogs, and 115 of 342 AD risk genes are outside the risk regions, likely under the regulation of transcriptional regulatory elements contained therein. Even more significantly, for 109 AD risk genes, we predicted 150 risk variants, of both coding and regulatory (in promoters or enhancers) types, and 85 (57%) of them are supported by functional annotation. In-depth functional analyses showed that AD risk genes were overrepresented in AD-related pathways or GO terms-e.g., the complement and coagulation cascade and phosphorylation and activation of immune response-and their expression was relatively enriched in microglia, endothelia, and pericytes of the human brain. We found nine AD risk genes-e.g., IL1RAP, PMAIP1, LAMTOR4-as predictors for the prognosis of AD survival and genes such as ARL6IP5 with altered network connectivity between AD patients and normal individuals involved in AD progression. Our findings open new strategies for developing therapeutics targeting AD risk genes or risk variants to influence AD pathogenesis.

Genetics of Extreme Human Longevity to Guide Drug Discovery for Healthy Ageing.
Zhang ZD, Milman S, Lin JR, Wierbowski S, Yu H, Barzilai N, Gorbunova V, Ladiges WC, Niedernhofer LJ, Suh Y, Robbins PD, Vijg J (2020) Nat Metab 2(8):663-672.  JOURNAL   PUBMED   REPRINT     AGING · DISEASE STUDY · GENETIC VARIATION · REVIEW / PERSPECTIVE
ABSTRACT: Ageing is the greatest risk factor for most common chronic human diseases, and it therefore is a logical target for developing interventions to prevent, mitigate or reverse multiple age-related morbidities. Over the past two decades, genetic and pharmacologic interventions targeting conserved pathways of growth and metabolism have consistently led to substantial extension of the lifespan and healthspan in model organisms as diverse as nematodes, flies and mice. Recent genetic analysis of long-lived individuals is revealing common and rare variants enriched in these same conserved pathways that significantly correlate with longevity. In this Perspective, we summarize recent insights into the genetics of extreme human longevity and propose the use of this rare phenotype to identify genetic variants as molecular targets for gaining insight into the physiology of healthy ageing and the development of new therapies to extend the human healthspan.

PGA: Post-Gwas Analysis for Disease Gene Identification.
Lin JR, Jaroslawicz D, Cai Y, Zhang Q, Wang Z, Zhang ZD (2018) Bioinformatics 34(10):1786-1788.  JOURNAL   PUBMED   REPRINT   WEBSITE     ANALYTICAL METHOD · DISEASE STUDY · GENETIC VARIATION
ABSTRACT: SUMMARY: Although the genome-wide association study (GWAS) is a powerful method to identify disease-associated variants, it does not directly address the biological mechanisms underlying such genetic association signals. Here, we present PGA, a Perl- and Java-based program for post-GWAS analysis that predicts likely disease genes given a list of GWAS-reported variants. Designed with a command line interface, PGA incorporates genomic and eQTL data in identifying disease gene candidates and uses gene network and ontology data to score them based upon the strength of their relationship to the disease in question. AVAILABILITY AND IMPLEMENTATION: http://zdzlab.einstein.yu.edu/1/pga.html. CONTACT: zhengdong.zhang@einstein.yu.edu. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.

Epigenetic Alterations to Polycomb Targets Precede Malignant Transition in a Mouse Model of Breast Cancer.
Cai Y, Lin JR, Zhang Q, O'Brien K, Montagna C, Zhang ZD (2018) Sci Rep 8(1):5535.  JOURNAL   PUBMED   REPRINT     DISEASE STUDY · EPIGENOMICS
ABSTRACT: Malignant breast cancer remains a major health threat to women of all ages worldwide and epigenetic variations on DNA methylation have been widely reported in cancers of different types. We profiled DNA methylation with ERRBS (Enhanced Reduced Representation Bisulfite Sequencing) across four main stages of tumor progression in the MMTV-PyMT mouse model (hyperplasia, adenoma/mammary intraepithelial neoplasia, early carcinoma and late carcinoma), during which malignant transition occurs. We identified a large number of differentially methylated cytosines (DMCs) in tumors relative to age-matched normal mammary glands from FVB mice. Despite similarities, the methylation differences of the premalignant stages were distinct from the malignant ones. Many differentially methylated loci were preserved from the first to the last stage throughout tumor progression. Genes affected by methylation gains were enriched in Polycomb repressive complex 2 (PRC2) targets, which may present biomarkers for early diagnosis and targets for treatment.

Integrated Rare Variant-Based Risk Gene Prioritization in Disease Case-Control Sequencing Studies.
Lin JR, Zhang Q, Cai Y, Morrow BE, Zhang ZD (2017) PLoS Genet 13(12):e1007142.  JOURNAL   PUBMED   REPRINT   WEBSITE     ANALYTICAL METHOD · DISEASE STUDY · GENETIC VARIATION
ABSTRACT: Rare variants of major effect play an important role in human complex diseases and can be discovered by sequencing-based genome-wide association studies. Here, we introduce an integrated approach that combines the rare variant association test with gene network and phenotype information to identify risk genes implicated by rare variants for human complex diseases. Our data integration method follows a 'discovery-driven' strategy without relying on prior knowledge about the disease and thus maintains the unbiased character of genome-wide association studies. Simulations reveal that our method can outperform a widely-used rare variant association test method by 2 to 3 times. In a case study of a small disease cohort, we uncovered putative risk genes and the corresponding rare variants that may act as genetic modifiers of congenital heart disease in 22q11.2 deletion syndrome patients. These variants were missed by a conventional approach that relied on the rare variant association test alone.

A Neurogenetic Model for the Study of Schizophrenia Spectrum Disorders: The International 22q11.2 Deletion Syndrome Brain Behavior Consortium.
Gur RE, Bassett AS, McDonald-McGinn DM, Bearden CE, Chow E, Emanuel BS, Owen M, Swillen A, Van den Bree M, Vermeesch J, Vorstman JAS, Warren S, Lehner T, Morrow B (2017) Mol Psychiatry 22(12):1664-1672.  JOURNAL   PUBMED   REPRINT   WEBSITE     DISEASE STUDY · REVIEW / PERSPECTIVE
ABSTRACT: Rare copy number variants contribute significantly to the risk for schizophrenia, with the 22q11.2 locus consistently implicated. Individuals with the 22q11.2 deletion syndrome (22q11DS) have an estimated 25-fold increased risk for schizophrenia spectrum disorders, compared to individuals in the general population. The International 22q11DS Brain Behavior Consortium is examining this highly informative neurogenetic syndrome phenotypically and genomically. Here we detail the procedures of the effort to characterize the neuropsychiatric and neurobehavioral phenotypes associated with 22q11DS, focusing on schizophrenia and subthreshold expression of psychosis. The genomic approach includes a combination of whole-genome sequencing and genome-wide microarray technologies, allowing the investigation of all possible DNA variation and gene pathways influencing the schizophrenia-relevant phenotypic expression. A phenotypically rich data set provides a psychiatrically well-characterized sample of unprecedented size (n=1616) that informs the neurobehavioral developmental course of 22q11DS. This combined set of phenotypic and genomic data will enable hypothesis testing to elucidate the mechanisms underlying the pathogenesis of schizophrenia spectrum disorders.

Transcriptomic Dynamics of Breast Cancer Progression in the MMTV-PyMT Mouse Model.
Cai Y, Nogales-Cadenas R, Zhang Q, Lin JR, Zhang W, O'Brien K, Montagna C, Zhang ZD (2017) BMC Genomics 18(1):185.  JOURNAL   PUBMED   REPRINT     DISEASE STUDY · FUNCTIONAL GENOMICS
ABSTRACT: BACKGROUND: Malignant breast cancer with complex molecular mechanisms of progression and metastasis remains a leading cause of death in women. To improve diagnosis and drug development, it is critical to identify panels of genes and molecular pathways involved in tumor progression and malignant transition. Using the PyMT mouse, a genetically engineered mouse model that has been widely used to study human breast cancer, we profiled and analyzed gene expression from four distinct stages of tumor progression (hyperplasia, adenoma/MIN, early carcinoma and late carcinoma) during which malignant transition occurs. RESULTS: We found remarkable expression similarity among the four stages, meaning genes altered in the later stages showed trace in the beginning of tumor progression. We identified a large number of differentially expressed genes in PyMT samples of all stages compared with normal mammary glands, enriched in cancer-related pathways. Using co-expression networks, we found panels of genes as signature modules with some hub genes that predict metastatic risk. Time-course analysis revealed genes with expression transition when shifting to malignant stages. These may provide additional insight into the molecular mechanisms beyond pathways. CONCLUSIONS: Thus, in this study, our various analyses with the PyMT mouse model shed new light on transcriptomic dynamics during breast cancer malignant progression.

Network Analysis of Mitonuclear GWAS Reveals Functional Networks and Tissue Expression Profiles of Disease-Associated Genes.
Johnson SC, Gonzalez B, Zhang Q, Milholland B, Zhang Z, Suh Y (2017) Hum Genet 136(1):55-65.  JOURNAL   PUBMED   REPRINT     DISEASE STUDY · SYSTEMS BIOLOGY
ABSTRACT: While mitochondria have been linked to many human diseases through genetic association and functional studies, the precise role of mitochondria in specific pathologies, such as cardiovascular, neurodegenerative, and metabolic diseases, is often unclear. Here, we take advantage of the catalog of human genome-wide associations, whole-genome tissue expression and expression quantitative trait loci datasets, and annotated mitochondrial proteome databases to examine the role of common genetic variation in mitonuclear genes in human disease. Through pathway-based analysis we identified distinct functional pathways and tissue expression profiles associated with each of the major human diseases. Among our most striking findings, we observe that mitonuclear genes associated with cancer are broadly expressed among human tissues and largely represent one functional process, intrinsic apoptosis, while mitonuclear genes associated with other diseases, such as neurodegenerative and metabolic diseases, show tissue-specific expression profiles and are associated with unique functional pathways. These results provide new insight into human diseases using unbiased genome-wide approaches.

◸ JOURNAL HIGHLIGHTED ARTICLE 
Integrated Post-Gwas Analysis Sheds New Light on the Disease Mechanisms of Schizophrenia.
Lin JR, Cai Y, Zhang Q, Zhang W, Nogales-Cadenas R, Zhang ZD (2016) Genetics 204(4):1587-1600.  JOURNAL   PUBMED   REPRINT   WEBSITE     ANALYTICAL METHOD · DISEASE STUDY
ABSTRACT: Schizophrenia is a severe mental disorder with a large genetic component. Recent genome-wide association studies (GWAS) have identified many schizophrenia-associated common variants. For most of the reported associations, however, the underlying biological mechanisms are not clear. The critical first step for their elucidation is to identify the most likely disease genes as the source of the association signals. Here, we describe a general computational framework of post-GWAS analysis for complex disease gene prioritization. We identify 132 putative schizophrenia risk genes in 76 risk regions spanning 120 schizophrenia-associated common variants, 78 of which have not been recognized as schizophrenia disease genes by previous GWAS. Even more significantly, 29 of them are outside the risk regions, likely under regulation of transcriptional regulatory elements contained therein. These putative schizophrenia risk genes are transcriptionally active in both brain and the immune system, and highly enriched among cellular pathways, consistent with leading pathophysiological hypotheses about the pathogenesis of schizophrenia. With their involvement in distinct biological processes, these putative schizophrenia risk genes, with different association strengths, show distinctive temporal expression patterns, and play specific biological roles during brain development.

MicroRNA Expression and Gene Regulation Drive Breast Cancer Progression and Metastasis in PyMT Mice.
Nogales-Cadenas R, Cai Y, Lin JR, Zhang Q, Zhang W, Montagna C, Zhang ZD (2016) Breast Cancer Res 18(1):75.  JOURNAL   PUBMED   REPRINT   WEBSITE     DISEASE STUDY · FUNCTIONAL GENOMICS
ABSTRACT: BACKGROUND: MicroRNAs (miRNAs) are small non-coding RNA molecules of about 22 nucleotides which function to silence the expression of their target genes. Numerous studies have shown that miRNAs are not only key regulators in important cellular processes but are also drivers in the development of many diseases, especially cancer. Estrogen receptor positive luminal B is the second most common but the least studied subtype of breast cancer. Only a few studies have examined the expression profiles of miRNAs in luminal B breast cancer, and their regulatory roles in cancer progression have yet to be investigated. METHODS: In this study, using polyoma middle T antigen (PyMT) mice, a widely used luminal B breast cancer model, we profiled microRNA (miRNA) expression at four time points that represent different key developmental stages of cancer progression. We considered the expression of both miRNAs and messenger RNAs (mRNAs) at these time points to improve the identification of regulatory targets of miRNAs. By combining gene functional and pathway annotation with miRNA-mRNA interactions, we created a PyMT-specific tripartite miRNA-mRNA-pathway network and identified novel functional regulatory programs (FRPs). RESULTS: We identified 151 differentially expressed miRNAs with a strict dual nature of either upregulation or downregulation during the whole course of disease progression. Among 82 newly discovered breast-cancer-related miRNAs, 35 can potentially regulate 271 protein-coding genes based on their sequence complementarity and expression profiles. We also identified miRNA-mRNA regulatory modules driving specific cancer-related biological processes. CONCLUSIONS: In this study we profiled the expression of miRNAs during breast cancer progression in the PyMT mouse model. By integrating miRNA and mRNA expression profiles, we identified differentially expressed miRNAs and their target genes involved in several hallmarks of cancer. We applied a novel clustering method to an annotated miRNA-mRNA regulatory network and identified network modules involved in specific cancer-related biological processes.

Prioritization of Schizophrenia Risk Genes by a Network-Regularized Logistic Regression Method.
Zhang W, Lin JR, Nogales-Cadenas R, Zhang Q, Cai Y, Zhang ZD (2016) Lecture Notes in Bioinformatics 9565:434-445.  JOURNAL   REPRINT     ANALYTICAL METHOD · DISEASE STUDY
ABSTRACT: Schizophrenia (SCZ) is a severe mental disorder with a large genetic component. While recent large-scale microarray- and sequencing-based genome wide association studies have made significant progress toward finding SCZ risk variants and genes of subtle effect, the interactions among them were not considered in those studies. Using a protein-protein interaction network both in our regression model and to generate a SCZ gene subnetwork, we developed an analytical framework with Logit-Lapnet, the graphical Laplacian-regularized logistic regression, for whole exome sequencing (WES) data analysis to detect SCZ gene subnetworks. Using simulated data from sequencing-based association study, we compared the performances of Logit-Lapnet with other logistic regression (LR)-based models. We use Logit-Lapnet to prioritize genes according to their coefficients and select top-ranked genes as seeds to generate the gene sub-network that is associated to SCZ. The comparison demonstrated not only the applicability but also better performance of Logit-Lapnet to score disease risk genes using sequencing-based association data. We applied our method to SCZ whole exome sequencing data and selected top-ranked risk genes, the majority of which are either known SCZ genes or genes potentially associated with SCZ. We then used the seed genes to construct SCZ gene subnetworks. This result demonstrates that by rank-ing gene according to their disease contributions our method scores and thus prioritiz-es disease risk genes for further investigation. An implementation of our approach in MATLAB is freely available for download at:
http://zdzlab.einstein.yu.edu/1/publications/LapNet-MATLAB.zip.

Whole-Genome Sequencing and Integrative Genomic Analysis Approach on Two 22q11.2 Deletion Syndrome Family Trios for Genotype to Phenotype Correlations.
Chung JH, Cai J, Suskin BG, Zhang Z, Coleman K, Morrow BE (2015) Hum Mutat 36(8):797-807.  JOURNAL   PUBMED   REPRINT     DISEASE STUDY · GENETIC VARIATION · GENOMIC SEQUENCING
ABSTRACT: The 22q11.2 deletion syndrome (22q11DS) affects 1:4,000 live births and presents with highly variable phenotype expressivity. In this study, we developed an analytical approach utilizing whole-genome sequencing (WGS) and integrative analysis to discover genetic modifiers. Our pipeline combined available tools in order to prioritize rare, predicted deleterious, coding and noncoding single-nucleotide variants (SNVs), and insertion/deletions from WGS. We sequenced two unrelated probands with 22q11DS, with contrasting clinical findings, and their unaffected parents. Proband P1 had cognitive impairment, psychotic episodes, anxiety, and tetralogy of Fallot (TOF), whereas proband P2 had juvenile rheumatoid arthritis but no other major clinical findings. In P1, we identified common variants in COMT and PRODH on 22q11.2 as well as rare potentially deleterious DNA variants in other behavioral/neurocognitive genes. We also identified a de novo SNV in ADNP2 (NM_014913.3:c.2243G>C), encoding a neuroprotective protein that may be involved in behavioral disorders. In P2, we identified a novel nonsynonymous SNV in ZFPM2 (NM_012082.3:c.1576C>T), a known causative gene for TOF, which may act as a protective variant downstream of TBX1, haploinsufficiency of which is responsible for congenital heart disease in individuals with 22q11DS.

Mosaic Epigenetic Dysregulation of Ectodermal Cells in Autism Spectrum Disorder.
Berko ER, Suzuki M, Beren F, Lemetre C, Alaimo CM, Calder RB, Ballaban-Gil K, Gounder B, Kampf K, Kirschen J, Maqbool SB, Momin Z, Reynolds DM, Russo N, Shulman L, Stasiek E, Tozour J, Valicenti-McDermott M, Wang S, Abrahams BS, Hargitai J, Inbar D, Zhang Z, Buxbaum JD, Molholm S, Foxe JJ, Marion RW, Auton A, Greally JM (2014) PLoS Genet 10(5):e1004402.  JOURNAL   PUBMED   REPRINT     DISEASE STUDY · EPIGENOMICS
ABSTRACT: DNA mutational events are increasingly being identified in autism spectrum disorder (ASD), but the potential additional role of dysregulation of the epigenome in the pathogenesis of the condition remains unclear. The epigenome is of interest as a possible mediator of environmental effects during development, encoding a cellular memory reflected by altered function of progeny cells. Advanced maternal age (AMA) is associated with an increased risk of having a child with ASD for reasons that are not understood. To explore whether AMA involves covert aneuploidy or epigenetic dysregulation leading to ASD in the offspring, we tested a homogeneous ectodermal cell type from 47 individuals with ASD compared with 48 typically developing (TD) controls born to mothers of ≥35 years, using a quantitative genome-wide DNA methylation assay. We show that DNA methylation patterns are dysregulated in ectodermal cells in these individuals, having accounted for confounding effects due to subject age, sex and ancestral haplotype. We did not find mosaic aneuploidy or copy number variability to occur at differentially-methylated regions in these subjects. Of note, the loci with distinctive DNA methylation were found at genes expressed in the brain and encoding protein products significantly enriched for interactions with those produced by known ASD-causing genes, representing a perturbation by epigenomic dysregulation of the same networks compromised by DNA mutational mechanisms. The results indicate the presence of a mosaic subpopulation of epigenetically-dysregulated, ectodermally-derived cells in subjects with ASD. The epigenetic dysregulation observed in these ASD subjects born to older mothers may be associated with aging parental gametes, environmental influences during embryogenesis or could be the consequence of mutations of the chromatin regulatory genes increasingly implicated in ASD. The results indicate that epigenetic dysregulatory mechanisms may complement and interact with DNA mutations in the pathogenesis of the disorder.

Epigenomics (4)

Global, Integrated Analysis of Methylomes and Transcriptomes from Laser Capture Microdissected Bronchial and Alveolar Cells in Human Lung.
Dong X, Shi M, Lee M, Toro R, Gravina S, Han W, Yasuda S, Wang T, Zhang Z, Vijg J, Suh Y, Spivack SD (2018) Epigenetics 13(3):264-274.  JOURNAL   PUBMED   REPRINT     CELL BIOLOGY · EPIGENOMICS
ABSTRACT: Gene regulatory analysis of highly diverse human tissues in vivo is essentially constrained by the challenge of performing genome-wide, integrated epigenetic and transcriptomic analysis in small selected groups of specific cell types. Here we performed genome-wide bisulfite sequencing and RNA-seq from the same small groups of bronchial and alveolar cells isolated by laser capture microdissection from flash-frozen lung tissue of 12 donors and their peripheral blood T cells. Methylation and transcriptome patterns differed between alveolar and bronchial cells, while each of these epithelia showed more differences from mesodermally-derived T cells. Differentially methylated regions (DMRs) between alveolar and bronchial cells tended to locate at regulatory regions affecting promoters of 4,350 genes. A large number of pathways enriched for these DMRs including GTPase signal transduction, cell death, and skeletal muscle. Similar patterns of transcriptome differences were observed: 4,108 differentially expressed genes (DEGs) enriched in GTPase signal transduction, inflammation, cilium assembly, and others. Prioritizing using DMR-DEG regulatory network, we highlighted genes, e.g., ETS1, PPARG, and RXRG, at prominent alveolar vs. bronchial cell discriminant nodes. Our results show that multi-omic analysis of small, highly specific cells is feasible and yields unique physiologic loci distinguishing human lung cell types in situ.

Epigenetic Alterations to Polycomb Targets Precede Malignant Transition in a Mouse Model of Breast Cancer.
Cai Y, Lin JR, Zhang Q, O'Brien K, Montagna C, Zhang ZD (2018) Sci Rep 8(1):5535.  JOURNAL   PUBMED   REPRINT     DISEASE STUDY · EPIGENOMICS
ABSTRACT: Malignant breast cancer remains a major health threat to women of all ages worldwide and epigenetic variations on DNA methylation have been widely reported in cancers of different types. We profiled DNA methylation with ERRBS (Enhanced Reduced Representation Bisulfite Sequencing) across four main stages of tumor progression in the MMTV-PyMT mouse model (hyperplasia, adenoma/mammary intraepithelial neoplasia, early carcinoma and late carcinoma), during which malignant transition occurs. We identified a large number of differentially methylated cytosines (DMCs) in tumors relative to age-matched normal mammary glands from FVB mice. Despite similarities, the methylation differences of the premalignant stages were distinct from the malignant ones. Many differentially methylated loci were preserved from the first to the last stage throughout tumor progression. Genes affected by methylation gains were enriched in Polycomb repressive complex 2 (PRC2) targets, which may present biomarkers for early diagnosis and targets for treatment.

RNA:DNA Hybrids in the Human Genome Have Distinctive Nucleotide Characteristics, Chromatin Composition, and Transcriptional Relationships.
Nadel J, Athanasiadou R, Lemetre C, Wijetunga NA, Ó Broin P, Sato H, Zhang Z, Jeddeloh J, Montagna C, Golden A, Seoighe C, Greally JM (2015) Epigenetics Chromatin 8:46.  JOURNAL   PUBMED   REPRINT     EPIGENOMICS · FUNCTIONAL GENOMICS
ABSTRACT: BACKGROUND: RNA:DNA hybrids represent a non-canonical nucleic acid structure that has been associated with a range of human diseases and potential transcriptional regulatory functions. Mapping of RNA:DNA hybrids in human cells reveals them to have a number of characteristics that give insights into their functions. RESULTS: We find RNA:DNA hybrids to occupy millions of base pairs in the human genome. A directional sequencing approach shows the RNA component of the RNA:DNA hybrid to be purine-rich, indicating a thermodynamic contribution to their in vivo stability. The RNA:DNA hybrids are enriched at loci with decreased DNA methylation and increased DNase hypersensitivity, and within larger domains with characteristics of heterochromatin formation, indicating potential transcriptional regulatory properties. Mass spectrometry studies of chromatin at RNA:DNA hybrids shows the presence of the ILF2 and ILF3 transcription factors, supporting a model of certain transcription factors binding preferentially to the RNA:DNA conformation. CONCLUSIONS: Overall, there is little to indicate a dependence for RNA:DNA hybrids forming co-transcriptionally, with results from the ribosomal DNA repeat unit instead supporting the intriguing model of RNA generating these structures in trans. The results of the study indicate heterogeneous functions of these genomic elements and new insights into their formation and stability in vivo.

Mosaic Epigenetic Dysregulation of Ectodermal Cells in Autism Spectrum Disorder.
Berko ER, Suzuki M, Beren F, Lemetre C, Alaimo CM, Calder RB, Ballaban-Gil K, Gounder B, Kampf K, Kirschen J, Maqbool SB, Momin Z, Reynolds DM, Russo N, Shulman L, Stasiek E, Tozour J, Valicenti-McDermott M, Wang S, Abrahams BS, Hargitai J, Inbar D, Zhang Z, Buxbaum JD, Molholm S, Foxe JJ, Marion RW, Auton A, Greally JM (2014) PLoS Genet 10(5):e1004402.  JOURNAL   PUBMED   REPRINT     DISEASE STUDY · EPIGENOMICS
ABSTRACT: DNA mutational events are increasingly being identified in autism spectrum disorder (ASD), but the potential additional role of dysregulation of the epigenome in the pathogenesis of the condition remains unclear. The epigenome is of interest as a possible mediator of environmental effects during development, encoding a cellular memory reflected by altered function of progeny cells. Advanced maternal age (AMA) is associated with an increased risk of having a child with ASD for reasons that are not understood. To explore whether AMA involves covert aneuploidy or epigenetic dysregulation leading to ASD in the offspring, we tested a homogeneous ectodermal cell type from 47 individuals with ASD compared with 48 typically developing (TD) controls born to mothers of ≥35 years, using a quantitative genome-wide DNA methylation assay. We show that DNA methylation patterns are dysregulated in ectodermal cells in these individuals, having accounted for confounding effects due to subject age, sex and ancestral haplotype. We did not find mosaic aneuploidy or copy number variability to occur at differentially-methylated regions in these subjects. Of note, the loci with distinctive DNA methylation were found at genes expressed in the brain and encoding protein products significantly enriched for interactions with those produced by known ASD-causing genes, representing a perturbation by epigenomic dysregulation of the same networks compromised by DNA mutational mechanisms. The results indicate the presence of a mosaic subpopulation of epigenetically-dysregulated, ectodermally-derived cells in subjects with ASD. The epigenetic dysregulation observed in these ASD subjects born to older mothers may be associated with aging parental gametes, environmental influences during embryogenesis or could be the consequence of mutations of the chromatin regulatory genes increasingly implicated in ASD. The results indicate that epigenetic dysregulatory mechanisms may complement and interact with DNA mutations in the pathogenesis of the disorder.

Evolutionary genomics (8)

Protocol for Gene Annotation, Prediction, and Validation of Genomic Gene Expansion.
Zhang Q, Zhang ZD (2022) STAR Protoc 3(4):101692.  JOURNAL   PUBMED   REPRINT     ANALYTICAL METHOD · COMPARATIVE GENOMICS · EVOLUTIONARY GENOMICS · SOFTWARE / PIPELINE / DATABASE
ABSTRACT: Although gene expansion plays an important role in evolution, its identification remains a challenge due to potential errors in genome assembly and annotation. Here, we describe a detailed step-by-step protocol for gene annotation, prediction of genomic gene expansion, and its computational and experimental validation. Finally, we also detail steps to discover functionality of each copy of replicated genes. For complete details on the use and execution of this protocol, please refer to Zhang et al. (2021).

Genomic Expansion of Aldh1a1 Protects Beavers Against High Metabolic Aldehydes from Lipid Oxidation.
Zhang Q, Tombline G, Ablaeva J, Zhang L, Zhou X, Smith Z, Zhao Y, Xiaoli AM, Wang Z, Lin JR, Jabalameli MR, Mitra J, Nguyen N, Vijg J, Seluanov A, Gladyshev VN, Gorbunova V, Zhang ZD (2021) Cell Rep 37(6):109965.  JOURNAL   PUBMED   REPRINT     AGING · COMPARATIVE GENOMICS · EVOLUTIONARY GENOMICS
ABSTRACT: The North American beaver is an exceptionally long-lived and cancer-resistant rodent species. Here, we report the evolutionary changes in its gene coding sequences, copy numbers, and expression. We identify changes that likely increase its ability to detoxify aldehydes, enhance tumor suppression and DNA repair, and alter lipid metabolism, potentially contributing to its longevity and cancer resistance. Hpgd, a tumor suppressor gene, is uniquely duplicated in beavers among rodents, and several genes associated with tumor suppression and longevity are under positive selection in beavers. Lipid metabolism genes show positive selection signals, changes in copy numbers, or altered gene expression in beavers. Aldh1a1, encoding an enzyme for aldehydes detoxification, is particularly notable due to its massive expansion in beavers, which enhances their cellular resistance to ethanol and capacity to metabolize diverse aldehyde substrates from lipid oxidation and their woody diet. We hypothesize that the amplification of Aldh1a1 may contribute to the longevity of beavers.

Beaver and Naked Mole Rat Genomes Reveal Common Paths to Longevity.
Zhou X, Dou Q, Fan G, Zhang Q, Sanderford M, Kaya A, Johnson J, Karlsson EK, Tian X, Mikhalchenko A, Kumar S, Seluanov A, Zhang ZD, Gorbunova V, Liu X, Gladyshev VN (2020) Cell Rep 32(4):107949.  JOURNAL   PUBMED   REPRINT     AGING · COMPARATIVE GENOMICS · EVOLUTIONARY GENOMICS
ABSTRACT: Long-lived rodents have become an attractive model for the studies on aging. To understand evolutionary paths to long life, we prepare chromosome-level genome assemblies of the two longest-lived rodents, Canadian beaver (Castor canadensis) and naked mole rat (NMR, Heterocephalus glaber), which were scaffolded with in vitro proximity ligation and chromosome conformation capture data and complemented with long-read sequencing. Our comparative genomic analyses reveal that amino acid substitutions at "disease-causing" sites are widespread in the rodent genomes and that identical substitutions in long-lived rodents are associated with common adaptive phenotypes, e.g., enhanced resistance to DNA damage and cellular stress. By employing a newly developed substitution model and likelihood ratio test, we find that energy and fatty acid metabolism pathways are enriched for signals of positive selection in both long-lived rodents. Thus, the high-quality genome resource of long-lived rodents can assist in the discovery of genetic factors that control longevity and adaptive evolution.

DNA Repair in Species with Extreme Lifespan Differences.
MacRae SL, Croken MM, Calder RB, Aliper A, Milholland B, White RR, Zhavoronkov A, Gladyshev VN, Seluanov A, Gorbunova V, Zhang ZD, Vijg J (2015) Aging (Albany NY) 7(12):1171-84.  JOURNAL   PUBMED   REPRINT     AGING · COMPARATIVE GENOMICS · EVOLUTIONARY GENOMICS
ABSTRACT: Differences in DNA repair capacity have been hypothesized to underlie the great range of maximum lifespans among mammals. However, measurements of individual DNA repair activities in cells and animals have not substantiated such a relationship because utilization of repair pathways among animals--depending on habitats, anatomical characteristics, and life styles--varies greatly between mammalian species. Recent advances in high-throughput genomics, in combination with increased knowledge of the genetic pathways involved in genome maintenance, now enable a comprehensive comparison of DNA repair transcriptomes in animal species with extreme lifespan differences. Here we compare transcriptomes of liver, an organ with high oxidative metabolism and abundant spontaneous DNA damage, from humans, naked mole rats, and mice, with maximum lifespans of ~120, 30, and 3 years, respectively, with a focus on genes involved in DNA repair. The results show that the longer-lived species, human and naked mole rat, share higher expression of DNA repair genes, including core genes in several DNA repair pathways. A more systematic approach of signaling pathway analysis indicates statistically significant upregulation of several DNA repair signaling pathways in human and naked mole rat compared with mouse. The results of this present work indicate, for the first time, that DNA repair is upregulated in a major metabolic organ in long-lived humans and naked mole rats compared with short-lived mice. These results strongly suggest that DNA repair can be considered a genuine longevity assurance system.

Comparative Analysis of Genome Maintenance Genes in Naked Mole Rat, Mouse, and Human.
MacRae SL, Zhang Q, Lemetre C, Seim I, Calder RB, Hoeijmakers J, Suh Y, Gladyshev VN, Seluanov A, Gorbunova V, Vijg J, Zhang ZD (2015) Aging Cell 14(2):288-91.  JOURNAL   PUBMED   REPRINT     AGING · COMPARATIVE GENOMICS · EVOLUTIONARY GENOMICS
ABSTRACT: Genome maintenance (GM) is an essential defense system against aging and cancer, as both are characterized by increased genome instability. Here, we compared the copy number variation and mutation rate of 518 GM-associated genes in the naked mole rat (NMR), mouse, and human genomes. GM genes appeared to be strongly conserved, with copy number variation in only four genes. Interestingly, we found NMR to have a higher copy number of CEBPG, a regulator of DNA repair, and TINF2, a protector of telomere integrity. NMR, as well as human, was also found to have a lower rate of germline nucleotide substitution than the mouse. Together, the data suggest that the long-lived NMR, as well as human, has more robust GM than mouse and identifies new targets for the analysis of the exceptional longevity of the NMR.

Naked Mole-Rat Has Increased Translational Fidelity Compared with the Mouse, as Well as a Unique 28S Ribosomal RNA Cleavage.
Azpurua J, Ke Z, Chen IX, Zhang Q, Ermolenko DN, Zhang ZD, Gorbunova V, Seluanov A (2013) Proc Natl Acad Sci U S A 110(43):17350-5.  JOURNAL   PUBMED   REPRINT     COMPARATIVE GENOMICS · EVOLUTIONARY GENOMICS
ABSTRACT: The naked mole-rat (Heterocephalus glaber) is a subterranean eusocial rodent with a markedly long lifespan and resistance to tumorigenesis. Multiple data implicate modulation of protein translation in longevity. Here we report that 28S ribosomal RNA (rRNA) of the naked mole-rat is processed into two smaller fragments of unequal size. The two breakpoints are located in the 28S rRNA divergent region 6 and excise a fragment of 263 nt. The excised fragment is unique to the naked mole-rat rRNA and does not show homology to other genomic regions. Because this hidden break site could alter ribosome structure, we investigated whether translation rate and amino acid incorporation fidelity were altered. We report that naked mole-rat fibroblasts have significantly increased translational fidelity despite having comparable translation rates with mouse fibroblasts. Although we cannot directly test whether the unique 28S rRNA structure contributes to the increased fidelity of translation, we speculate that it may change the folding or dynamics of the large ribosomal subunit, altering the rate of GTP hydrolysis and/or interaction of the large subunit with tRNA during accommodation, thus affecting the fidelity of protein synthesis. In summary, our results show that naked mole-rat cells produce fewer aberrant proteins, supporting the hypothesis that the more stable proteome of the naked mole-rat contributes to its longevity.

Identification and Analysis of Unitary Pseudogenes: Historic and Contemporary Gene Losses in Humans and Other Primates.
Zhang ZD, Frankish A, Hunt T, Harrow J, Gerstein M (2010) Genome Biol 11(3):R26.  JOURNAL   PUBMED   REPRINT     COMPARATIVE GENOMICS · EVOLUTIONARY GENOMICS · PSEUDOGENE · SOFTWARE / PIPELINE / DATABASE
ABSTRACT: BACKGROUND: Unitary pseudogenes are a class of unprocessed pseudogenes without functioning counterparts in the genome. They constitute only a small fraction of annotated pseudogenes in the human genome. However, as they represent distinct functional losses over time, they shed light on the unique features of humans in primate evolution. RESULTS: We have developed a pipeline to detect human unitary pseudogenes through analyzing the global inventory of orthologs between the human genome and its mammalian relatives. We focus on gene losses along the human lineage after the divergence from rodents about 75 million years ago. In total, we identify 76 unitary pseudogenes, including previously annotated ones, and many novel ones. By comparing each of these to its functioning ortholog in other mammals, we can approximately date the creation of each unitary pseudogene (that is, the gene 'death date') and show that for our group of 76, the functional genes appear to be disabled at a fairly uniform rate throughout primate evolution - not all at once, correlated, for instance, with the 'Alu burst'. Furthermore, we identify 11 unitary pseudogenes that are polymorphic - that is, they have both nonfunctional and functional alleles currently segregating in the human population. Comparing them with their orthologs in other primates, we find that two of them are in fact pseudogenes in non-human primates, suggesting that they represent cases of a gene being resurrected in the human lineage. CONCLUSIONS: This analysis of unitary pseudogenes provides insights into the evolutionary constraints faced by different organisms and the timescales of functional gene loss in humans.

Rapid Evolution by Positive Darwinian Selection in T-Cell Antigen CD4 in Primates.
Zhang ZD, Weinstock G, Gerstein M (2008) J Mol Evol 66(5):446-56.  JOURNAL   PUBMED   REPRINT     COMPARATIVE GENOMICS · EVOLUTIONARY GENOMICS
ABSTRACT: CD4, an integral membrane glycoprotein, plays a critical role in the immune response and in the life cycle of simian and human immunodeficiency virus (SIV and HIV). Pairwise comparisons of orthologous human and mouse genes show that CD4 is evolving much faster than the majority of mammalian genes. The acceleration is too great to be attributed to a simple relaxation of the action of purifying selection alone. Here we show that the selective pressure acting on CD4 is highly variable between regions in the protein and identify codon sites under strong positive selection. We reconstruct the coding sequences for ancestral primate CD4s and model tertiary structures of all ancestral and extant sequences. Structural mapping of positively selected sites shows they distribute on the surface of the D1 domain of CD4, where the exogenous SIV gp120 protein binds. Moreover, structural models of the ancestral sequences show substantially larger variation in the interfacial electrostatic charge on CD4 and in the surface complementary between CD4 and gp120 in CD4 lineages from primates with natural SIV infections than those without. Thus, positive selection on CD4 among primates may reflect forces driven by SIV infection and could provide a link between changes in sequence and structure of CD4 during evolution and the interaction with the immunodeficiency virus.

Functional genomics (18)

Enhancer Release and Retargeting Activates Disease-Susceptibility Genes.
Oh S, Shao J, Mitra J, Xiong F, D'Antonio M, Wang R, Garcia-Bassets I, Ma Q, Zhu X, Lee JH, Nair SJ, Yang F, Ohgi K, Frazer KA, Zhang ZD, Li W, Rosenfeld MG (2021) Nature 595(7869):735-740.  JOURNAL   PUBMED   REPRINT     DISEASE STUDY · FUNCTIONAL GENOMICS · GENETIC VARIATION
ABSTRACT: The functional engagement between an enhancer and its target promoter ensures precise gene transcription1. Understanding the basis of promoter choice by enhancers has important implications for health and disease. Here we report that functional loss of a preferred promoter can release its partner enhancer to loop to and activate an alternative promoter (or alternative promoters) in the neighbourhood. We refer to this target-switching process as 'enhancer release and retargeting'. Genetic deletion, motif perturbation or mutation, and dCas9-mediated CTCF tethering reveal that promoter choice by an enhancer can be determined by the binding of CTCF at promoters, in a cohesin-dependent manner-consistent with a model of 'enhancer scanning' inside the contact domain. Promoter-associated CTCF shows a lower affinity than that at chromatin domain boundaries and often lacks a preferred motif orientation or a partnering CTCF at the cognate enhancer, suggesting properties distinct from boundary CTCF. Analyses of cancer mutations, data from the GTEx project and risk loci from genome-wide association studies, together with a focused CRISPR interference screen, reveal that enhancer release and retargeting represents an overlooked mechanism that underlies the activation of disease-susceptibility genes, as exemplified by a risk locus for Parkinson's disease (NUCKS1-RAB7L1) and three loci associated with cancer (CLPTM1L-TERT, ZCCHC7-PAX5 and PVT1-MYC).

Transposon-Triggered Innate Immune Response Confers Cancer Resistance to the Blind Mole Rat.
Zhao Y, Oreskovic E, Zhang Q, Lu Q, Gilman A, Lin YS, He J, Zheng Z, Lu JY, Lee J, Ke Z, Ablaeva J, Sweet MJ, Horvath S, Zhang Z, Nevo E, Seluanov A, Gorbunova V (2021) Nat Immunol 22(10):1219-1230.  JOURNAL   PUBMED   REPRINT     FUNCTIONAL GENOMICS
ABSTRACT: Blind mole rats (BMRs) are small rodents, characterized by an exceptionally long lifespan (>21 years) and resistance to both spontaneous and induced tumorigenesis. Here we report that cancer resistance in the BMR is mediated by retrotransposable elements (RTEs). Cells and tissues of BMRs express very low levels of DNA methyltransferase 1. Following cell hyperplasia, the BMR genome DNA loses methylation, resulting in the activation of RTEs. Upregulated RTEs form cytoplasmic RNA-DNA hybrids, which activate the cGAS-STING pathway to induce cell death. Although this mechanism is enhanced in the BMR, we show that it functions in mice and humans. We propose that RTEs were co-opted to serve as tumor suppressors that monitor cell proliferation and are activated in premalignant cells to trigger cell death via activation of the innate immune response. Activation of RTEs is a double-edged sword, serving as a tumor suppressor but contributing to aging in late life via the induction of sterile inflammation.

HEDD: Human Enhancer Disease Database.
Wang Z, Zhang Q, Zhang W, Lin JR, Cai Y, Mitra J, Zhang ZD (2018) Nucleic Acids Res 46(D1):D113-D120.  JOURNAL   PUBMED   REPRINT   WEBSITE     FUNCTIONAL GENOMICS · SOFTWARE / PIPELINE / DATABASE
ABSTRACT: Enhancers, as specialized genomic cis-regulatory elements, activate transcription of their target genes and play an important role in pathogenesis of many human complex diseases. Despite recent systematic identification of them in the human genome, currently there is an urgent need for comprehensive annotation databases of human enhancers with a focus on their disease connections. In response, we built the Human Enhancer Disease Database (HEDD) to facilitate studies of enhancers and their potential roles in human complex diseases. HEDD currently provides comprehensive genomic information for ∼2.8 million human enhancers identified by ENCODE, FANTOM5 and RoadMap with disease association scores based on enhancer-gene and gene-disease connections. It also provides Web-based analytical tools to visualize enhancer networks and score enhancers given a set of selected genes in a specific gene network. HEDD is freely accessible at http://zdzlab.einstein.yu.edu/1/hedd.php.

Transcriptomic Dynamics of Breast Cancer Progression in the MMTV-PyMT Mouse Model.
Cai Y, Nogales-Cadenas R, Zhang Q, Lin JR, Zhang W, O'Brien K, Montagna C, Zhang ZD (2017) BMC Genomics 18(1):185.  JOURNAL   PUBMED   REPRINT     DISEASE STUDY · FUNCTIONAL GENOMICS
ABSTRACT: BACKGROUND: Malignant breast cancer with complex molecular mechanisms of progression and metastasis remains a leading cause of death in women. To improve diagnosis and drug development, it is critical to identify panels of genes and molecular pathways involved in tumor progression and malignant transition. Using the PyMT mouse, a genetically engineered mouse model that has been widely used to study human breast cancer, we profiled and analyzed gene expression from four distinct stages of tumor progression (hyperplasia, adenoma/MIN, early carcinoma and late carcinoma) during which malignant transition occurs. RESULTS: We found remarkable expression similarity among the four stages, meaning genes altered in the later stages showed trace in the beginning of tumor progression. We identified a large number of differentially expressed genes in PyMT samples of all stages compared with normal mammary glands, enriched in cancer-related pathways. Using co-expression networks, we found panels of genes as signature modules with some hub genes that predict metastatic risk. Time-course analysis revealed genes with expression transition when shifting to malignant stages. These may provide additional insight into the molecular mechanisms beyond pathways. CONCLUSIONS: Thus, in this study, our various analyses with the PyMT mouse model shed new light on transcriptomic dynamics during breast cancer malignant progression.

MicroRNA Expression and Gene Regulation Drive Breast Cancer Progression and Metastasis in PyMT Mice.
Nogales-Cadenas R, Cai Y, Lin JR, Zhang Q, Zhang W, Montagna C, Zhang ZD (2016) Breast Cancer Res 18(1):75.  JOURNAL   PUBMED   REPRINT   WEBSITE     DISEASE STUDY · FUNCTIONAL GENOMICS
ABSTRACT: BACKGROUND: MicroRNAs (miRNAs) are small non-coding RNA molecules of about 22 nucleotides which function to silence the expression of their target genes. Numerous studies have shown that miRNAs are not only key regulators in important cellular processes but are also drivers in the development of many diseases, especially cancer. Estrogen receptor positive luminal B is the second most common but the least studied subtype of breast cancer. Only a few studies have examined the expression profiles of miRNAs in luminal B breast cancer, and their regulatory roles in cancer progression have yet to be investigated. METHODS: In this study, using polyoma middle T antigen (PyMT) mice, a widely used luminal B breast cancer model, we profiled microRNA (miRNA) expression at four time points that represent different key developmental stages of cancer progression. We considered the expression of both miRNAs and messenger RNAs (mRNAs) at these time points to improve the identification of regulatory targets of miRNAs. By combining gene functional and pathway annotation with miRNA-mRNA interactions, we created a PyMT-specific tripartite miRNA-mRNA-pathway network and identified novel functional regulatory programs (FRPs). RESULTS: We identified 151 differentially expressed miRNAs with a strict dual nature of either upregulation or downregulation during the whole course of disease progression. Among 82 newly discovered breast-cancer-related miRNAs, 35 can potentially regulate 271 protein-coding genes based on their sequence complementarity and expression profiles. We also identified miRNA-mRNA regulatory modules driving specific cancer-related biological processes. CONCLUSIONS: In this study we profiled the expression of miRNAs during breast cancer progression in the PyMT mouse model. By integrating miRNA and mRNA expression profiles, we identified differentially expressed miRNAs and their target genes involved in several hallmarks of cancer. We applied a novel clustering method to an annotated miRNA-mRNA regulatory network and identified network modules involved in specific cancer-related biological processes.

RNA:DNA Hybrids in the Human Genome Have Distinctive Nucleotide Characteristics, Chromatin Composition, and Transcriptional Relationships.
Nadel J, Athanasiadou R, Lemetre C, Wijetunga NA, Ó Broin P, Sato H, Zhang Z, Jeddeloh J, Montagna C, Golden A, Seoighe C, Greally JM (2015) Epigenetics Chromatin 8:46.  JOURNAL   PUBMED   REPRINT     EPIGENOMICS · FUNCTIONAL GENOMICS
ABSTRACT: BACKGROUND: RNA:DNA hybrids represent a non-canonical nucleic acid structure that has been associated with a range of human diseases and potential transcriptional regulatory functions. Mapping of RNA:DNA hybrids in human cells reveals them to have a number of characteristics that give insights into their functions. RESULTS: We find RNA:DNA hybrids to occupy millions of base pairs in the human genome. A directional sequencing approach shows the RNA component of the RNA:DNA hybrid to be purine-rich, indicating a thermodynamic contribution to their in vivo stability. The RNA:DNA hybrids are enriched at loci with decreased DNA methylation and increased DNase hypersensitivity, and within larger domains with characteristics of heterochromatin formation, indicating potential transcriptional regulatory properties. Mass spectrometry studies of chromatin at RNA:DNA hybrids shows the presence of the ILF2 and ILF3 transcription factors, supporting a model of certain transcription factors binding preferentially to the RNA:DNA conformation. CONCLUSIONS: Overall, there is little to indicate a dependence for RNA:DNA hybrids forming co-transcriptionally, with results from the ribosomal DNA repeat unit instead supporting the intriguing model of RNA generating these structures in trans. The results of the study indicate heterogeneous functions of these genomic elements and new insights into their formation and stability in vivo.

An Integrated Encyclopedia of DNA Elements in the Human Genome.
ENCODE Project Consortium (2012) Nature 489(7414):57-74.  JOURNAL   PUBMED   REPRINT   WEBSITE     FUNCTIONAL GENOMICS
ABSTRACT: The human genome encodes the blueprint of life, but the function of the vast majority of its nearly three billion bases is unknown. The Encyclopedia of DNA Elements (ENCODE) project has systematically mapped regions of transcription, transcription factor association, chromatin structure and histone modification. These data enabled us to assign biochemical functions for 80% of the genome, in particular outside of the well-studied protein-coding regions. Many discovered candidate regulatory elements are physically associated with one another and with expressed genes, providing new insights into the mechanisms of gene regulation. The newly identified elements also show a statistical correspondence to sequence variants linked to human disease, and can thereby guide interpretation of this variation. Overall, the project provides new insights into the organization and regulation of our genes and genome, and is an expansive resource of functional annotations for biomedical research.

EBNA1 Regulates Cellular Gene Expression by Binding Cellular Promoters.
Canaan A, Haviv I, Urban AE, Schulz VP, Hartman S, Zhang Z, Palejev D, Deisseroth AB, Lacy J, Snyder M, Gerstein M, Weissman SM (2009) Proc Natl Acad Sci U S A 106(52):22421-6.  JOURNAL   PUBMED   REPRINT     FUNCTIONAL GENOMICS
ABSTRACT: Epstein-Barr virus (EBV) is associated with several types of lymphomas and epithelial tumors including Burkitt's lymphoma (BL), HIV-associated lymphoma, posttransplant lymphoproliferative disorder, and nasopharyngeal carcinoma. EBV nuclear antigen 1 (EBNA1) is expressed in all EBV associated tumors and is required for latency and transformation. EBNA1 initiates latent viral replication in B cells, maintains the viral genome copy number, and regulates transcription of other EBV-encoded latent genes. These activities are mediated through the ability of EBNA1 to bind viral-DNA. To further elucidate the role of EBNA1 in the host cell, we have examined the effect of EBNA1 on cellular gene expression by microarray analysis using the B cell BJAB and the epithelial 293 cell lines transfected with EBNA1. Analysis of the data revealed distinct profiles of cellular gene changes in BJAB and 293 cell lines. Subsequently, chromatin immune-precipitation revealed a direct binding of EBNA1 to cellular promoters. We have correlated EBNA1 bound promoters with changes in gene expression. Sequence analysis of the 100 promoters most enriched revealed a DNA motif that differs from the EBNA1 binding site in the EBV genome.

Systematic Analysis of Transcribed Loci in ENCODE Regions Using RACE Sequencing Reveals Extensive Transcription in the Human Genome.
Wu JQ, Du J, Rozowsky J, Zhang Z, Urban AE, Euskirchen G, Weissman S, Gerstein M, Snyder M (2008) Genome Biol 9(1):R3.  JOURNAL   PUBMED   REPRINT     FUNCTIONAL GENOMICS
ABSTRACT: BACKGROUND: Recent studies of the mammalian transcriptome have revealed a large number of additional transcribed regions and extraordinary complexity in transcript diversity. However, there is still much uncertainty regarding precisely what portion of the genome is transcribed, the exact structures of these novel transcripts, and the levels of the transcripts produced. RESULTS: We have interrogated the transcribed loci in 420 selected ENCyclopedia Of DNA Elements (ENCODE) regions using rapid amplification of cDNA ends (RACE) sequencing. We analyzed annotated known gene regions, but primarily we focused on novel transcriptionally active regions (TARs), which were previously identified by high-density oligonucleotide tiling arrays and on random regions that were not believed to be transcribed. We found RACE sequencing to be very sensitive and were able to detect low levels of transcripts in specific cell types that were not detectable by microarrays. We also observed many instances of sense-antisense transcripts; further analysis suggests that many of the antisense transcripts (but not all) may be artifacts generated from the reverse transcription reaction. Our results show that the majority of the novel TARs analyzed (60%) are connected to other novel TARs or known exons. Of previously unannotated random regions, 17% were shown to produce overlapping transcripts. Furthermore, it is estimated that 9% of the novel transcripts encode proteins. CONCLUSION: We conclude that RACE sequencing is an efficient, sensitive, and highly accurate method for characterization of the transcriptome of specific cell/tissue types. Using this method, it appears that much of the genome is represented in polyA+ RNA. Moreover, a fraction of the novel RNAs can encode protein and are likely to be functional.

Analysis of Nuclear Receptor Pseudogenes in Vertebrates: How the Silent Tell Their Stories.
Zhang ZD, Cayting P, Weinstock G, Gerstein M (2008) Mol Biol Evol 25(1):131-43.  JOURNAL   PUBMED   REPRINT   WEBSITE     FUNCTIONAL GENOMICS · PSEUDOGENE
ABSTRACT: Transcription factor pseudogenes have not been systematically studied before. Nuclear receptors (NRs) constitute one of the largest groups of transcription factors in animals (e.g., 48 NRs in human). The availability of whole-genome sequences enables a global inventory of the NR pseudogenes in a number of vertebrate model organisms. Here we identify the NR pseudogenes in 8 vertebrate organisms and make our results available online at http://www.pseudogene.org/nr. The assignments reveal that NR pseudogenes as a group have characteristics related to generation and distribution contrary to expectations derived from previous large-scale pseudogene studies. In particular, 1) despite its large size, the NR gene family has only a very small number of pseudogenes in each of the vertebrate genomes examined; 2) despite the low transcription levels of NR genes, except for one, all other NR pseudogenes identified in this study are retropseudogenes; and 3) no duplicated NR pseudogenes are found, contrary to the fact that the NR gene family was expanded through several waves of gene duplication events. Our analyses further reveal a number of interesting aspects of NR pseudogenes. Specifically, through careful sequence analysis, we identify remnant introns in 2 mouse retropseudogenes, psiRev-erbbeta and psiLRH1. Generated from partially processed pre-mRNAs, they appear to be rare examples of highly unusual "semiprocessed" pseudogenes. Second, by comparing the genomic sequences, we uncover a pseudogene that is unique to the human lineage relative to chimpanzee. Generated by a recent duplication of a segment in the human genome, this pseudogene is a "duplicated-processed" pseudogene, belonging to a new pseudogene species. Finally, FXRbeta was nonfunctionalized in the human lineage and thus appears to be an example of a rare unitary pseudogene. By comparing orthologous sequences, we dated the FXR-FXRbeta duplication and the nonfunctionalization of FXRbeta in primates.

Divergence of Transcription Factor Binding Sites Across Related Yeast Species.
Borneman AR, Gianoulis TA, Zhang ZD, Yu H, Rozowsky J, Seringhaus MR, Wang LY, Gerstein M, Snyder M (2007) Science 317(5839):815-9.  JOURNAL   PUBMED   REPRINT   WEBSITE     FUNCTIONAL GENOMICS
ABSTRACT: Characterization of interspecies differences in gene regulation is crucial for understanding the molecular basis of both phenotypic diversity and evolution. By means of chromatin immunoprecipitation and DNA microarray analysis, the divergence in the binding sites of the pseudohyphal regulators Ste12 and Tec1 was determined in the yeasts Saccharomyces cerevisiae, S. mikatae, and S. bayanus under pseudohyphal conditions. We have shown that most of these sites have diverged across these species, far exceeding the interspecies variation in orthologous genes. A group of Ste12 targets was shown to be bound only in S. mikatae and S. bayanus under pseudohyphal conditions. Many of these genes are targets of Ste12 during mating in S. cerevisiae, indicating that specialization between the two pathways has occurred in this species. Transcription factor binding sites have therefore diverged substantially faster than ortholog content. Thus, gene regulation resulting from transcription factor binding is likely to be a major cause of divergence between related species.

Transcription Factor Binding Site Identification in Yeast: A Comparison of High-Density Oligonucleotide and PCR-based Microarray Platforms.
Borneman AR, Zhang ZD, Rozowsky J, Seringhaus MR, Gerstein M, Snyder M (2007) Funct Integr Genomics 7(4):335-45.  JOURNAL   PUBMED   REPRINT     FUNCTIONAL GENOMICS
ABSTRACT: In recent years, techniques have been developed to map transcription factor binding sites using chromatin immunoprecipitation combined with DNA microarrays (chIP chip). Initially, polymerase chain reaction (PCR)-based DNA arrays were used for the chIP chip procedure, however, high-density oligonucleotide (HDO) arrays, which allow for the production of thousands more features per array, have emerged as a competing array platform. To compare the two platforms, data from chIP chip analysis performed for three factors (Tec1, Ste12, and Sok2) using both HDO and PCR arrays under identical experimental conditions were compared. HDO arrays provided increased reproducibility and sensitivity, detecting approximately three times more binding events than the PCR arrays while also showing increased accuracy. The increased resolution provided by the HDO arrays also allowed for the identification of multiple binding peaks in close proximity and of novel binding events such as binding within ORFs. The HDO array platform provides a far more robust array system by all measures than PCR-based arrays, all of which is directly attributable to the large number of probes available.

Identification and Analysis of Functional Elements in 1% of the Human Genome by the ENCODE Pilot Project.
ENCODE Project Consortium, Birney E, Stamatoyannopoulos JA, Dutta A, Guigó R, Gingeras TR, Margulies EH, Weng Z, Snyder M, Dermitzakis ET, Thurman RE, Kuehn MS, Taylor CM, Neph S, Koch CM, Asthana S, Malhotra A, Adzhubei I, Greenbaum JA, Andrews RM, Flicek P, Boyle PJ, Cao H, Carter NP, Clelland GK, Davis S, Day N, Dhami P, Dillon SC, Dorschner MO, Fiegler H, Giresi PG, Goldy J, Hawrylycz M, Haydock A, Humbert R, James KD, Johnson BE, Johnson EM, Frum TT, Rosenzweig ER, Karnani N, Lee K, Lefebvre GC, Navas PA, Neri F, Parker SC, Sabo PJ, Sandstrom R, Shafer A, Vetrie D, Weaver M, Wilcox S, Yu M, Collins FS, Dekker J, Lieb JD, Tullius TD, Crawford GE, Sunyaev S, Noble WS, Dunham I, Denoeud F, Reymond A, Kapranov P, Rozowsky J, Zheng D, Castelo R, Frankish A, Harrow J, Ghosh S, Sandelin A, Hofacker IL, Baertsch R, Keefe D, Dike S, Cheng J, Hirsch HA, Sekinger EA, Lagarde J, Abril JF, Shahab A, Flamm C, Fried C, Hackermüller J, Hertel J, Lindemeyer M, Missal K, Tanzer A, Washietl S, Korbel J, Emanuelsson O, Pedersen JS, Holroyd N, Taylor R, Swarbreck D, Matthews N, Dickson MC, Thomas DJ, Weirauch MT, Gilbert J, Drenkow J, Bell I, Zhao X, Srinivasan KG, Sung WK, Ooi HS, Chiu KP, Foissac S, Alioto T, Brent M, Pachter L, Tress ML, Valencia A, Choo SW, Choo CY, Ucla C, Manzano C, Wyss C, Cheung E, Clark TG, Brown JB, Ganesh M, Patel S, Tammana H, Chrast J, Henrichsen CN, Kai C, Kawai J, Nagalakshmi U, Wu J, Lian Z, Lian J, Newburger P, Zhang X, Bickel P, Mattick JS, Carninci P, Hayashizaki Y, Weissman S, Hubbard T, Myers RM, Rogers J, Stadler PF, Lowe TM, Wei CL, Ruan Y, Struhl K, Gerstein M, Antonarakis SE, Fu Y, Green ED, Karaöz U, Siepel A, Taylor J, Liefer LA, Wetterstrand KA, Good PJ, Feingold EA, Guyer MS, Cooper GM, Asimenos G, Dewey CN, Hou M, Nikolaev S, Montoya-Burgos JI, Löytynoja A, Whelan S, Pardi F, Massingham T, Huang H, Zhang NR, Holmes I, Mullikin JC, Ureta-Vidal A, Paten B, Seringhaus M, Church D, Rosenbloom K, Kent WJ, Stone EA, NISC Comparative Sequencing Program, Baylor College of Medicine Human Genome Sequencing Center, Washington University Genome Sequencing Center, Broad Institute, Children's Hospital Oakland Research Institute, Batzoglou S, Goldman N, Hardison RC, Haussler D, Miller W, Sidow A, Trinklein ND, Zhang ZD, Barrera L, Stuart R, King DC, Ameur A, Enroth S, Bieda MC, Kim J, Bhinge AA, Jiang N, Liu J, Yao F, Vega VB, Lee CW, Ng P, Shahab A, Yang A, Moqtaderi Z, Zhu Z, Xu X, Squazzo S, Oberley MJ, Inman D, Singer MA, Richmond TA, Munn KJ, Rada-Iglesias A, Wallerman O, Komorowski J, Fowler JC, Couttet P, Bruce AW, Dovey OM, Ellis PD, Langford CF, Nix DA, Euskirchen G, Hartman S, Urban AE, Kraus P, Van Calcar S, Heintzman N, Kim TH, Wang K, Qu C, Hon G, Luna R, Glass CK, Rosenfeld MG, Aldred SF, Cooper SJ, Halees A, Lin JM, Shulha HP, Zhang X, Xu M, Haidar JN, Yu Y, Ruan Y, Iyer VR, Green RD, Wadelius C, Farnham PJ, Ren B, Harte RA, Hinrichs AS, Trumbower H, Clawson H, Hillman-Jackson J, Zweig AS, Smith K, Thakkapallayil A, Barber G, Kuhn RM, Karolchik D, Armengol L, Bird CP, de Bakker PI, Kern AD, Lopez-Bigas N, Martin JD, Stranger BE, Woodroffe A, Davydov E, Dimas A, Eyras E, Hallgrímsdóttir IB, Huppert J, Zody MC, Abecasis GR, Estivill X, Bouffard GG, Guan X, Hansen NF, Idol JR, Maduro VV, Maskeri B, McDowell JC, Park M, Thomas PJ, Young AC, Blakesley RW, Muzny DM, Sodergren E, Wheeler DA, Worley KC, Jiang H, Weinstock GM, Gibbs RA, Graves T, Fulton R, Mardis ER, Wilson RK, Clamp M, Cuff J, Gnerre S, Jaffe DB, Chang JL, Lindblad-Toh K, Lander ES, Koriabine M, Nefedov M, Osoegawa K, Yoshinaga Y, Zhu B, de Jong PJ (2007) Nature 447(7146):799-816.  JOURNAL   PUBMED   REPRINT   WEBSITE     FUNCTIONAL GENOMICS
ABSTRACT: We report the generation and analysis of functional data from multiple, diverse experiments performed on a targeted 1% of the human genome as part of the pilot phase of the ENCODE Project. These data have been further integrated and augmented by a number of evolutionary and computational analyses. Together, our results advance the collective knowledge about human genome function in several major areas. First, our studies provide convincing evidence that the genome is pervasively transcribed, such that the majority of its bases can be found in primary transcripts, including non-protein-coding transcripts, and those that extensively overlap one another. Second, systematic examination of transcriptional regulation has yielded new understanding about transcription start sites, including their relationship to specific regulatory sequences and features of chromatin accessibility and histone modification. Third, a more sophisticated view of chromatin structure has emerged, including its inter-relationship with DNA replication and transcriptional regulation. Finally, integration of these new sources of information, in particular with respect to mammalian evolution based on inter- and intra-species sequence comparisons, has yielded new mechanistic and evolutionary insights concerning the functional landscape of the human genome. Together, these studies are defining a path for pursuit of a more comprehensive characterization of human genome function.

Mapping of Transcription Factor Binding Regions in Mammalian Cells by ChIP: Comparison of Array- and Sequencing-Based Technologies.
Euskirchen GM, Rozowsky JS, Wei CL, Lee WH, Zhang ZD, Hartman S, Emanuelsson O, Stolc V, Weissman S, Gerstein MB, Ruan Y, Snyder M (2007) Genome Res 17(6):898-909.  JOURNAL   PUBMED   REPRINT     FUNCTIONAL GENOMICS
ABSTRACT: Recent progress in mapping transcription factor (TF) binding regions can largely be credited to chromatin immunoprecipitation (ChIP) technologies. We compared strategies for mapping TF binding regions in mammalian cells using two different ChIP schemes: ChIP with DNA microarray analysis (ChIP-chip) and ChIP with DNA sequencing (ChIP-PET). We first investigated parameters central to obtaining robust ChIP-chip data sets by analyzing STAT1 targets in the ENCODE regions of the human genome, and then compared ChIP-chip to ChIP-PET. We devised methods for scoring and comparing results among various tiling arrays and examined parameters such as DNA microarray format, oligonucleotide length, hybridization conditions, and the use of competitor Cot-1 DNA. The best performance was achieved with high-density oligonucleotide arrays, oligonucleotides >/=50 bases (b), the presence of competitor Cot-1 DNA and hybridizations conducted in microfluidics stations. When target identification was evaluated as a function of array number, 80%-86% of targets were identified with three or more arrays. Comparison of ChIP-chip with ChIP-PET revealed strong agreement for the highest ranked targets with less overlap for the low ranked targets. With advantages and disadvantages unique to each approach, we found that ChIP-chip and ChIP-PET are frequently complementary in their relative abilities to detect STAT1 targets for the lower ranked targets; each method detected validated targets that were missed by the other method. The most comprehensive list of STAT1 binding regions is obtained by merging results from ChIP-chip and ChIP-sequencing. Overall, this study provides information for robust identification, scoring, and validation of TF targets using ChIP-based technologies.

Statistical Analysis of the Genomic Distribution and Correlation of Regulatory Elements in the ENCODE Regions.
Zhang ZD, Paccanaro A, Fu Y, Weissman S, Weng Z, Chang J, Snyder M, Gerstein MB (2007) Genome Res 17(6):787-97.  JOURNAL   PUBMED   REPRINT   POSTER   WEBSITE     ANALYTICAL METHOD · FUNCTIONAL GENOMICS
ABSTRACT: The comprehensive inventory of functional elements in 44 human genomic regions carried out by the ENCODE Project Consortium enables for the first time a global analysis of the genomic distribution of transcriptional regulatory elements. In this study we developed an intuitive and yet powerful approach to analyze the distribution of regulatory elements found in many different ChIP-chip experiments on a 10 approximately 100-kb scale. First, we focus on the overall chromosomal distribution of regulatory elements in the ENCODE regions and show that it is highly nonuniform. We demonstrate, in fact, that regulatory elements are associated with the location of known genes. Further examination on a local, single-gene scale shows an enrichment of regulatory elements near both transcription start and end sites. Our results indicate that overall these elements are clustered into regulatory rich "islands" and poor "deserts." Next, we examine how consistent the nonuniform distribution is between different transcription factors. We perform on all the factors a multivariate analysis in the framework of a biplot, which enhances biological signals in the experiments. This groups transcription factors into sequence-specific and sequence-nonspecific clusters. Moreover, with experimental variation carefully controlled, detailed correlations show that the distribution of sites was generally reproducible for a specific factor between different laboratories and microarray platforms. Data sets associated with histone modifications have particularly strong correlations. Finally, we show how the correlations between factors change when only regulatory elements far from the transcription start sites are considered.

Integrated Analysis of Experimental Data Sets Reveals Many Novel Promoters in 1% of the Human Genome.
Trinklein ND, Karaöz U, Wu J, Halees A, Force Aldred S, Collins PJ, Zheng D, Zhang ZD, Gerstein MB, Snyder M, Myers RM, Weng Z (2007) Genome Res 17(6):720-31.  JOURNAL   PUBMED   REPRINT     FUNCTIONAL GENOMICS
ABSTRACT: The regulation of transcriptional initiation in the human genome is a critical component of global gene regulation, but a complete catalog of human promoters currently does not exist. In order to identify regulatory regions, we developed four computational methods to integrate 129 sets of ENCODE-wide chromatin immunoprecipitation data. They collectively predicted 1393 regions. Roughly 47% of the regions were unique to one method, as each method makes different assumptions about the data. Overall, predicted regions tend to localize to highly conserved, DNase I hypersensitive, and actively transcribed regions in the genome. Interestingly, a significant portion of the regions overlaps with annotated 3'-UTRs, suggesting that some of them might regulate anti-sense transcription. The majority of the predicted regions are >2 kb away from the 5'-ends of previously annotated human cDNAs and hence are novel. These novel regions may regulate unannotated transcripts or may represent new alternative transcription start sites of known genes. We tested 163 such regions for promoter activity in four cell lines using transient transfection assays, and 25% of them showed transcriptional activity above background in at least one cell line. We also performed 5'-RACE experiments on 62 novel regions, and 76% of the regions were associated with the 5'-ends of at least two RACE products. Our results suggest that there are at least 35% more functional promoters in the human genome than currently annotated.

What Is a Gene, Post-Encode? History and Updated Definition.
Gerstein MB, Bruce C, Rozowsky JS, Zheng D, Du J, Korbel JO, Emanuelsson O, Zhang ZD, Weissman S, Snyder M (2007) Genome Res 17(6):669-81.  JOURNAL   PUBMED   REPRINT   POSTER     FUNCTIONAL GENOMICS · REVIEW / PERSPECTIVE
ABSTRACT: While sequencing of the human genome surprised us with how many protein-coding genes there are, it did not fundamentally change our perspective on what a gene is. In contrast, the complex patterns of dispersed regulation and pervasive transcription uncovered by the ENCODE project, together with non-genic conservation and the abundance of noncoding RNA genes, have challenged the notion of the gene. To illustrate this, we review the evolution of operational definitions of a gene over the past century--from the abstract elements of heredity of Mendel and Morgan to the present-day ORFs enumerated in the sequence databanks. We then summarize the current ENCODE findings and provide a computational metaphor for the complexity. Finally, we propose a tentative update to the definition of a gene: A gene is a union of genomic sequences encoding a coherent set of potentially overlapping functional products. Our definition side-steps the complexities of regulation and transcription by removing the former altogether from the definition and arguing that final, functional gene products (rather than intermediate transcripts) should be used to group together entities associated with a single gene. It also manifests how integral the concept of biological function is in defining genes.

Genomic Analysis of the Nuclear Receptor Family: New Insights into Structure, Regulation, and Evolution from the Rat Genome.
Zhang Z, Burch PE, Cooney AJ, Lanz RB, Pereira FA, Wu J, Gibbs RA, Weinstock G, Wheeler DA (2004) Genome Res 14(4):580-90.  JOURNAL   PUBMED   REPRINT   POSTER     COMPARATIVE GENOMICS · FUNCTIONAL GENOMICS
ABSTRACT: Completion of the Rattus norvegicus genome sequence enabled a global inventory and analysis of the nuclear receptors (NRs) in three mammalian species. Forty-nine NR members were found in mouse, 48 in human. Forty-seven were found in the rat, with gaps at the locations expected for the other two. Pairwise comparisons of their distribution in rat, mouse, and human identified 11 syntenic NR gene blocks, including three small clusters of two or three closely related genes, each spanning 40 kb to 1700 kb. The exon structure of the ligand-binding domain suggests that exon shuffling has played a role in the evolution of this family. An invariant splice junction in all members of the NR family except LXRbeta suggests a functional role for the intron. The ligand-binding domains of PXR and CAR are among the most divergent in the family. Their higher nucleotide substitution rates may be related to the central role played by these two NRs in the metabolism of the foreign compounds and may have resulted from limited positive selection.

Genetic variation (18)

Deletion Size and Background Genetic Variation Shape Congenital Heart Disease Phenotypes in 3,016 Individuals with 22q11.2 Deletion Syndrome
Lin J-R, Miller D, Luong D, Nelson T, Crowley TB, Tran OT, Thiruvahindrapuram B, Hajianpour A, Campbell L, Busa T, Heine-Suner D, Garcia-Minaur S, Fernandez L, Murphy KC, Murphy D, Hawula W, Angkustsiri K, Shashi V, Schoch K, Bearden CE, Tomita Mitchell A, Mitchell ME, Carmel M, Weizman A, Michaelovsky E, Gothelf D, van den Bree MBM, Owen MJ, Vorstman JAS, Boot E, Vingerhoets C, van Amelsvoort T, Swillen A, Breckpot J, Vermeesch JR, Devriendt K, Schneider M, Eliez S, Digilio MC, Unolt M, Putotto C, Marino B, Pontillo M, Armando M, Vicari S, Repetto GM, Kates WR, Shprintzen RJ, Gur RE, Zackai EH, Goldmuntz E, Wang T, Raj S, Emanuel BS, McDonald-McGinn DM, Scherer SC, Bassett AS, Zhang ZD, Morrow BE (2026) medRxiv 2026.02.23.26346918.  JOURNAL     DISEASE STUDY · GENETIC VARIATION
ABSTRACT: Congenital heart disease (CHD) occurs in over half of individuals with 22q11.2 deletion syndrome (22q11.2DS) and the types of lesions range from mild to severe. To determine the basis of variation in cardiac phenotypes we analyzed demographic data from 3,016 unrelated individuals with 22q11.2DS from centers in the Northeast US, Canada, Europe, South America, Israel and Australia. Most individuals in this cohort had a 3 million base pair hemizygous deletion between low copy repeat, LCR22 A-D (87.2%), while some had nested deletions. We performed multivariable mixed-effects logistic regression and uncovered significant differences between CHD phenotypes and basic demographic features. Individuals with the A-D deletion had a lower risk of persistent truncus arteriosus (OR = 0.37, 95% CI 0.18-0.75) but a higher risk of septal defects (OR = 4.7, 95% CI 1.7-12.8) compared to those with the smaller A-B deletion, suggesting distinct developmental pathways sensitive to 22q11.2 gene dosage. In addition, genome-wide genetic principal components (PCs) were associated with specific CHD subtypes, including reduced risk of pulmonary stenosis or atresia with other heart lesions (PC2; OR = 0.73, 95% CI 0.61-0.87) and increased risk of abnormal origin of the subclavian arteries (PC4; OR = 2.6, 95% CI 1.4-4.9), indicating that background genetic variation modifies heart lesion-specific susceptibility. Together, these results suggest that both deletion size and background genetic variation shape the highly variable cardiac phenotypes in 22q11.2DS.

Prevalence and Spectrum of Congenital Heart Disease in Individuals with Distal Chromosome 22q11.22-23 Deletions.
Nelson TJ, McGinn DE, Crowley TB, Rockart L, Green A, Giunta V, Tran O, Miller D, Breckpot J, Swillen A, Digilio MC, Unolt M, Putotto C, Pulvirenti F, Marino B, Emanuel BS, Zackai EH, Zhang ZD, Goldmuntz E, Boot E, Bassett AS, Morrow BE, McDonald-McGinn DM (2026) Clin Genet 109(5):859-868.  JOURNAL   PUBMED   REPRINT     DISEASE STUDY · GENETIC VARIATION
ABSTRACT: This study is aimed at determining the spectrum of congenital heart disease associated with distal 22q11.22-23 deletions flanked by low copy repeats, LCR22 D-H. We analyzed cardiology findings in 128 unrelated individuals with distal LCR22 D-H deletions. A total of 62 were newly described and 66 were derived from previous reports. We found that deletions which included LCR22-D as the proximal endpoint were the most prevalent in the cohort (104/128, 81.3%). Clinically relevant congenital heart disease was identified in 48 individuals (37.5%, 95% CI 29%-46%), which is lower than the prevalence reported for typical, proximal LCR22 A-D deletions (p = 3.7E-4), especially for conotruncal defects (13/128, 10.2%; p = 7.1E-13). Mild to moderate CHD predominated, including ventricular septal defects (22/128), bicuspid aortic valve (9/128) and mild cardiomyopathy (3/128). Persistent truncus arteriosus was the most prevalent (n = 8/13) conotruncal heart defect, but other anomalies also occurred in singleton cases. These findings support the need for cardiac evaluation in all individuals with distal 22q11.22-23 deletions, increased use of clinical genetic testing in syndromic individuals with these findings, and molecular studies in model systems. The results demonstrate that reduced gene dosage of distal 22q11.21-23, particularly within the D-E region including MAPK1 and HIC2 convey risk for CHD.

Genetic Variants Associated with Age-Related Episodic Memory Decline Implicate Distinct Memory Pathologies.
Ali A, Milman S, Weiss EF, Gao T, Napolioni V, Barzilai N, Zhang ZD, Lin JR (2025) Alzheimers Dement 21(1):e14379.  JOURNAL   PUBMED   REPRINT     AGING · DISEASE STUDY · GENETIC VARIATION
ABSTRACT: BACKGROUND: Approximately 40% of people aged ≥ 65 experience memory loss, particularly in episodic memory. Identifying the genetic basis of episodic memory decline is crucial for uncovering its underlying causes. METHODS: We investigated common and rare genetic variants associated with episodic memory decline in 742 (632 for rare variants) Ashkenazi Jewish individuals (mean age 75) from the LonGenity study. All-atom molecular dynamics simulations were performed to uncover mechanistic insights underlying rare variants associated with episodic memory decline. RESULTS: In addition to the common polygenic risk of Alzheimer's disease, we identified and replicated rare variant associations in ITSN1 and CRHR2. Structural analyses revealed distinct memory pathologies mediated by interfacial rare coding variants such as impaired receptor activation of corticotropin releasing hormone and dysregulated L-serine synthesis. DISCUSSION: Our study uncovers novel risk loci for episodic memory decline. The identified underlying mechanisms point toward heterogenous memory pathologies mediated by rare coding variants. HIGHLIGHTS: We demonstrated the contribution of the common polygenic risk of Alzheimer's disease to episodic memory decline. We discovered and replicated two risk genes associated with episodic memory decline implicated by rare variants, were discovered and replicated. We demonstrated molecular mechanisms and potential novel memory pathologies underlying interfacial rare coding variants. Molecular dynamics simulations were performed to understand the downstream effects of risk rare coding variants.

Polygenic Prediction of Human Longevity on the Supposition of Pervasive Pleiotropy.
Jabalameli MR, Lin JR, Zhang Q, Wang Z, Mitra J, Nguyen N, Gao T, Khusidman M, Sathyan S, Atzmon G, Milman S, Vijg J, Barzilai N, Zhang ZD (2024) Sci Rep 14(1):19981.  JOURNAL   PUBMED   REPRINT     AGING · GENETIC VARIATION
ABSTRACT: The highly polygenic nature of human longevity renders pleiotropy an indispensable feature of its genetic architecture. Leveraging the genetic correlation between aging-related traits (ARTs), we aimed to model the additive variance in lifespan as a function of the cumulative liability from pleiotropic segregating variants. We tracked allele frequency changes as a function of viability across different age bins and prioritized 34 variants with an immediate implication on lipid metabolism, body mass index (BMI), and cognitive performance, among other traits, revealed by PheWAS analysis in the UK Biobank. Given the highly complex and non-linear interactions between the genetic determinants of longevity, we reasoned that a composite polygenic score would approximate a substantial portion of the variance in lifespan and developed the integrated longevity genetic scores (iLGSs) for distinguishing exceptional survival. We showed that coefficients derived from our ensemble model could potentially reveal an interesting pattern of genomic pleiotropy specific to lifespan. We assessed the predictive performance of our model for distinguishing the enrichment of exceptional longevity among long-lived individuals in two replication cohorts (the Scripps Wellderly cohort and the Medical Genome Reference Bank (MRGB)) and showed that the median lifespan in the highest decile of our composite prognostic index is up to 4.8 years longer. Finally, using the proteomic correlates of iLGS, we identified protein markers associated with exceptional longevity irrespective of chronological age and prioritized drugs with repurposing potentials for gerotherapeutics. Together, our approach demonstrates a promising framework for polygenic modeling of additive liability conferred by ARTs in defining exceptional longevity and assisting the identification of individuals at a higher risk of mortality for targeted lifestyle modifications earlier in life. Furthermore, the proteomic signature associated with iLGS highlights the functional pathway upstream of the PI3K-Akt that can be effectively targeted to slow down aging and extend lifespan.

Rare Coding Variants as Risk Modifiers of the 22q11.2 Deletion Implicate Postnatal Cortical Development in Syndromic Schizophrenia.
Lin JR, Zhao Y, Jabalameli MR, Nguyen N, Mitra J, International 22q11.DS Brain and Behavior Consortium, Swillen A, Vorstman JAS, Chow EWC, van den Bree M, Emanuel BS, Vermeesch JR, Owen MJ, Williams NM, Bassett AS, McDonald-McGinn DM, Gur RE, Bearden CE, Morrow BE, Lachman HM, Zhang ZD (2023) Mol Psychiatry 28(5):2071-2080.  JOURNAL   PUBMED   REPRINT   WEBSITE     ANALYTICAL METHOD · DISEASE STUDY · GENETIC VARIATION
ABSTRACT: 22q11.2 deletion is one of the strongest known genetic risk factors for schizophrenia. Recent whole-genome sequencing of schizophrenia cases and controls with this deletion provided an unprecedented opportunity to identify risk modifying genetic variants and investigate their contribution to the pathogenesis of schizophrenia in 22q11.2 deletion syndrome. Here, we apply a novel analytic framework that integrates gene network and phenotype data to investigate the aggregate effects of rare coding variants and identified modifier genes in this etiologically homogenous cohort (223 schizophrenia cases and 233 controls of European descent). Our analyses revealed significant additive genetic components of rare nonsynonymous variants in 110 modifier genes (adjusted P = 9.4E-04) that overall accounted for 4.6% of the variance in schizophrenia status in this cohort, of which 4.0% was independent of the common polygenic risk for schizophrenia. The modifier genes affected by rare coding variants were enriched with genes involved in synaptic function and developmental disorders. Spatiotemporal transcriptomic analyses identified an enrichment of coexpression between modifier and 22q11.2 genes in cortical brain regions from late infancy to young adulthood. Corresponding gene coexpression modules are enriched with brain-specific protein-protein interactions of SLC25A1, COMT, and PI4KA in the 22q11.2 deletion region. Overall, our study highlights the contribution of rare coding variants to the SCZ risk. They not only complement common variants in disease genetics but also pinpoint brain regions and developmental stages critical to the etiology of syndromic schizophrenia.

◸ NEWS AND VIEWS 
Unravelling Genetic Components of Longevity.
Jabalameli MR, Zhang ZD (2022) Nat Aging 2(1):5-6.  JOURNAL   PUBMED   REPRINT     AGING · DISEASE STUDY · GENETIC VARIATION · REVIEW / PERSPECTIVE
ABSTRACT: Many aging-related traits share a common genetic component. How to disentangle it from the trait-specific effects has remained largely unexplored. A new study in Nature Aging uses an analysis framework for isolating the shared genetic component in genome-wide association studies of aging-related traits and identifies genomic loci that contribute to aging.

Substance Abuse and the Risk of Severe COVID-19: Mendelian Randomization Confirms the Causal Role of Opioids but Hints a Negative Causal Effect for Cannabinoids.
Jabalameli MR, Zhang ZD (2022) Front Genet 13:1070428.  JOURNAL   PUBMED   REPRINT     DISEASE STUDY · GENETIC VARIATION
ABSTRACT: Since the start of the COVID-19 global pandemic, our understanding of the underlying disease mechanism and factors associated with the disease severity has dramatically increased. A recent study investigated the relationship between substance use disorders (SUD) and the risk of severe COVID-19 in the United States and concluded that the risk of hospitalization and death due to COVID-19 is directly correlated with substance abuse, including opioid use disorder (OUD) and cannabis use disorder (CUD). While we found this analysis fascinating, we believe this observation may be biased due to comorbidities (such as hypertension, diabetes, and cardiovascular disease) confounding the direct effect of SUD on severe COVID-19 illness. To answer this question, we sought to investigate the causal relationship between substance abuse and medication-taking history (as a proxy trait for comorbidities) with the risk of COVID-19 adverse outcomes. Our Mendelian randomization analysis confirms the causal relationship between OUD and severe COVID-19 illness but suggests an inverse causal effect for cannabinoids. Considering that COVID-19 mortality is largely attributed to disturbed immune regulation, the possible modulatory impact of cannabinoids in alleviating cytokine storms merits further investigation.

Rare Genetic Coding Variants Associated with Human Longevity and Protection Against Age-Related Diseases.
Lin JR, Sin-Chan P, Napolioni V, Torres GG, Mitra J, Zhang Q, Jabalameli MR, Wang Z, Nguyen N, Gao T, Regeneron Genetics Center, Laudes M, Görg S, Franke A, Nebel A, Greicius MD, Atzmon G, Ye K, Gorbunova V, Ladiges WC, Shuldiner AR, Niedernhofer LJ, Robbins PD, Milman S, Suh Y, Vijg J, Barzilai N, Zhang ZD (2021) Nat Aging 1(9):783-794.  JOURNAL   PUBMED   REPRINT   WEBSITE     AGING · DISEASE STUDY · GENETIC VARIATION
ABSTRACT: Extreme longevity in humans has a strong genetic component, but whether this involves genetic variation in the same longevity pathways as found in model organisms is unclear. Using whole-exome sequences of a large cohort of Ashkenazi Jewish centenarians to examine enrichment for rare coding variants, we found most longevity-associated rare coding variants converge upon conserved insulin/insulin-like growth factor 1 signaling and AMP-activating protein kinase signaling pathways. Centenarians have a number of pathogenic rare coding variants similar to control individuals, suggesting that rare variants detected in the conserved longevity pathways are protective against age-related pathology. Indeed, we detected a pro-longevity effect of rare coding variants in the Wnt signaling pathway on individuals harboring the known common risk allele APOE4. The genetic component of extreme human longevity constitutes, at least in part, rare coding variants in pathways that protect against aging, including those that control longevity in model organisms.

Enhancer Release and Retargeting Activates Disease-Susceptibility Genes.
Oh S, Shao J, Mitra J, Xiong F, D'Antonio M, Wang R, Garcia-Bassets I, Ma Q, Zhu X, Lee JH, Nair SJ, Yang F, Ohgi K, Frazer KA, Zhang ZD, Li W, Rosenfeld MG (2021) Nature 595(7869):735-740.  JOURNAL   PUBMED   REPRINT     DISEASE STUDY · FUNCTIONAL GENOMICS · GENETIC VARIATION
ABSTRACT: The functional engagement between an enhancer and its target promoter ensures precise gene transcription1. Understanding the basis of promoter choice by enhancers has important implications for health and disease. Here we report that functional loss of a preferred promoter can release its partner enhancer to loop to and activate an alternative promoter (or alternative promoters) in the neighbourhood. We refer to this target-switching process as 'enhancer release and retargeting'. Genetic deletion, motif perturbation or mutation, and dCas9-mediated CTCF tethering reveal that promoter choice by an enhancer can be determined by the binding of CTCF at promoters, in a cohesin-dependent manner-consistent with a model of 'enhancer scanning' inside the contact domain. Promoter-associated CTCF shows a lower affinity than that at chromatin domain boundaries and often lacks a preferred motif orientation or a partnering CTCF at the cognate enhancer, suggesting properties distinct from boundary CTCF. Analyses of cancer mutations, data from the GTEx project and risk loci from genome-wide association studies, together with a focused CRISPR interference screen, reveal that enhancer release and retargeting represents an overlooked mechanism that underlies the activation of disease-susceptibility genes, as exemplified by a risk locus for Parkinson's disease (NUCKS1-RAB7L1) and three loci associated with cancer (CLPTM1L-TERT, ZCCHC7-PAX5 and PVT1-MYC).

Deep Post-Gwas Analysis Identifies Potential Risk Genes and Risk Variants for Alzheimer's Disease, Providing New Insights into Its Disease Mechanisms.
Wang Z, Zhang Q, Lin JR, Jabalameli MR, Mitra J, Nguyen N, Zhang ZD (2021) Sci Rep 11(1):20511.  JOURNAL   PUBMED   REPRINT     ANALYTICAL METHOD · DISEASE STUDY · GENETIC VARIATION
ABSTRACT: Alzheimer's disease (AD) is a genetically complex, multifactorial neurodegenerative disease. It affects more than 45 million people worldwide and currently remains untreatable. Although genome-wide association studies (GWAS) have identified many AD-associated common variants, only about 25 genes are currently known to affect the risk of developing AD, despite its highly polygenic nature. Moreover, the risk variants underlying GWAS AD-association signals remain unknown. Here, we describe a deep post-GWAS analysis of AD-associated variants, using an integrated computational framework for predicting both disease genes and their risk variants. We identified 342 putative AD risk genes in 203 risk regions spanning 502 AD-associated common variants. 246 AD risk genes have not been identified as AD risk genes by previous GWAS collected in GWAS catalogs, and 115 of 342 AD risk genes are outside the risk regions, likely under the regulation of transcriptional regulatory elements contained therein. Even more significantly, for 109 AD risk genes, we predicted 150 risk variants, of both coding and regulatory (in promoters or enhancers) types, and 85 (57%) of them are supported by functional annotation. In-depth functional analyses showed that AD risk genes were overrepresented in AD-related pathways or GO terms-e.g., the complement and coagulation cascade and phosphorylation and activation of immune response-and their expression was relatively enriched in microglia, endothelia, and pericytes of the human brain. We found nine AD risk genes-e.g., IL1RAP, PMAIP1, LAMTOR4-as predictors for the prognosis of AD survival and genes such as ARL6IP5 with altered network connectivity between AD patients and normal individuals involved in AD progression. Our findings open new strategies for developing therapeutics targeting AD risk genes or risk variants to influence AD pathogenesis.

Genetics of Extreme Human Longevity to Guide Drug Discovery for Healthy Ageing.
Zhang ZD, Milman S, Lin JR, Wierbowski S, Yu H, Barzilai N, Gorbunova V, Ladiges WC, Niedernhofer LJ, Suh Y, Robbins PD, Vijg J (2020) Nat Metab 2(8):663-672.  JOURNAL   PUBMED   REPRINT     AGING · DISEASE STUDY · GENETIC VARIATION · REVIEW / PERSPECTIVE
ABSTRACT: Ageing is the greatest risk factor for most common chronic human diseases, and it therefore is a logical target for developing interventions to prevent, mitigate or reverse multiple age-related morbidities. Over the past two decades, genetic and pharmacologic interventions targeting conserved pathways of growth and metabolism have consistently led to substantial extension of the lifespan and healthspan in model organisms as diverse as nematodes, flies and mice. Recent genetic analysis of long-lived individuals is revealing common and rare variants enriched in these same conserved pathways that significantly correlate with longevity. In this Perspective, we summarize recent insights into the genetics of extreme human longevity and propose the use of this rare phenotype to identify genetic variants as molecular targets for gaining insight into the physiology of healthy ageing and the development of new therapies to extend the human healthspan.

PGA: Post-Gwas Analysis for Disease Gene Identification.
Lin JR, Jaroslawicz D, Cai Y, Zhang Q, Wang Z, Zhang ZD (2018) Bioinformatics 34(10):1786-1788.  JOURNAL   PUBMED   REPRINT   WEBSITE     ANALYTICAL METHOD · DISEASE STUDY · GENETIC VARIATION
ABSTRACT: SUMMARY: Although the genome-wide association study (GWAS) is a powerful method to identify disease-associated variants, it does not directly address the biological mechanisms underlying such genetic association signals. Here, we present PGA, a Perl- and Java-based program for post-GWAS analysis that predicts likely disease genes given a list of GWAS-reported variants. Designed with a command line interface, PGA incorporates genomic and eQTL data in identifying disease gene candidates and uses gene network and ontology data to score them based upon the strength of their relationship to the disease in question. AVAILABILITY AND IMPLEMENTATION: http://zdzlab.einstein.yu.edu/1/pga.html. CONTACT: zhengdong.zhang@einstein.yu.edu. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.

Integrated Rare Variant-Based Risk Gene Prioritization in Disease Case-Control Sequencing Studies.
Lin JR, Zhang Q, Cai Y, Morrow BE, Zhang ZD (2017) PLoS Genet 13(12):e1007142.  JOURNAL   PUBMED   REPRINT   WEBSITE     ANALYTICAL METHOD · DISEASE STUDY · GENETIC VARIATION
ABSTRACT: Rare variants of major effect play an important role in human complex diseases and can be discovered by sequencing-based genome-wide association studies. Here, we introduce an integrated approach that combines the rare variant association test with gene network and phenotype information to identify risk genes implicated by rare variants for human complex diseases. Our data integration method follows a 'discovery-driven' strategy without relying on prior knowledge about the disease and thus maintains the unbiased character of genome-wide association studies. Simulations reveal that our method can outperform a widely-used rare variant association test method by 2 to 3 times. In a case study of a small disease cohort, we uncovered putative risk genes and the corresponding rare variants that may act as genetic modifiers of congenital heart disease in 22q11.2 deletion syndrome patients. These variants were missed by a conventional approach that relied on the rare variant association test alone.

Whole-Genome Sequencing and Integrative Genomic Analysis Approach on Two 22q11.2 Deletion Syndrome Family Trios for Genotype to Phenotype Correlations.
Chung JH, Cai J, Suskin BG, Zhang Z, Coleman K, Morrow BE (2015) Hum Mutat 36(8):797-807.  JOURNAL   PUBMED   REPRINT     DISEASE STUDY · GENETIC VARIATION · GENOMIC SEQUENCING
ABSTRACT: The 22q11.2 deletion syndrome (22q11DS) affects 1:4,000 live births and presents with highly variable phenotype expressivity. In this study, we developed an analytical approach utilizing whole-genome sequencing (WGS) and integrative analysis to discover genetic modifiers. Our pipeline combined available tools in order to prioritize rare, predicted deleterious, coding and noncoding single-nucleotide variants (SNVs), and insertion/deletions from WGS. We sequenced two unrelated probands with 22q11DS, with contrasting clinical findings, and their unaffected parents. Proband P1 had cognitive impairment, psychotic episodes, anxiety, and tetralogy of Fallot (TOF), whereas proband P2 had juvenile rheumatoid arthritis but no other major clinical findings. In P1, we identified common variants in COMT and PRODH on 22q11.2 as well as rare potentially deleterious DNA variants in other behavioral/neurocognitive genes. We also identified a de novo SNV in ADNP2 (NM_014913.3:c.2243G>C), encoding a neuroprotective protein that may be involved in behavioral disorders. In P2, we identified a novel nonsynonymous SNV in ZFPM2 (NM_012082.3:c.1576C>T), a known causative gene for TOF, which may act as a protective variant downstream of TBX1, haploinsufficiency of which is responsible for congenital heart disease in individuals with 22q11DS.

The Origin, Evolution, and Functional Impact of Short Insertion-Deletion Variants Identified in 179 Human Genomes.
Montgomery SB, Goode DL, Kvikstad E, Albers CA, Zhang ZD, Mu XJ, Ananda G, Howie B, Karczewski KJ, Smith KS, Anaya V, Richardson R, Davis J, 1000 Genomes Project Consortium, MacArthur DG, Sidow A, Duret L, Gerstein M, Makova KD, Marchini J, McVean G, Lunter G (2013) Genome Res 23(5):749-61.  JOURNAL   PUBMED   REPRINT     GENETIC VARIATION · GENOMIC SEQUENCING
ABSTRACT: Short insertions and deletions (indels) are the second most abundant form of human genetic variation, but our understanding of their origins and functional effects lags behind that of other types of variants. Using population-scale sequencing, we have identified a high-quality set of 1.6 million indels from 179 individuals representing three diverse human populations. We show that rates of indel mutagenesis are highly heterogeneous, with 43%-48% of indels occurring in 4.03% of the genome, whereas in the remaining 96% their prevalence is 16 times lower than SNPs. Polymerase slippage can explain upwards of three-fourths of all indels, with the remainder being mostly simple deletions in complex sequence. However, insertions do occur and are significantly associated with pseudo-palindromic sequence features compatible with the fork stalling and template switching (FoSTeS) mechanism more commonly associated with large structural variations. We introduce a quantitative model of polymerase slippage, which enables us to identify indel-hypermutagenic protein-coding genes, some of which are associated with recurrent mutations leading to disease. Accounting for mutational rate heterogeneity due to sequence context, we find that indels across functional sequence are generally subject to stronger purifying selection than SNPs. We find that indel length modulates selection strength, and that indels affecting multiple functionally constrained nucleotides undergo stronger purifying selection. We further find that indels are enriched in associations with gene expression and find evidence for a contribution of nonsense-mediated decay. Finally, we show that indels can be integrated in existing genome-wide association studies (GWAS); although we do not find direct evidence that potentially causal protein-coding indels are enriched with associations to known disease-associated SNPs, our findings suggest that the causal variant underlying some of these associations may be indels.

A Systematic Survey of Loss-of-Function Variants in Human Protein-Coding Genes.
MacArthur DG, Balasubramanian S, Frankish A, Huang N, Morris J, Walter K, Jostins L, Habegger L, Pickrell JK, Montgomery SB, Albers CA, Zhang ZD, Conrad DF, Lunter G, Zheng H, Ayub Q, DePristo MA, Banks E, Hu M, Handsaker RE, Rosenfeld JA, Fromer M, Jin M, Mu XJ, Khurana E, Ye K, Kay M, Saunders GI, Suner MM, Hunt T, Barnes IH, Amid C, Carvalho-Silva DR, Bignell AH, Snow C, Yngvadottir B, Bumpstead S, Cooper DN, Xue Y, Romero IG, 1000 Genomes Project Consortium, Wang J, Li Y, Gibbs RA, McCarroll SA, Dermitzakis ET, Pritchard JK, Barrett JC, Harrow J, Hurles ME, Gerstein MB, Tyler-Smith C (2012) Science 335(6070):823-8.  JOURNAL   PUBMED   REPRINT     GENETIC VARIATION · GENOMIC SEQUENCING
ABSTRACT: Genome-sequencing studies indicate that all humans carry many genetic variants predicted to cause loss of function (LoF) of protein-coding genes, suggesting unexpected redundancy in the human genome. Here we apply stringent filters to 2951 putative LoF variants obtained from 185 human genomes to determine their true prevalence and properties. We estimate that human genomes typically contain ~100 genuine LoF variants with ~20 genes completely inactivated. We identify rare and likely deleterious LoF alleles, including 26 known and 21 predicted severe disease-causing variants, as well as common LoF variants in nonessential genes. We describe functional and evolutionary differences between LoF-tolerant and recessive disease genes and a method for using these differences to prioritize candidate genes found in clinical sequencing studies.

Mapping Copy Number Variation by Population-Scale Genome Sequencing.
Mills RE, Walter K, Stewart C, Handsaker RE, Chen K, Alkan C, Abyzov A, Yoon SC, Ye K, Cheetham RK, Chinwalla A, Conrad DF, Fu Y, Grubert F, Hajirasouliha I, Hormozdiari F, Iakoucheva LM, Iqbal Z, Kang S, Kidd JM, Konkel MK, Korn J, Khurana E, Kural D, Lam HY, Leng J, Li R, Li Y, Lin CY, Luo R, Mu XJ, Nemesh J, Peckham HE, Rausch T, Scally A, Shi X, Stromberg MP, Stütz AM, Urban AE, Walker JA, Wu J, Zhang Y, Zhang ZD, Batzer MA, Ding L, Marth GT, McVean G, Sebat J, Snyder M, Wang J, Ye K, Eichler EE, Gerstein MB, Hurles ME, Lee C, McCarroll SA, Korbel JO, 1000 Genomes Project (2011) Nature 470(7332):59-65.  JOURNAL   PUBMED   REPRINT     GENETIC VARIATION · GENOMIC SEQUENCING
ABSTRACT: Genomic structural variants (SVs) are abundant in humans, differing from other forms of variation in extent, origin and functional impact. Despite progress in SV characterization, the nucleotide resolution architecture of most SVs remains unknown. We constructed a map of unbalanced SVs (that is, copy number variants) based on whole genome DNA sequencing data from 185 human genomes, integrating evidence from complementary SV discovery approaches with extensive experimental validations. Our map encompassed 22,025 deletions and 6,000 additional SVs, including insertions and tandem duplications. Most SVs (53%) were mapped to nucleotide resolution, which facilitated analysing their origin and functional impact. We examined numerous whole and partial gene deletions with a genotyping approach and observed a depletion of gene disruptions amongst high frequency deletions. Furthermore, we observed differences in the size spectra of SVs originating from distinct formation mechanisms, and constructed a map of SV hotspots formed by common mechanisms. Our analytical framework and SV map serves as a resource for sequencing-based association studies.

A Map of Human Genome Variation from Population-Scale Sequencing.
1000 Genomes Project Consortium, Abecasis GR, Altshuler D, Auton A, Brooks LD, Durbin RM, Gibbs RA, Hurles ME, McVean GA (2010) Nature 467(7319):1061-73.  JOURNAL   PUBMED   REPRINT   WEBSITE     GENETIC VARIATION · GENOMIC SEQUENCING
ABSTRACT: The 1000 Genomes Project aims to provide a deep characterization of human genome sequence variation as a foundation for investigating the relationship between genotype and phenotype. Here we present results of the pilot phase of the project, designed to develop and compare different strategies for genome-wide sequencing with high-throughput platforms. We undertook three projects: low-coverage whole-genome sequencing of 179 individuals from four populations; high-coverage sequencing of two mother-father-child trios; and exon-targeted sequencing of 697 individuals from seven populations. We describe the location, allele frequency and local haplotype structure of approximately 15 million single nucleotide polymorphisms, 1 million short insertions and deletions, and 20,000 structural variants, most of which were previously undescribed. We show that, because we have catalogued the vast majority of common variation, over 95% of the currently accessible variants found in any individual are present in this data set. On average, each person is found to carry approximately 250 to 300 loss-of-function variants in annotated genes and 50 to 100 variants previously implicated in inherited disorders. We demonstrate how these results can be used to inform association and functional studies. From the two trios, we directly estimate the rate of de novo germline base substitution mutations to be approximately 10(-8) per base pair per generation. We explore the data with regard to signatures of natural selection, and identify a marked reduction of genetic variation in the neighbourhood of genes, due to selection at linked sites. These methods and public data will support the next phase of human genetic research.

Genomic sequencing (8)

Whole-Genome Sequencing and Integrative Genomic Analysis Approach on Two 22q11.2 Deletion Syndrome Family Trios for Genotype to Phenotype Correlations.
Chung JH, Cai J, Suskin BG, Zhang Z, Coleman K, Morrow BE (2015) Hum Mutat 36(8):797-807.  JOURNAL   PUBMED   REPRINT     DISEASE STUDY · GENETIC VARIATION · GENOMIC SEQUENCING
ABSTRACT: The 22q11.2 deletion syndrome (22q11DS) affects 1:4,000 live births and presents with highly variable phenotype expressivity. In this study, we developed an analytical approach utilizing whole-genome sequencing (WGS) and integrative analysis to discover genetic modifiers. Our pipeline combined available tools in order to prioritize rare, predicted deleterious, coding and noncoding single-nucleotide variants (SNVs), and insertion/deletions from WGS. We sequenced two unrelated probands with 22q11DS, with contrasting clinical findings, and their unaffected parents. Proband P1 had cognitive impairment, psychotic episodes, anxiety, and tetralogy of Fallot (TOF), whereas proband P2 had juvenile rheumatoid arthritis but no other major clinical findings. In P1, we identified common variants in COMT and PRODH on 22q11.2 as well as rare potentially deleterious DNA variants in other behavioral/neurocognitive genes. We also identified a de novo SNV in ADNP2 (NM_014913.3:c.2243G>C), encoding a neuroprotective protein that may be involved in behavioral disorders. In P2, we identified a novel nonsynonymous SNV in ZFPM2 (NM_012082.3:c.1576C>T), a known causative gene for TOF, which may act as a protective variant downstream of TBX1, haploinsufficiency of which is responsible for congenital heart disease in individuals with 22q11DS.

The Origin, Evolution, and Functional Impact of Short Insertion-Deletion Variants Identified in 179 Human Genomes.
Montgomery SB, Goode DL, Kvikstad E, Albers CA, Zhang ZD, Mu XJ, Ananda G, Howie B, Karczewski KJ, Smith KS, Anaya V, Richardson R, Davis J, 1000 Genomes Project Consortium, MacArthur DG, Sidow A, Duret L, Gerstein M, Makova KD, Marchini J, McVean G, Lunter G (2013) Genome Res 23(5):749-61.  JOURNAL   PUBMED   REPRINT     GENETIC VARIATION · GENOMIC SEQUENCING
ABSTRACT: Short insertions and deletions (indels) are the second most abundant form of human genetic variation, but our understanding of their origins and functional effects lags behind that of other types of variants. Using population-scale sequencing, we have identified a high-quality set of 1.6 million indels from 179 individuals representing three diverse human populations. We show that rates of indel mutagenesis are highly heterogeneous, with 43%-48% of indels occurring in 4.03% of the genome, whereas in the remaining 96% their prevalence is 16 times lower than SNPs. Polymerase slippage can explain upwards of three-fourths of all indels, with the remainder being mostly simple deletions in complex sequence. However, insertions do occur and are significantly associated with pseudo-palindromic sequence features compatible with the fork stalling and template switching (FoSTeS) mechanism more commonly associated with large structural variations. We introduce a quantitative model of polymerase slippage, which enables us to identify indel-hypermutagenic protein-coding genes, some of which are associated with recurrent mutations leading to disease. Accounting for mutational rate heterogeneity due to sequence context, we find that indels across functional sequence are generally subject to stronger purifying selection than SNPs. We find that indel length modulates selection strength, and that indels affecting multiple functionally constrained nucleotides undergo stronger purifying selection. We further find that indels are enriched in associations with gene expression and find evidence for a contribution of nonsense-mediated decay. Finally, we show that indels can be integrated in existing genome-wide association studies (GWAS); although we do not find direct evidence that potentially causal protein-coding indels are enriched with associations to known disease-associated SNPs, our findings suggest that the causal variant underlying some of these associations may be indels.

A Systematic Survey of Loss-of-Function Variants in Human Protein-Coding Genes.
MacArthur DG, Balasubramanian S, Frankish A, Huang N, Morris J, Walter K, Jostins L, Habegger L, Pickrell JK, Montgomery SB, Albers CA, Zhang ZD, Conrad DF, Lunter G, Zheng H, Ayub Q, DePristo MA, Banks E, Hu M, Handsaker RE, Rosenfeld JA, Fromer M, Jin M, Mu XJ, Khurana E, Ye K, Kay M, Saunders GI, Suner MM, Hunt T, Barnes IH, Amid C, Carvalho-Silva DR, Bignell AH, Snow C, Yngvadottir B, Bumpstead S, Cooper DN, Xue Y, Romero IG, 1000 Genomes Project Consortium, Wang J, Li Y, Gibbs RA, McCarroll SA, Dermitzakis ET, Pritchard JK, Barrett JC, Harrow J, Hurles ME, Gerstein MB, Tyler-Smith C (2012) Science 335(6070):823-8.  JOURNAL   PUBMED   REPRINT     GENETIC VARIATION · GENOMIC SEQUENCING
ABSTRACT: Genome-sequencing studies indicate that all humans carry many genetic variants predicted to cause loss of function (LoF) of protein-coding genes, suggesting unexpected redundancy in the human genome. Here we apply stringent filters to 2951 putative LoF variants obtained from 185 human genomes to determine their true prevalence and properties. We estimate that human genomes typically contain ~100 genuine LoF variants with ~20 genes completely inactivated. We identify rare and likely deleterious LoF alleles, including 26 known and 21 predicted severe disease-causing variants, as well as common LoF variants in nonessential genes. We describe functional and evolutionary differences between LoF-tolerant and recessive disease genes and a method for using these differences to prioritize candidate genes found in clinical sequencing studies.

Mapping Copy Number Variation by Population-Scale Genome Sequencing.
Mills RE, Walter K, Stewart C, Handsaker RE, Chen K, Alkan C, Abyzov A, Yoon SC, Ye K, Cheetham RK, Chinwalla A, Conrad DF, Fu Y, Grubert F, Hajirasouliha I, Hormozdiari F, Iakoucheva LM, Iqbal Z, Kang S, Kidd JM, Konkel MK, Korn J, Khurana E, Kural D, Lam HY, Leng J, Li R, Li Y, Lin CY, Luo R, Mu XJ, Nemesh J, Peckham HE, Rausch T, Scally A, Shi X, Stromberg MP, Stütz AM, Urban AE, Walker JA, Wu J, Zhang Y, Zhang ZD, Batzer MA, Ding L, Marth GT, McVean G, Sebat J, Snyder M, Wang J, Ye K, Eichler EE, Gerstein MB, Hurles ME, Lee C, McCarroll SA, Korbel JO, 1000 Genomes Project (2011) Nature 470(7332):59-65.  JOURNAL   PUBMED   REPRINT     GENETIC VARIATION · GENOMIC SEQUENCING
ABSTRACT: Genomic structural variants (SVs) are abundant in humans, differing from other forms of variation in extent, origin and functional impact. Despite progress in SV characterization, the nucleotide resolution architecture of most SVs remains unknown. We constructed a map of unbalanced SVs (that is, copy number variants) based on whole genome DNA sequencing data from 185 human genomes, integrating evidence from complementary SV discovery approaches with extensive experimental validations. Our map encompassed 22,025 deletions and 6,000 additional SVs, including insertions and tandem duplications. Most SVs (53%) were mapped to nucleotide resolution, which facilitated analysing their origin and functional impact. We examined numerous whole and partial gene deletions with a genotyping approach and observed a depletion of gene disruptions amongst high frequency deletions. Furthermore, we observed differences in the size spectra of SVs originating from distinct formation mechanisms, and constructed a map of SV hotspots formed by common mechanisms. Our analytical framework and SV map serves as a resource for sequencing-based association studies.

A Map of Human Genome Variation from Population-Scale Sequencing.
1000 Genomes Project Consortium, Abecasis GR, Altshuler D, Auton A, Brooks LD, Durbin RM, Gibbs RA, Hurles ME, McVean GA (2010) Nature 467(7319):1061-73.  JOURNAL   PUBMED   REPRINT   WEBSITE     GENETIC VARIATION · GENOMIC SEQUENCING
ABSTRACT: The 1000 Genomes Project aims to provide a deep characterization of human genome sequence variation as a foundation for investigating the relationship between genotype and phenotype. Here we present results of the pilot phase of the project, designed to develop and compare different strategies for genome-wide sequencing with high-throughput platforms. We undertook three projects: low-coverage whole-genome sequencing of 179 individuals from four populations; high-coverage sequencing of two mother-father-child trios; and exon-targeted sequencing of 697 individuals from seven populations. We describe the location, allele frequency and local haplotype structure of approximately 15 million single nucleotide polymorphisms, 1 million short insertions and deletions, and 20,000 structural variants, most of which were previously undescribed. We show that, because we have catalogued the vast majority of common variation, over 95% of the currently accessible variants found in any individual are present in this data set. On average, each person is found to carry approximately 250 to 300 loss-of-function variants in annotated genes and 50 to 100 variants previously implicated in inherited disorders. We demonstrate how these results can be used to inform association and functional studies. From the two trios, we directly estimate the rate of de novo germline base substitution mutations to be approximately 10(-8) per base pair per generation. We explore the data with regard to signatures of natural selection, and identify a marked reduction of genetic variation in the neighbourhood of genes, due to selection at linked sites. These methods and public data will support the next phase of human genetic research.

The DNA Sequence, Annotation and Analysis of Human Chromosome 3.
Muzny DM, Scherer SE, Kaul R, Wang J, Yu J, Sudbrak R, Buhay CJ, Chen R, Cree A, Ding Y, Dugan-Rocha S, Gill R, Gunaratne P, Harris RA, Hawes AC, Hernandez J, Hodgson AV, Hume J, Jackson A, Khan ZM, Kovar-Smith C, Lewis LR, Lozado RJ, Metzker ML, Milosavljevic A, Miner GR, Morgan MB, Nazareth LV, Scott G, Sodergren E, Song XZ, Steffen D, Wei S, Wheeler DA, Wright MW, Worley KC, Yuan Y, Zhang Z, Adams CQ, Ansari-Lari MA, Ayele M, Brown MJ, Chen G, Chen Z, Clendenning J, Clerc-Blankenburg KP, Chen R, Chen Z, Davis C, Delgado O, Dinh HH, Dong W, Draper H, Ernst S, Fu G, Gonzalez-Garay ML, Garcia DK, Gillett W, Gu J, Hao B, Haugen E, Havlak P, He X, Hennig S, Hu S, Huang W, Jackson LR, Jacob LS, Kelly SH, Kube M, Levy R, Li Z, Liu B, Liu J, Liu W, Lu J, Maheshwari M, Nguyen BV, Okwuonu GO, Palmeiri A, Pasternak S, Perez LM, Phelps KA, Plopper FJ, Qiang B, Raymond C, Rodriguez R, Saenphimmachak C, Santibanez J, Shen H, Shen Y, Subramanian S, Tabor PE, Verduzco D, Waldron L, Wang J, Wang J, Wang Q, Williams GA, Wong GK, Yao Z, Zhang J, Zhang X, Zhao G, Zhou J, Zhou Y, Nelson D, Lehrach H, Reinhardt R, Naylor SL, Yang H, Olson M, Weinstock G, Gibbs RA (2006) Nature 440(7088):1194-8.  JOURNAL   PUBMED   REPRINT     GENOMIC SEQUENCING
ABSTRACT: After the completion of a draft human genome sequence, the International Human Genome Sequencing Consortium has proceeded to finish and annotate each of the 24 chromosomes comprising the human genome. Here we describe the sequencing and analysis of human chromosome 3, one of the largest human chromosomes. Chromosome 3 comprises just four contigs, one of which currently represents the longest unbroken stretch of finished DNA sequence known so far. The chromosome is remarkable in having the lowest rate of segmental duplication in the genome. It also includes a chemokine receptor gene cluster as well as numerous loci involved in multiple human cancers such as the gene encoding FHIT, which contains the most common constitutive fragile site in the genome, FRA3B. Using genomic sequence from chimpanzee and rhesus macaque, we were able to characterize the breakpoints defining a large pericentric inversion that occurred some time after the split of Homininae from Ponginae, and propose an evolutionary history of the inversion.

The Finished DNA Sequence of Human Chromosome 12.
Scherer SE, Muzny DM, Buhay CJ, Chen R, Cree A, Ding Y, Dugan-Rocha S, Gill R, Gunaratne P, Harris RA, Hawes AC, Hernandez J, Hodgson AV, Hume J, Jackson A, Khan ZM, Kovar-Smith C, Lewis LR, Lozado RJ, Metzker ML, Milosavljevic A, Miner GR, Montgomery KT, Morgan MB, Nazareth LV, Scott G, Sodergren E, Song XZ, Steffen D, Lovering RC, Wheeler DA, Worley KC, Yuan Y, Zhang Z, Adams CQ, Ansari-Lari MA, Ayele M, Brown MJ, Chen G, Chen Z, Clerc-Blankenburg KP, Davis C, Delgado O, Dinh HH, Draper H, Gonzalez-Garay ML, Havlak P, Jackson LR, Jacob LS, Kelly SH, Li L, Li Z, Liu J, Liu W, Lu J, Maheshwari M, Nguyen BV, Okwuonu GO, Pasternak S, Perez LM, Plopper FJ, Santibanez J, Shen H, Tabor PE, Verduzco D, Waldron L, Wang Q, Williams GA, Zhang J, Zhou J, Allen CC, Amin AG, Anyalebechi V, Bailey M, Barbaria JA, Bimage KE, Bryant NP, Burch PE, Burkett CE, Burrell KL, Calderon E, Cardenas V, Carter K, Casias K, Cavazos I, Cavazos SR, Ceasar H, Chacko J, Chan SN, Chavez D, Christopoulos C, Chu J, Cockrell R, Cox CD, Dang M, Dathorne SR, David R, Davis CM, Davy-Carroll L, Deshazo DR, Donlin JE, D'Souza L, Eaves KA, Simons R, Emery-Cohen AJ, Escotto M, Flagg N, Forbes LD, Gabisi AM, Garza M, Hamilton C, Henderson N, Hernandez O, Hines S, Hogues ME, Huang M, Idlebird DG, Johnson R, Jolivet A, Jones S, Kagan R, King LM, Leal B, Lebow H, Lee S, LeVan JM, Lewis LC, London P, Lorensuhewa LM, Loulseged H, Lovett DA, Lucier A, Lucier RL, Ma J, Madu RC, Mapua P, Martindale AD, Martinez E, Massey E, Mawhiney S, Meador MG, Mendez S, Mercado C, Mercado IC, Merritt CE, Miner ZL, Minja E, Mitchell T, Mohabbat F, Mohabbat K, Montgomery B, Moore N, Morris S, Munidasa M, Ngo RN, Nguyen NB, Nickerson E, Nwaokelemeh OO, Nwokenkwo S, Obregon M, Oguh M, Oragunye N, Oviedo RJ, Parish BJ, Parker DN, Parrish J, Parks KL, Paul HA, Payton BA, Perez A, Perrin W, Pickens A, Primus EL, Pu LL, Puazo M, Quiles MM, Quiroz JB, Rabata D, Reeves K, Ruiz SJ, Shao H, Sisson I, Sonaike T, Sorelle RP, Sutton AE, Svatek AF, Svetz LA, Tamerisa KS, Taylor TR, Teague B, Thomas N, Thorn RD, Trejos ZY, Trevino BK, Ukegbu ON, Urban JB, Vasquez LI, Vera VA, Villasana DM, Wang L, Ward-Moore S, Warren JT, Wei X, White F, Williamson AL, Wleczyk R, Wooden HS, Wooden SH, Yen J, Yoon L, Yoon V, Zorrilla SE, Nelson D, Kucherlapati R, Weinstock G, Gibbs RA, Baylor College of Medicine Human Genome Sequencing Center Sequence Production Team (2006) Nature 440(7082):346-51.  JOURNAL   PUBMED   REPRINT     GENOMIC SEQUENCING
ABSTRACT: Human chromosome 12 contains more than 1,400 coding genes and 487 loci that have been directly implicated in human disease. The q arm of chromosome 12 contains one of the largest blocks of linkage disequilibrium found in the human genome. Here we present the finished sequence of human chromosome 12, which has been finished to high quality and spans approximately 132 megabases, representing approximately 4.5% of the human genome. Alignment of the human chromosome 12 sequence across vertebrates reveals the origin of individual segments in chicken, and a unique history of rearrangement through rodent and primate lineages. The rate of base substitutions in recent evolutionary history shows an overall slowing in hominids compared with primates and rodents.

Genome Sequence of the Brown Norway Rat Yields Insights into Mammalian Evolution.
Gibbs RA, Weinstock GM, Metzker ML, Muzny DM, Sodergren EJ, Scherer S, Scott G, Steffen D, Worley KC, Burch PE, Okwuonu G, Hines S, Lewis L, DeRamo C, Delgado O, Dugan-Rocha S, Miner G, Morgan M, Hawes A, Gill R, Celera, Holt RA, Adams MD, Amanatides PG, Baden-Tillson H, Barnstead M, Chin S, Evans CA, Ferriera S, Fosler C, Glodek A, Gu Z, Jennings D, Kraft CL, Nguyen T, Pfannkoch CM, Sitter C, Sutton GG, Venter JC, Woodage T, Smith D, Lee HM, Gustafson E, Cahill P, Kana A, Doucette-Stamm L, Weinstock K, Fechtel K, Weiss RB, Dunn DM, Green ED, Blakesley RW, Bouffard GG, De Jong PJ, Osoegawa K, Zhu B, Marra M, Schein J, Bosdet I, Fjell C, Jones S, Krzywinski M, Mathewson C, Siddiqui A, Wye N, McPherson J, Zhao S, Fraser CM, Shetty J, Shatsman S, Geer K, Chen Y, Abramzon S, Nierman WC, Havlak PH, Chen R, Durbin KJ, Simons R, Ren Y, Song XZ, Li B, Liu Y, Qin X, Cawley S, Worley KC, Cooney AJ, D'Souza LM, Martin K, Wu JQ, Gonzalez-Garay ML, Jackson AR, Kalafus KJ, McLeod MP, Milosavljevic A, Virk D, Volkov A, Wheeler DA, Zhang Z, Bailey JA, Eichler EE, Tuzun E, Birney E, Mongin E, Ureta-Vidal A, Woodwark C, Zdobnov E, Bork P, Suyama M, Torrents D, Alexandersson M, Trask BJ, Young JM, Huang H, Wang H, Xing H, Daniels S, Gietzen D, Schmidt J, Stevens K, Vitt U, Wingrove J, Camara F, Mar Albà M, Abril JF, Guigo R, Smit A, Dubchak I, Rubin EM, Couronne O, Poliakov A, Hübner N, Ganten D, Goesele C, Hummel O, Kreitler T, Lee YA, Monti J, Schulz H, Zimdahl H, Himmelbauer H, Lehrach H, Jacob HJ, Bromberg S, Gullings-Handley J, Jensen-Seaman MI, Kwitek AE, Lazar J, Pasko D, Tonellato PJ, Twigger S, Ponting CP, Duarte JM, Rice S, Goodstadt L, Beatson SA, Emes RD, Winter EE, Webber C, Brandt P, Nyakatura G, Adetobi M, Chiaromonte F, Elnitski L, Eswara P, Hardison RC, Hou M, Kolbe D, Makova K, Miller W, Nekrutenko A, Riemer C, Schwartz S, Taylor J, Yang S, Zhang Y, Lindpaintner K, Andrews TD, Caccamo M, Clamp M, Clarke L, Curwen V, Durbin R, Eyras E, Searle SM, Cooper GM, Batzoglou S, Brudno M, Sidow A, Stone EA, Venter JC, Payseur BA, Bourque G, López-Otín C, Puente XS, Chakrabarti K, Chatterji S, Dewey C, Pachter L, Bray N, Yap VB, Caspi A, Tesler G, Pevzner PA, Haussler D, Roskin KM, Baertsch R, Clawson H, Furey TS, Hinrichs AS, Karolchik D, Kent WJ, Rosenbloom KR, Trumbower H, Weirauch M, Cooper DN, Stenson PD, Ma B, Brent M, Arumugam M, Shteynberg D, Copley RR, Taylor MS, Riethman H, Mudunuri U, Peterson J, Guyer M, Felsenfeld A, Old S, Mockrin S, Collins F, Rat Genome Sequencing Project Consortium (2004) Nature 428(6982):493-521.  JOURNAL   PUBMED   REPRINT     GENOMIC SEQUENCING
ABSTRACT: The laboratory rat (Rattus norvegicus) is an indispensable tool in experimental medicine and drug development, having made inestimable contributions to human health. We report here the genome sequence of the Brown Norway (BN) rat strain. The sequence represents a high-quality 'draft' covering over 90% of the genome. The BN rat sequence is the third complete mammalian genome to be deciphered, and three-way comparisons with the human and mouse genomes resolve details of mammalian evolution. This first comprehensive analysis includes genes and proteins and their relation to human disease, repeated sequences, comparative genome-wide studies of mammalian orthologous chromosomal regions and rearrangement breakpoints, reconstruction of ancestral karyotypes and the events leading to existing species, rates of variation, and lineage-specific and lineage-independent evolutionary events such as expansion of gene families, orthology relations and protein evolution.

Pseudogene (2)

Identification and Analysis of Unitary Pseudogenes: Historic and Contemporary Gene Losses in Humans and Other Primates.
Zhang ZD, Frankish A, Hunt T, Harrow J, Gerstein M (2010) Genome Biol 11(3):R26.  JOURNAL   PUBMED   REPRINT     COMPARATIVE GENOMICS · EVOLUTIONARY GENOMICS · PSEUDOGENE · SOFTWARE / PIPELINE / DATABASE
ABSTRACT: BACKGROUND: Unitary pseudogenes are a class of unprocessed pseudogenes without functioning counterparts in the genome. They constitute only a small fraction of annotated pseudogenes in the human genome. However, as they represent distinct functional losses over time, they shed light on the unique features of humans in primate evolution. RESULTS: We have developed a pipeline to detect human unitary pseudogenes through analyzing the global inventory of orthologs between the human genome and its mammalian relatives. We focus on gene losses along the human lineage after the divergence from rodents about 75 million years ago. In total, we identify 76 unitary pseudogenes, including previously annotated ones, and many novel ones. By comparing each of these to its functioning ortholog in other mammals, we can approximately date the creation of each unitary pseudogene (that is, the gene 'death date') and show that for our group of 76, the functional genes appear to be disabled at a fairly uniform rate throughout primate evolution - not all at once, correlated, for instance, with the 'Alu burst'. Furthermore, we identify 11 unitary pseudogenes that are polymorphic - that is, they have both nonfunctional and functional alleles currently segregating in the human population. Comparing them with their orthologs in other primates, we find that two of them are in fact pseudogenes in non-human primates, suggesting that they represent cases of a gene being resurrected in the human lineage. CONCLUSIONS: This analysis of unitary pseudogenes provides insights into the evolutionary constraints faced by different organisms and the timescales of functional gene loss in humans.

Analysis of Nuclear Receptor Pseudogenes in Vertebrates: How the Silent Tell Their Stories.
Zhang ZD, Cayting P, Weinstock G, Gerstein M (2008) Mol Biol Evol 25(1):131-43.  JOURNAL   PUBMED   REPRINT   WEBSITE     FUNCTIONAL GENOMICS · PSEUDOGENE
ABSTRACT: Transcription factor pseudogenes have not been systematically studied before. Nuclear receptors (NRs) constitute one of the largest groups of transcription factors in animals (e.g., 48 NRs in human). The availability of whole-genome sequences enables a global inventory of the NR pseudogenes in a number of vertebrate model organisms. Here we identify the NR pseudogenes in 8 vertebrate organisms and make our results available online at http://www.pseudogene.org/nr. The assignments reveal that NR pseudogenes as a group have characteristics related to generation and distribution contrary to expectations derived from previous large-scale pseudogene studies. In particular, 1) despite its large size, the NR gene family has only a very small number of pseudogenes in each of the vertebrate genomes examined; 2) despite the low transcription levels of NR genes, except for one, all other NR pseudogenes identified in this study are retropseudogenes; and 3) no duplicated NR pseudogenes are found, contrary to the fact that the NR gene family was expanded through several waves of gene duplication events. Our analyses further reveal a number of interesting aspects of NR pseudogenes. Specifically, through careful sequence analysis, we identify remnant introns in 2 mouse retropseudogenes, psiRev-erbbeta and psiLRH1. Generated from partially processed pre-mRNAs, they appear to be rare examples of highly unusual "semiprocessed" pseudogenes. Second, by comparing the genomic sequences, we uncover a pseudogene that is unique to the human lineage relative to chimpanzee. Generated by a recent duplication of a segment in the human genome, this pseudogene is a "duplicated-processed" pseudogene, belonging to a new pseudogene species. Finally, FXRbeta was nonfunctionalized in the human lineage and thus appears to be an example of a rare unitary pseudogene. By comparing orthologous sequences, we dated the FXR-FXRbeta duplication and the nonfunctionalization of FXRbeta in primates.

Review / Perspective (9)

Chromosome Instability and Aneuploidy in the Mammalian Brain.
Albert O, Sun S, Huttner A, Zhang Z, Suh Y, Campisi J, Vijg J, Montagna C (2023) Chromosome Res 31(4):32.  JOURNAL   PUBMED   REPRINT     REVIEW / PERSPECTIVE
ABSTRACT: This review investigates the role of aneuploidy and chromosome instability (CIN) in the aging brain. Aneuploidy refers to an abnormal chromosomal count, deviating from the normal diploid set. It can manifest as either a deficiency or excess of chromosomes. CIN encompasses a broader range of chromosomal alterations, including aneuploidy as well as structural modifications in DNA. We provide an overview of the state-of-the-art methodologies utilized for studying aneuploidy and CIN in non-tumor somatic tissues devoid of clonally expanded populations of aneuploid cells.CIN and aneuploidy, well-established hallmarks of cancer cells, are also associated with the aging process. In non-transformed cells, aneuploidy can contribute to functional impairment and developmental disorders. Despite the importance of understanding the prevalence and specific consequences of aneuploidy and CIN in the aging brain, these aspects remain incompletely understood, emphasizing the need for further scientific investigations.This comprehensive review consolidates the present understanding, addresses discrepancies in the literature, and provides valuable insights for future research efforts.

◸ NEWS AND VIEWS 
Unravelling Genetic Components of Longevity.
Jabalameli MR, Zhang ZD (2022) Nat Aging 2(1):5-6.  JOURNAL   PUBMED   REPRINT     AGING · DISEASE STUDY · GENETIC VARIATION · REVIEW / PERSPECTIVE
ABSTRACT: Many aging-related traits share a common genetic component. How to disentangle it from the trait-specific effects has remained largely unexplored. A new study in Nature Aging uses an analysis framework for isolating the shared genetic component in genome-wide association studies of aging-related traits and identifies genomic loci that contribute to aging.

Genetics of Extreme Human Longevity to Guide Drug Discovery for Healthy Ageing.
Zhang ZD, Milman S, Lin JR, Wierbowski S, Yu H, Barzilai N, Gorbunova V, Ladiges WC, Niedernhofer LJ, Suh Y, Robbins PD, Vijg J (2020) Nat Metab 2(8):663-672.  JOURNAL   PUBMED   REPRINT     AGING · DISEASE STUDY · GENETIC VARIATION · REVIEW / PERSPECTIVE
ABSTRACT: Ageing is the greatest risk factor for most common chronic human diseases, and it therefore is a logical target for developing interventions to prevent, mitigate or reverse multiple age-related morbidities. Over the past two decades, genetic and pharmacologic interventions targeting conserved pathways of growth and metabolism have consistently led to substantial extension of the lifespan and healthspan in model organisms as diverse as nematodes, flies and mice. Recent genetic analysis of long-lived individuals is revealing common and rare variants enriched in these same conserved pathways that significantly correlate with longevity. In this Perspective, we summarize recent insights into the genetics of extreme human longevity and propose the use of this rare phenotype to identify genetic variants as molecular targets for gaining insight into the physiology of healthy ageing and the development of new therapies to extend the human healthspan.

Schizophrenia Risk Genes.
Zhang ZD (2017) Access Science.  JOURNAL   WEBSITE     REVIEW / PERSPECTIVE
ABSTRACT: Schizophrenia risk genes increase the likelihood of developing schizophrenia, a severe mental disorder with a large genetic component. Crucial to providing insights into the underlying disease mechanisms and for identifying new drug targets, they have been identified in neurotransmitter systems and more recently by large-scale genetic studies.

A Neurogenetic Model for the Study of Schizophrenia Spectrum Disorders: The International 22q11.2 Deletion Syndrome Brain Behavior Consortium.
Gur RE, Bassett AS, McDonald-McGinn DM, Bearden CE, Chow E, Emanuel BS, Owen M, Swillen A, Van den Bree M, Vermeesch J, Vorstman JAS, Warren S, Lehner T, Morrow B (2017) Mol Psychiatry 22(12):1664-1672.  JOURNAL   PUBMED   REPRINT   WEBSITE     DISEASE STUDY · REVIEW / PERSPECTIVE
ABSTRACT: Rare copy number variants contribute significantly to the risk for schizophrenia, with the 22q11.2 locus consistently implicated. Individuals with the 22q11.2 deletion syndrome (22q11DS) have an estimated 25-fold increased risk for schizophrenia spectrum disorders, compared to individuals in the general population. The International 22q11DS Brain Behavior Consortium is examining this highly informative neurogenetic syndrome phenotypically and genomically. Here we detail the procedures of the effort to characterize the neuropsychiatric and neurobehavioral phenotypes associated with 22q11DS, focusing on schizophrenia and subthreshold expression of psychosis. The genomic approach includes a combination of whole-genome sequencing and genome-wide microarray technologies, allowing the investigation of all possible DNA variation and gene pathways influencing the schizophrenia-relevant phenotypic expression. A phenotypically rich data set provides a psychiatrically well-characterized sample of unprecedented size (n=1616) that informs the neurobehavioral developmental course of 22q11DS. This combined set of phenotypic and genomic data will enable hypothesis testing to elucidate the mechanisms underlying the pathogenesis of schizophrenia spectrum disorders.

From Gene Expression to Disease Phenotypes: Network-Based Approaches to Study Complex Human Diseases.
Zhang Q, Zhang W, Nogales-Cadenas R, Lin JR, Cai Y, Zhang ZD (2015) Transcriptomics and Gene Regulation Translational Bioinformatics 9:115-140.  JOURNAL   REPRINT     REVIEW / PERSPECTIVE
ABSTRACT: Gene expression is a fundamental biological process under tight regulation at all levels in normal cells. Its dysregulation can cause abnormal cell behaviors and result in diseases, and thus gene expression profiling and analysis have been widely used to provide the first clue about the molecular mechanisms of human diseases. Because genes and their products interact with and regulate one another, it is essential to analyze gene expression data and understand the genetics of disease in a biological network context. In this chapter, we first introduce the state-of-the-art gene expression analysis with network integration and the joint analysis of mRNA and miRNA expression to understand disease regulatory mechanisms, and then discuss how disease genes are predicted by incorporating knowledge of gene regulation and characterized in biological networks.

Comparative Genetics of Longevity and Cancer: Insights from Long-Lived Rodents.
Gorbunova V, Seluanov A, Zhang Z, Gladyshev VN, Vijg J (2014) Nat Rev Genet 15(8):531-40.  JOURNAL   PUBMED   REPRINT     AGING · REVIEW / PERSPECTIVE
ABSTRACT: Mammals have evolved a remarkable diversity of ageing rates. Within the single order of Rodentia, maximum lifespans range from 4 years in mice to 32 years in naked mole rats. Cancer rates also differ substantially between cancer-prone mice and almost cancer-proof naked mole rats and blind mole rats. Recent progress in rodent comparative biology, together with the emergence of whole-genome sequence information, has opened opportunities for the discovery of genetic factors that control longevity and cancer susceptibility.

A Brief Introduction to Tiling Microarrays: Principles, Concepts, and Applications.
Lemetre C, Zhang ZD (2013) Methods Mol Biol 1067:3-19.  JOURNAL   PUBMED   REPRINT     REVIEW / PERSPECTIVE
ABSTRACT: Technological achievements have always contributed to the advancement of biomedical research. It has never been more so than in recent times, when the development and application of innovative cutting-edge technologies have transformed biology into a data-rich quantitative science. This stunning revolution in biology primarily ensued from the emergence of microarrays over two decades ago. The completion of whole-genome sequencing projects and the advance in microarray manufacturing technologies enabled the development of tiling microarrays, which gave unprecedented genomic coverage. Since their first description, several types of application of tiling arrays have emerged, each aiming to tackle a different biological problem. Although numerous algorithms have already been developed to analyze microarray data, new method development is still needed not only for better performance but also for integration of available microarray data sets, which without doubt constitute one of the largest collections of biological data ever generated. In this chapter we first introduce the principles behind the emergence and the development of tiling microarrays, and then discuss with some examples how they are used to investigate different biological problems.

What Is a Gene, Post-Encode? History and Updated Definition.
Gerstein MB, Bruce C, Rozowsky JS, Zheng D, Du J, Korbel JO, Emanuelsson O, Zhang ZD, Weissman S, Snyder M (2007) Genome Res 17(6):669-81.  JOURNAL   PUBMED   REPRINT   POSTER     FUNCTIONAL GENOMICS · REVIEW / PERSPECTIVE
ABSTRACT: While sequencing of the human genome surprised us with how many protein-coding genes there are, it did not fundamentally change our perspective on what a gene is. In contrast, the complex patterns of dispersed regulation and pervasive transcription uncovered by the ENCODE project, together with non-genic conservation and the abundance of noncoding RNA genes, have challenged the notion of the gene. To illustrate this, we review the evolution of operational definitions of a gene over the past century--from the abstract elements of heredity of Mendel and Morgan to the present-day ORFs enumerated in the sequence databanks. We then summarize the current ENCODE findings and provide a computational metaphor for the complexity. Finally, we propose a tentative update to the definition of a gene: A gene is a union of genomic sequences encoding a coherent set of potentially overlapping functional products. Our definition side-steps the complexities of regulation and transcription by removing the former altogether from the definition and arguing that final, functional gene products (rather than intermediate transcripts) should be used to group together entities associated with a single gene. It also manifests how integral the concept of biological function is in defining genes.

Software / Pipeline / Database (15)

Direct Probabilistic Quantification of Mosaic Loss of Chromosome Y from Sequencing Data
Lin J-R, Chang Y-C, Maslov AY, Song Y, Gao T, Shan J, Bennett D, Milman S, Barzilai N, Vijg J, Montagna C, Zhang ZD (2026) bioRxiv 2026.06.26.734767.  JOURNAL     ANALYTICAL METHOD · SOFTWARE / PIPELINE / DATABASE
ABSTRACT: Loss of chromosome Y (LOY) is the most common aneuploidy in aging men and is increasingly recognized as a marker of aging and genomic instability. Because LOY occurs in mosaic form, its degree reflects the fraction of cells lacking the Y chromosome. Existing SNP-array- and sequencing-based methods rely largely on single genomic features and indirect transformations to estimate this fraction. We developed BaySeq-Y, a Bayesian method that directly estimates LOY mosaicism from sequencing data using VCF files with read depth (DP) and allelic depth (AD). Within a rigorous Bayesian framework, BaySeq-Y integrates complementary LOY-associated genomic features, including decreased read depth and allelic imbalance, and can additionally leverage haplotype phasing to improve precision. In simulations and fluorescence in situ hybridization validation (FISH), BaySeq-Y provided accurate estimates and outperformed existing methods. Applications to ROSMAP and GTEx supported its biological relevance through transcriptomic validation, demonstrating its utility for quantifying LOY across diverse sequencing datasets.

Bayesian Estimation of Mosaic Loss of Chromosome Y from Bulk RNA Sequencing Data
Lin J-R, Zhang ZD (2026) bioRxiv 2026.05.20.726153.  JOURNAL     ANALYTICAL METHOD · SOFTWARE / PIPELINE / DATABASE
ABSTRACT: Mosaic loss of chromosome Y (LOY) is a common age-associated somatic alteration in men and is typically measured from DNA-based assays. Many cohorts, however, contain bulk RNA-seq data without matched DNA-based LOY measurements. We developed a Bayesian framework to estimate the fraction of cells with LOY from male bulk RNA-seq by modeling reduced Y-linked gene expression relative to expected expression after adjustment for age, expression covariates, and autosomal/X-linked control genes. In 377 male GTEx samples, individual Y-linked genes showed negative correlations with separately obtained DNA-based LOY measurements, supporting a shared Y-expression depletion signal. The primary fast empirical Bayes estimator achieved a Pearson correlation of 0.678 with measured LOY, a mean absolute error of 1.79%, a root mean squared error of 3.72%, and 95.2% empirical coverage of measured LOY. Performance was strongest for identifying large LOY events, with an AUC of 0.964 for measured LOY greater than 20%, while fine ranking among low-LOY samples remained uncertain. A mixture/PCA hierarchical Bayesian sensitivity model provided similar validation performance and interpretable posterior quantities but did not improve point estimation. Leave-one-Y-gene-out and prior-sensitivity analyses showed that the signal was distributed across multiple Y-linked transcripts and that prior shrinkage affected calibration. In an external whole-blood RNA-seq dataset without measured LOY, estimated LOY showed a modest age-related increase, but ex vivo immune stimulation shifted RNA-derived LOY estimates and reduced multiple Y-linked transcripts, indicating transcriptional confounding. These results show that bulk RNA-seq contains usable information about LOY, especially for larger events, but RNA-derived LOY should be interpreted as a probabilistic transcriptome-based estimate rather than a direct substitute for DNA-based mosaicism measurement.

Protocol for Gene Annotation, Prediction, and Validation of Genomic Gene Expansion.
Zhang Q, Zhang ZD (2022) STAR Protoc 3(4):101692.  JOURNAL   PUBMED   REPRINT     ANALYTICAL METHOD · COMPARATIVE GENOMICS · EVOLUTIONARY GENOMICS · SOFTWARE / PIPELINE / DATABASE
ABSTRACT: Although gene expansion plays an important role in evolution, its identification remains a challenge due to potential errors in genome assembly and annotation. Here, we describe a detailed step-by-step protocol for gene annotation, prediction of genomic gene expansion, and its computational and experimental validation. Finally, we also detail steps to discover functionality of each copy of replicated genes. For complete details on the use and execution of this protocol, please refer to Zhang et al. (2021).

HEDD: Human Enhancer Disease Database.
Wang Z, Zhang Q, Zhang W, Lin JR, Cai Y, Mitra J, Zhang ZD (2018) Nucleic Acids Res 46(D1):D113-D120.  JOURNAL   PUBMED   REPRINT   WEBSITE     FUNCTIONAL GENOMICS · SOFTWARE / PIPELINE / DATABASE
ABSTRACT: Enhancers, as specialized genomic cis-regulatory elements, activate transcription of their target genes and play an important role in pathogenesis of many human complex diseases. Despite recent systematic identification of them in the human genome, currently there is an urgent need for comprehensive annotation databases of human enhancers with a focus on their disease connections. In response, we built the Human Enhancer Disease Database (HEDD) to facilitate studies of enhancers and their potential roles in human complex diseases. HEDD currently provides comprehensive genomic information for ∼2.8 million human enhancers identified by ENCODE, FANTOM5 and RoadMap with disease association scores based on enhancer-gene and gene-disease connections. It also provides Web-based analytical tools to visualize enhancer networks and score enhancers given a set of selected genes in a specific gene network. HEDD is freely accessible at http://zdzlab.einstein.yu.edu/1/hedd.php.

SubNet: A Java Application for Subnetwork Extraction.
Lemetre C, Zhang Q, Zhang ZD (2013) Bioinformatics 29(19):2509-11.  JOURNAL   PUBMED   REPRINT   WEBSITE     ANALYTICAL METHOD · SOFTWARE / PIPELINE / DATABASE · SYSTEMS BIOLOGY
ABSTRACT: SUMMARY: The extraction of targeted subnetworks is a powerful way to identify functional modules and pathways within complex networks. Here, we present SubNet, a Java-based stand-alone program for extracting subnetworks, given a basal network and a set of selected nodes. Designed with a graphical user-friendly interface, SubNet combines four different extraction methods, which offer the possibility to interrogate a biological network according to the question investigated. Of note, we developed a method based on the highly successful Google PageRank algorithm to extract the subnetwork using the node centrality metric, to which possible node weights of the selected genes can be incorporated. AVAILABILITY: http://www.zdzlab.org/1/subnet.html

The Einstein Genome Gateway Using WASP - a High Throughput Multi-Layered Life Sciences Portal for XSEDE.
Golden A, McLellan AS, Dubin RA, Jing Q, O Broin P, Moskowitz D, Zhang Z, Suzuki M, Hargitai J, Calder RB, Greally JM (2012) Stud Health Technol Inform 175:182-91.  PUBMED     SOFTWARE / PIPELINE / DATABASE
ABSTRACT: Massively-parallel sequencing (MPS) technologies and their diverse applications in genomics and epigenomics research have yielded enormous new insights into the physiology and pathophysiology of the human genome. The biggest hurdle remains the magnitude and diversity of the datasets generated, compromising our ability to manage, organize, process and ultimately analyse data. The Wiki-based Automated Sequence Processor (WASP), developed at the Albert Einstein College of Medicine (hereafter Einstein), uniquely manages to tightly couple the sequencing platform, the sequencing assay, sample metadata and the automated workflows deployed on a heterogeneous high performance computing cluster infrastructure that yield sequenced, quality-controlled and 'mapped' sequence data, all within the one operating environment accessible by a web-based GUI interface. WASP at Einstein processes 4-6 TB of data per week and since its production cycle commenced it has processed ~ 1 PB of data overall and has revolutionized user interactivity with these new genomic technologies, who remain blissfully unaware of the data storage, management and most importantly processing services they request. The abstraction of such computational complexity for the user in effect makes WASP an ideal middleware solution, and an appropriate basis for the development of a grid-enabled resource - the Einstein Genome Gateway - as part of the Extreme Science and Engineering Discovery Environment (XSEDE) program. In this paper we discuss the existing WASP system, its proposed middleware role, and its planned interaction with XSEDE to form the Einstein Genome Gateway.

Identification of Genomic Indels and Structural Variations Using Split Reads.
Zhang ZD, Du J, Lam H, Abyzov A, Urban AE, Snyder M, Gerstein M (2011) BMC Genomics 12:375.  JOURNAL   PUBMED   REPRINT     ANALYTICAL METHOD · SOFTWARE / PIPELINE / DATABASE
ABSTRACT: BACKGROUND: Recent studies have demonstrated the genetic significance of insertions, deletions, and other more complex structural variants (SVs) in the human population. With the development of the next-generation sequencing technologies, high-throughput surveys of SVs on the whole-genome level have become possible. Here we present split-read identification, calibrated (SRiC), a sequence-based method for SV detection. RESULTS: We start by mapping each read to the reference genome in standard fashion using gapped alignment. Then to identify SVs, we score each of the many initial mappings with an assessment strategy designed to take into account both sequencing and alignment errors (e.g. scoring more highly events gapped in the center of a read). All current SV calling methods have multilevel biases in their identifications due to both experimental and computational limitations (e.g. calling more deletions than insertions). A key aspect of our approach is that we calibrate all our calls against synthetic data sets generated from simulations of high-throughput sequencing (with realistic error models). This allows us to calculate sensitivity and the positive predictive value under different parameter-value scenarios and for different classes of events (e.g. long deletions vs. short insertions). We run our calculations on representative data from the 1000 Genomes Project. Coupling the observed numbers of events on chromosome 1 with the calibrations gleaned from the simulations (for different length events) allows us to construct a relatively unbiased estimate for the total number of SVs in the human genome across a wide range of length scales. We estimate in particular that an individual genome contains ~670,000 indels/SVs. CONCLUSIONS: Compared with the existing read-depth and read-pair approaches for SV identification, our method can pinpoint the exact breakpoints of SV events, reveal the actual sequence content of insertions, and cover the whole size spectrum for deletions. Moreover, with the advent of the third-generation sequencing technologies that produce longer reads, we expect our method to be even more useful.

ACT: Aggregation and Correlation Toolbox for Analyses of Genome Tracks.
Jee J, Rozowsky J, Yip KY, Lochovsky L, Bjornson R, Zhong G, Zhang Z, Fu Y, Wang J, Weng Z, Gerstein M (2011) Bioinformatics 27(8):1152-4.  JOURNAL   PUBMED   REPRINT   WEBSITE     ANALYTICAL METHOD · SOFTWARE / PIPELINE / DATABASE
ABSTRACT: UNLABELLED: We have implemented aggregation and correlation toolbox (ACT), an efficient, multifaceted toolbox for analyzing continuous signal and discrete region tracks from high-throughput genomic experiments, such as RNA-seq or ChIP-chip signal profiles from the ENCODE and modENCODE projects, or lists of single nucleotide polymorphisms from the 1000 genomes project. It is able to generate aggregate profiles of a given track around a set of specified anchor points, such as transcription start sites. It is also able to correlate related tracks and analyze them for saturation--i.e. how much of a certain feature is covered with each new succeeding experiment. The ACT site contains downloadable code in a variety of formats, interactive web servers (for use on small quantities of data), example datasets, documentation and a gallery of outputs. Here, we explain the components of the toolbox in more detail and apply them in various contexts. AVAILABILITY: ACT is available at http://act.gersteinlab.org CONTACT: pi@gersteinlab.org.

Detection of Copy Number Variation from Array Intensity and Sequencing Read Depth Using a Stepwise Bayesian Model.
Zhang ZD, Gerstein MB (2010) BMC Bioinformatics 11:539.  JOURNAL   PUBMED   REPRINT     ANALYTICAL METHOD · SOFTWARE / PIPELINE / DATABASE
ABSTRACT: BACKGROUND: Copy number variants (CNVs) have been demonstrated to occur at a high frequency and are now widely believed to make a significant contribution to the phenotypic variation in human populations. Array-based comparative genomic hybridization (array-CGH) and newly developed read-depth approach through ultrahigh throughput genomic sequencing both provide rapid, robust, and comprehensive methods to identify CNVs on a whole-genome scale. RESULTS: We developed a Bayesian statistical analysis algorithm for the detection of CNVs from both types of genomic data. The algorithm can analyze such data obtained from PCR-based bacterial artificial chromosome arrays, high-density oligonucleotide arrays, and more recently developed high-throughput DNA sequencing. Treating parameters--e.g., the number of CNVs, the position of each CNV, and the data noise level--that define the underlying data generating process as random variables, our approach derives the posterior distribution of the genomic CNV structure given the observed data. Sampling from the posterior distribution using a Markov chain Monte Carlo method, we get not only best estimates for these unknown parameters but also Bayesian credible intervals for the estimates. We illustrate the characteristics of our algorithm by applying it to both synthetic and experimental data sets in comparison to other segmentation algorithms. CONCLUSIONS: In particular, the synthetic data comparison shows that our method is more sensitive than other approaches at low false positive rates. Furthermore, given its Bayesian origin, our method can also be seen as a technique to refine CNVs identified by fast point-estimate methods and also as a framework to integrate array-CGH and sequencing data with other CNV-related biological knowledge, all through informative priors.

Identification and Analysis of Unitary Pseudogenes: Historic and Contemporary Gene Losses in Humans and Other Primates.
Zhang ZD, Frankish A, Hunt T, Harrow J, Gerstein M (2010) Genome Biol 11(3):R26.  JOURNAL   PUBMED   REPRINT     COMPARATIVE GENOMICS · EVOLUTIONARY GENOMICS · PSEUDOGENE · SOFTWARE / PIPELINE / DATABASE
ABSTRACT: BACKGROUND: Unitary pseudogenes are a class of unprocessed pseudogenes without functioning counterparts in the genome. They constitute only a small fraction of annotated pseudogenes in the human genome. However, as they represent distinct functional losses over time, they shed light on the unique features of humans in primate evolution. RESULTS: We have developed a pipeline to detect human unitary pseudogenes through analyzing the global inventory of orthologs between the human genome and its mammalian relatives. We focus on gene losses along the human lineage after the divergence from rodents about 75 million years ago. In total, we identify 76 unitary pseudogenes, including previously annotated ones, and many novel ones. By comparing each of these to its functioning ortholog in other mammals, we can approximately date the creation of each unitary pseudogene (that is, the gene 'death date') and show that for our group of 76, the functional genes appear to be disabled at a fairly uniform rate throughout primate evolution - not all at once, correlated, for instance, with the 'Alu burst'. Furthermore, we identify 11 unitary pseudogenes that are polymorphic - that is, they have both nonfunctional and functional alleles currently segregating in the human population. Comparing them with their orthologs in other primates, we find that two of them are in fact pseudogenes in non-human primates, suggesting that they represent cases of a gene being resurrected in the human lineage. CONCLUSIONS: This analysis of unitary pseudogenes provides insights into the evolutionary constraints faced by different organisms and the timescales of functional gene loss in humans.

PEMer: A Computational Framework with Simulation-Based Error Models for Inferring Genomic Structural Variants from Massive Paired-End Sequencing Data.
Korbel JO, Abyzov A, Mu XJ, Carriero N, Cayting P, Zhang Z, Snyder M, Gerstein MB (2009) Genome Biol 10(2):R23.  JOURNAL   PUBMED   REPRINT   WEBSITE     ANALYTICAL METHOD · SOFTWARE / PIPELINE / DATABASE
ABSTRACT: Personal-genomics endeavors, such as the 1000 Genomes project, are generating maps of genomic structural variants by analyzing ends of massively sequenced genome fragments. To process these we developed Paired-End Mapper (PEMer; http://sv.gersteinlab.org/pemer). This comprises an analysis pipeline, compatible with several next-generation sequencing platforms; simulation-based error models, yielding confidence-values for each structural variant; and a back-end database. The simulations demonstrated high structural variant reconstruction efficiency for PEMer's coverage-adjusted multi-cutoff scoring-strategy and showed its relative insensitivity to base-calling errors.

PeakSeq Enables Systematic Scoring of ChIP-seq Experiments Relative to Controls.
Rozowsky J, Euskirchen G, Auerbach RK, Zhang ZD, Gibson T, Bjornson R, Carriero N, Snyder M, Gerstein MB (2009) Nat Biotechnol 27(1):66-75.  JOURNAL   PUBMED   REPRINT   WEBSITE     ANALYTICAL METHOD · SOFTWARE / PIPELINE / DATABASE
ABSTRACT: Chromatin immunoprecipitation (ChIP) followed by tag sequencing (ChIP-seq) using high-throughput next-generation instrumentation is fast, replacing chromatin immunoprecipitation followed by genome tiling array analysis (ChIP-chip) as the preferred approach for mapping of sites of transcription-factor binding and chromatin modification. Using two deeply sequenced data sets for human RNA polymerase II and STAT1, each with matching input-DNA controls, we describe a general scoring approach to address unique challenges in ChIP-seq data analysis. Our approach is based on the observation that sites of potential binding are strongly correlated with signal peaks in the control, likely revealing features of open chromatin. We develop a two-pass strategy called PeakSeq to compensate for this. A two-pass strategy compensates for signal caused by open chromatin, as revealed by inclusion of the controls. The first pass identifies putative binding sites and compensates for genomic variation in the 'mappability' of sequences. The second pass filters out sites not significantly enriched compared to the normalized control, computing precise enrichments and significances. Our scoring procedure enables us to optimize experimental design by estimating the depth of sequencing required for a desired level of coverage and demonstrating that more than two replicates provides only a marginal gain in information.

Tilescope: Online Analysis Pipeline for High-Density Tiling Microarray Data.
Zhang ZD, Rozowsky J, Lam HY, Du J, Snyder M, Gerstein M (2007) Genome Biol 8(5):R81.  JOURNAL   PUBMED   REPRINT   POSTER   WEBSITE     SOFTWARE / PIPELINE / DATABASE
ABSTRACT: We developed Tilescope, a fully integrated data processing pipeline for analyzing high-density tiling-array data http://tilescope.gersteinlab.org. In a completely automated fashion, Tilescope will normalize signals between channels and across arrays, combine replicate experiments, score each array element, and identify genomic features. The program is designed with a modular, three-tiered architecture, facilitating parallelism, and a graphic user-friendly interface, presenting results in an organized web page, downloadable for further analysis.

NCIR: A Database of Non-Canonical Interactions in Known RNA Structures.
Nagaswamy U, Larios-Sanz M, Hury J, Collins S, Zhang Z, Zhao Q, Fox GE (2002) Nucleic Acids Res 30(1):395-7.  JOURNAL   PUBMED   REPRINT   WEBSITE     SOFTWARE / PIPELINE / DATABASE · STRUCTURAL RNAS
ABSTRACT: The secondary and tertiary structure of an RNA molecule typically includes a number of non-canonical base-base interactions. The known occurrences of these interactions are tabulated in the NCIR database, which can be accessed from http://prion.bchs.uh.edu/bp_type/. The number of examples is now over 1400, which is an increase of >700% since the database was first published. This dramatic increase reflects the addition of data from the recently published crystal structures of the 50S (2.4 A) and 30S (3.0 A) ribosomal subunits. In addition, non-canonical interactions observed in published crystal and NMR structures of tRNAs, group I introns, ribozymes, RNA aptamers and synthetic oligonucleotides are included. Properties associated with these interactions, such as sequence context, sugar pucker conformation, glycosidic angle conformation, melting temperature, chemical shift and free energy, are also reported when available. Out of the 29 anticipated pairs with at least two hydrogen bonds, 28 have been observed to date. In addition, several novel examples, not generally predicted, have also been encountered, bringing the total of such pairs to 36. Added to this list are a variety of single, bifurcated, triple and quadruple interactions. The most common non-canonical pairs are the sheared GA, GA imino, AU reverse Hoogsteen, and the GU and AC wobble pairs. The most frequent triple interaction connects N3 of an A with the amino of a G that is also involved in a standard Watson-Crick pair.

Database of Non-Canonical Base Pairs Found in Known RNA Structures.
Nagaswamy U, Voss N, Zhang Z, Fox GE, New Collective Author (2000) Nucleic Acids Res 28(1):375-6.  JOURNAL   PUBMED   REPRINT   WEBSITE     SOFTWARE / PIPELINE / DATABASE · STRUCTURAL RNAS
ABSTRACT: Atomic resolution RNA structures are being published at an increasing rate. It is common to find a modest number of non-canonical base pairs in these structures in addition to the usual Watson-Crick pairs. This database summarizes the occurrence of these rare base pairs in accordance with standard nomenclature. The database, http://prion.bchs.uh.edu/, contains information such as sequence context, sugar pucker conformation, anti / syn base conformations, chemical shift, p K (a)values, melting temperature and free energy. Of the 29 anticipated pairs with two or more hydrogen bonds, 20 have been encountered to date. In addition, four unexpected pairs with two hydrogen bonds have been reported bringing the total to 24. Single hydrogen bond versions of five of the expected geometries have been encountered among the single hydrogen bond interactions. In addition, 18 different types of base triplets have been encountered, each of which involves three to six hydrogen bonds. The vast majority of the rare base pairs are antiparallel with the bases in the anti configuration relative to the ribose. The most common are the GU wobble, the Sheared GA pair, the Reverse Hoogsteen pair and the GA imino pair.

Structural RNAs (6)

Rapid in Vivo Exploration of a 5S rRNA Neutral Network.
Zhang ZD, Nayar M, Ammons D, Rampersad J, Fox GE (2009) J Microbiol Methods 76(2):181-7.  JOURNAL   PUBMED   REPRINT     STRUCTURAL RNAS
ABSTRACT: A partial knockout compensation method to screen 5S ribosomal RNA sequence variants in vivo is described. The system utilizes an Escherichia coli strain in which five of eight genomic 5S rRNA genes were deleted in conjunction with a plasmid which is compensatory when carrying a functionally active 5S rRNA. The partial knockout strain is transformed with a population of potentially compensatory plasmids each carrying a randomly generated 5S rRNA gene variant. a The ability to compensate the slow growth rate of the knockout strain is used in conjunction with sequencing to rapidly identify variant 5S rRNAs that are functional as well as those that likely are not. The assay is validated by showing that the growth rate of 15 variants separately expressed in the partial knockout strain can be accurately correlated with in vivo assessments of the potential validity of the same variants. A region of 5S rRNA was mutagenized with this approach and nine novel variants were recovered and characterized. Unlike a complete knockout system, the method allows recovery of both deleterious and functional variants.. The method can be used to study variants of any 5S rRNA in the E. coli context including those of E. coli.

Microbial Identification by Mass Cataloging.
Zhang Z, Jackson GW, Fox GE, Willson RC (2006) BMC Bioinformatics 7:117.  JOURNAL   PUBMED   REPRINT     STRUCTURAL RNAS
ABSTRACT: BACKGROUND: The public availability of over 180,000 bacterial 16S ribosomal RNA (rRNA) sequences has facilitated microbial identification and classification using hybridization and other molecular approaches. In their usual format, such assays are based on the presence of unique subsequences in the target RNA and require a prior knowledge of what organisms are likely to be in a sample. They are thus limited in generality when analyzing an unknown sample.Herein, we demonstrate the utility of catalogs of masses to characterize the bacterial 16S rRNA(s) in any sample. Sample nucleic acids are digested with a nuclease of known specificity and the products characterized using mass spectrometry. The resulting catalogs of masses can subsequently be compared to the masses known to occur in previously-sequenced 16S rRNAs allowing organism identification. Alternatively, if the organism is not in the existing database, it will still be possible to determine its genetic affinity relative to the known organisms. RESULTS: Ribonuclease T1 and ribonuclease A digestion patterns were calculated for 1,921 complete 16S rRNAs. Oligoribonucleotides generated by RNase T1 of length 9 and longer produce sufficient diversity of masses to be informative. In addition, individual fragments or combinations thereof can be used to recognize the presence of specific organisms in a complex sample. In this regard, 140 strains out of 1,921 organisms (7.3%) could be identified by the presence of a unique RNase T1-generated oligoribonucleotide mass. Combinations of just two and three oligoribonucleotide masses allowed 54% and 72% of the specific strains to be identified, respectively. An initial algorithm for recovering likely organisms present in complex samples is also described. CONCLUSION: The use of catalogs of compositions (masses) of characteristic oligoribonucleotides for microbial identification appears extremely promising. RNase T1 is more useful than ribonuclease A in generating characteristic masses, though RNase A produces oligomers which are more readily distinguished due to the large mass difference between A and G. Identification of multiple species in mixtures is also feasible. Practical applicability of the method depends on high performance mass spectrometric determination, and/or use of methods that increase the one dalton (Da) mass difference between uracil and cytosine.

Common 5S rRNA Variants Are Likely to Be Accepted in Many Sequence Contexts.
Zhang Z, D'Souza LM, Lee YH, Fox GE, New Collective Author (2003) J Mol Evol 56(1):69-76.  JOURNAL   PUBMED   REPRINT     STRUCTURAL RNAS
ABSTRACT: Over evolutionary time RNA sequences which are successfully fixed in a population are selected from among those that satisfy the structural and chemical requirements imposed by the function of the RNA. These sequences together comprise the structure space of the RNA. In principle, a comprehensive understanding of RNA structure and function would make it possible to enumerate which specific RNA sequences belong to a particular structure space and which do not. We are using bacterial 5S rRNA as a model system to attempt to identify principles that can be used to predict which sequences do or do not belong to the 5S rRNA structure space. One promising idea is the very intuitive notion that frequently seen sequence changes in an aligned data set of naturally occurring 5S rRNAs would be widely accepted in many other 5S rRNA sequence contexts. To test this hypothesis, we first developed well-defined operational definitions for a Vibrio region of the 5S rRNA structure space and what is meant by a highly variable position. Fourteen sequence variants (10 point changes and 4 base-pair changes) were identified in this way, which, by the hypothesis, would be expected to incorporate successfully in any of the known sequences in the Vibrio region. All 14 of these changes were constructed and separately introduced into the Vibrio proteolyticus 5S rRNA sequence where they are not normally found. Each variant was evaluated for its ability to function as a valid 5S rRNA in an E. coli cellular context. It was found that 93% (13/14) of the variants tested are likely valid 5S rRNAs in this context. In addition, seven variants were constructed that, although present in the Vibrio region, did not meet the stringent criteria for a highly variable position. In this case, 86% (6/7) are likely valid. As a control we also examined seven variants that are seldom or never seen in the Vibrio region of 5S rRNA sequence space. In this case only two of seven were found to be potentially valid. The results demonstrate that changes that occur multiple times in a local region of RNA sequence space in fact usually will be accepted in any sequence context in that same local region.

Identification of Characteristic Oligonucleotides in the Bacterial 16S Ribosomal RNA Sequence Dataset.
Zhang Z, Willson RC, Fox GE, New Collective Author (2002) Bioinformatics 18(2):244-50.  JOURNAL   PUBMED   REPRINT     STRUCTURAL RNAS
ABSTRACT: MOTIVATION: The phylogenetic structure of the bacterial world has been intensively studied by comparing sequences of 16S ribosomal RNA (16S rRNA). This database of sequences is now widely used to design probes for the detection of specific bacteria or groups of bacteria one at a time. The success of such methods reflects the fact that there are local sequence segments that are highly characteristic of particular organisms or groups of organisms. It is not clear, however, the extent to which such signature sequences exist in the 16S rRNA dataset. A better understanding of the numbers and distribution of highly informative oligonucleotide sequences may facilitate the design of hybridization arrays that can characterize the phylogenetic position of an unknown organism or serve as the basis for the development of novel approaches for use in bacterial identification. RESULTS: A computer-based algorithm that characterizes the extent to which any individual oligonucleotide sequence in 16S rRNA is characteristic of any particular bacterial grouping was developed. A measure of signature quality, Q(s), was formulated and subsequently calculated for every individual oligonucleotide sequence in the size range of 5-11 nucleotides and for 15mers with reference to each cluster and subcluster in a 929 organism representative phylogenetic tree. Subsequently, the perfect signature sequences were compared to the full set of 7322 sequences to see how common false positives were. The work completed here establishes beyond any doubt that highly characteristic oligonucleotides exist in the bacterial 16S rRNA sequence dataset in large numbers. Over 16,000 15mers were identified that might be useful as signatures. Signature oligonucleotides are available for over 80% of the nodes in the representative tree.

NCIR: A Database of Non-Canonical Interactions in Known RNA Structures.
Nagaswamy U, Larios-Sanz M, Hury J, Collins S, Zhang Z, Zhao Q, Fox GE (2002) Nucleic Acids Res 30(1):395-7.  JOURNAL   PUBMED   REPRINT   WEBSITE     SOFTWARE / PIPELINE / DATABASE · STRUCTURAL RNAS
ABSTRACT: The secondary and tertiary structure of an RNA molecule typically includes a number of non-canonical base-base interactions. The known occurrences of these interactions are tabulated in the NCIR database, which can be accessed from http://prion.bchs.uh.edu/bp_type/. The number of examples is now over 1400, which is an increase of >700% since the database was first published. This dramatic increase reflects the addition of data from the recently published crystal structures of the 50S (2.4 A) and 30S (3.0 A) ribosomal subunits. In addition, non-canonical interactions observed in published crystal and NMR structures of tRNAs, group I introns, ribozymes, RNA aptamers and synthetic oligonucleotides are included. Properties associated with these interactions, such as sequence context, sugar pucker conformation, glycosidic angle conformation, melting temperature, chemical shift and free energy, are also reported when available. Out of the 29 anticipated pairs with at least two hydrogen bonds, 28 have been observed to date. In addition, several novel examples, not generally predicted, have also been encountered, bringing the total of such pairs to 36. Added to this list are a variety of single, bifurcated, triple and quadruple interactions. The most common non-canonical pairs are the sheared GA, GA imino, AU reverse Hoogsteen, and the GU and AC wobble pairs. The most frequent triple interaction connects N3 of an A with the amino of a G that is also involved in a standard Watson-Crick pair.

Database of Non-Canonical Base Pairs Found in Known RNA Structures.
Nagaswamy U, Voss N, Zhang Z, Fox GE, New Collective Author (2000) Nucleic Acids Res 28(1):375-6.  JOURNAL   PUBMED   REPRINT   WEBSITE     SOFTWARE / PIPELINE / DATABASE · STRUCTURAL RNAS
ABSTRACT: Atomic resolution RNA structures are being published at an increasing rate. It is common to find a modest number of non-canonical base pairs in these structures in addition to the usual Watson-Crick pairs. This database summarizes the occurrence of these rare base pairs in accordance with standard nomenclature. The database, http://prion.bchs.uh.edu/, contains information such as sequence context, sugar pucker conformation, anti / syn base conformations, chemical shift, p K (a)values, melting temperature and free energy. Of the 29 anticipated pairs with two or more hydrogen bonds, 20 have been encountered to date. In addition, four unexpected pairs with two hydrogen bonds have been reported bringing the total to 24. Single hydrogen bond versions of five of the expected geometries have been encountered among the single hydrogen bond interactions. In addition, 18 different types of base triplets have been encountered, each of which involves three to six hydrogen bonds. The vast majority of the rare base pairs are antiparallel with the bases in the anti configuration relative to the ribose. The most common are the GU wobble, the Sheared GA pair, the Reverse Hoogsteen pair and the GA imino pair.

Systems biology (3)

Network Analysis of Mitonuclear GWAS Reveals Functional Networks and Tissue Expression Profiles of Disease-Associated Genes.
Johnson SC, Gonzalez B, Zhang Q, Milholland B, Zhang Z, Suh Y (2017) Hum Genet 136(1):55-65.  JOURNAL   PUBMED   REPRINT     DISEASE STUDY · SYSTEMS BIOLOGY
ABSTRACT: While mitochondria have been linked to many human diseases through genetic association and functional studies, the precise role of mitochondria in specific pathologies, such as cardiovascular, neurodegenerative, and metabolic diseases, is often unclear. Here, we take advantage of the catalog of human genome-wide associations, whole-genome tissue expression and expression quantitative trait loci datasets, and annotated mitochondrial proteome databases to examine the role of common genetic variation in mitonuclear genes in human disease. Through pathway-based analysis we identified distinct functional pathways and tissue expression profiles associated with each of the major human diseases. Among our most striking findings, we observe that mitonuclear genes associated with cancer are broadly expressed among human tissues and largely represent one functional process, intrinsic apoptosis, while mitonuclear genes associated with other diseases, such as neurodegenerative and metabolic diseases, show tissue-specific expression profiles and are associated with unique functional pathways. These results provide new insight into human diseases using unbiased genome-wide approaches.

Systems-Level Analysis of Human Aging Genes Shed New Light on Mechanisms of Aging.
Zhang Q, Nogales-Cadenas R, Lin JR, Zhang W, Cai Y, Vijg J, Zhang ZD (2016) Hum Mol Genet 25(14):2934-2947.  JOURNAL   PUBMED   REPRINT     AGING · SYSTEMS BIOLOGY
ABSTRACT: Although studies over the last decades have firmly connected a number of genes and molecular pathways to aging, the aging process as a whole still remains poorly understood. To gain novel insights into the mechanisms underlying aging, instead of considering aging genes individually, we studied their characteristics at the systems level in the context of biological networks. We calculated a comprehensive set of network characteristics for human aging-related genes from the GenAge database. By comparing them with other functional groups of genes, we identified a robust group of aging-specific network characteristics. To find the structural basis and the molecular mechanisms underlying this aging-related network specificity, we also analyzed protein domain interactions and gene expression patterns across different tissues. Our study revealed that aging genes not only tend to be network hubs, playing important roles in communication among different functional modules or pathways, but also are more likely to physically interact and be co-expressed with essential genes. The high expression of aging genes across a large number of tissue types also points to a high level of connectivity among aging genes. Unexpectedly, contrary to the depletion of interactions among hub genes in biological networks, we observed close interactions among aging hubs, which renders the aging subnetworks vulnerable to random attacks and thus may contribute to the aging process. Comparison across species reveals the evolution process of the aging subnetwork. As the organisms become more complex, the complexity of its aging mechanisms increases and their aging hub genes are more functionally connected.

SubNet: A Java Application for Subnetwork Extraction.
Lemetre C, Zhang Q, Zhang ZD (2013) Bioinformatics 29(19):2509-11.  JOURNAL   PUBMED   REPRINT   WEBSITE     ANALYTICAL METHOD · SOFTWARE / PIPELINE / DATABASE · SYSTEMS BIOLOGY
ABSTRACT: SUMMARY: The extraction of targeted subnetworks is a powerful way to identify functional modules and pathways within complex networks. Here, we present SubNet, a Java-based stand-alone program for extracting subnetworks, given a basal network and a set of selected nodes. Designed with a graphical user-friendly interface, SubNet combines four different extraction methods, which offer the possibility to interrogate a biological network according to the question investigated. Of note, we developed a method based on the highly successful Google PageRank algorithm to extract the subnetwork using the node centrality metric, to which possible node weights of the selected genes can be incorporated. AVAILABILITY: http://www.zdzlab.org/1/subnet.html