← All posts

DNA Methylation Data Analysis with EpiNexus

Updated 3 October 2026 to match the current EpiNexus app.

DNA methylation — the covalent addition of a methyl group to the 5-carbon of cytosine, predominantly at CpG dinucleotides — is one of the most extensively studied epigenetic modifications in biology. It plays a central role in gene regulation, genomic imprinting, X-chromosome inactivation, transposon silencing, and genome stability. Aberrant methylation patterns are hallmarks of cancer, neurodevelopmental disorders, and aging.

Modern technologies now allow genome-wide methylation profiling at single-base resolution. But for most wet lab researchers, turning those raw sequencing files into biological insight remains a serious computational challenge. In this post, we walk through what methylation analysis involves, why it's hard, and how EpiNexus makes it accessible.

How methylation is measured

There are several experimental approaches to profiling DNA methylation, each with trade-offs in resolution, coverage, and cost:

Whole-genome bisulfite sequencing (WGBS) treats genomic DNA with sodium bisulfite, which converts unmethylated cytosines to uracils while leaving methylated cytosines intact. After sequencing and alignment, you get the methylation status of virtually every CpG in the genome — roughly 28 million sites in humans. WGBS is the gold standard for unbiased, comprehensive methylome profiling, but requires deep sequencing (30x or more per sample).

Reduced representation bisulfite sequencing (RRBS) uses restriction enzyme digestion (typically MspI) to enrich for CpG-dense regions before bisulfite treatment. This captures around 1–3 million CpGs concentrated at promoters and CpG islands, at roughly a tenth of the sequencing cost of WGBS — making it practical for large cohort studies.

Enzymatic methyl-seq (EM-seq) is a newer alternative that replaces the harsh bisulfite conversion with enzymatic reactions (TET2 oxidation followed by APOBEC deamination). The result is the same base-resolution methylation readout, but with less DNA degradation, better coverage uniformity, and improved performance on limited or degraded input DNA such as FFPE samples or cell-free DNA.

Methylation arrays (Illumina 450K, EPIC) measure methylation at hundreds of thousands of predefined CpG sites using probe intensities, reporting beta values (proportion methylated) or M-values (logit-transformed). While not sequencing-based, they remain widely used in clinical and epidemiological studies.

All sequencing-based approaches produce FASTQ files that look like standard data — but require specialized alignment and quantification to handle the bisulfite-converted bases.

The traditional analysis pipeline

A typical methylation analysis chains together several command-line tools, each with its own installation headaches, parameter choices, and output formats:

1. Quality control and trimming. FastQC and Trim Galore assess read quality and remove adapters. Methylation data requires additional attention to end-of-read bias — bisulfite conversion introduces systematic artifacts at read termini that can skew methylation estimates if not properly trimmed.

2. Bisulfite-aware alignment. Tools like Bismark or bwa-meth align reads to an in-silico converted reference genome. The aligner must consider both C-to-T and G-to-A converted versions of the genome simultaneously. This makes alignment significantly slower and more memory-intensive than standard DNA-seq mapping. Post-alignment QC includes checking mapping rates, coverage uniformity, and M-bias plots to detect position-specific artifacts.

3. Methylation extraction. Bismark's methylation extractor or MethylDackel calculates the methylation percentage at each cytosine from the ratio of unconverted (C) to converted (T) reads. Strand-specific quantification is important for proper downstream analysis.

4. Filtering and summarization. CpGs with insufficient coverage (typically fewer than 10 reads) are removed, along with sites overlapping known SNPs. Methylation values can be aggregated over genomic windows, CpG islands, promoters, or other features of interest.

5. Exploratory analysis. PCA, hierarchical clustering, and global methylation distributions help detect batch effects, sample swaps, and broad biological patterns before diving into differential analysis.

6. Differential methylation. Statistical tools like DSS, methylKit, or dmrseq identify differentially methylated CpGs (DMCs) and differentially methylated regions (DMRs) between conditions. Region-level methods often provide higher statistical power and more biologically interpretable results by smoothing or aggregating signals from neighboring CpGs.

7. Annotation and interpretation. DMRs are mapped to genomic features — promoters, gene bodies, enhancers, CpG islands, shores, and shelves — and tested for enrichment in pathways, gene ontology terms, and regulatory element databases.

For a researcher who just received sequencing data from the core facility, assembling and validating this pipeline can take weeks.

What makes methylation data challenging

Beyond the usual bioinformatics hurdles, methylation data has specific complexities that make analysis particularly demanding:

Reduced sequence complexity. Bisulfite conversion turns most unmethylated C's into T's, effectively reducing the four-letter DNA alphabet. This makes read alignment slower, more ambiguous, and prone to mapping artifacts — especially in repetitive regions.

Coverage is uneven and expensive. WGBS requires deep sequencing for reliable single-CpG resolution, and even at 30x average coverage, some regions will have gaps. RRBS is cheaper but only covers a subset of CpGs. Researchers must make informed decisions about which sites to include and how to handle low-coverage data.

Subtle biological signals. Unlike histone modifications where a peak is present or absent, methylation changes are often quantitative shifts — from 80% to 60% methylated, for example. Detecting these requires statistical models that properly account for the binomial nature of count data and biological variability across replicates. Methods must handle overdispersion and apply careful multiple-testing correction across millions of sites.

Genomic context matters enormously. A CpG losing methylation in a promoter has very different biological implications than one in a gene body, intergenic region, or transposable element. CpG islands, shores (2 kb flanking regions), and shelves (the next 2 kb beyond) each carry distinct functional significance. Analysis that ignores context misses much of the biology.

Cell-type composition confounds bulk data. Bulk tissue samples mix cells with different methylation profiles. An apparent methylation change between conditions might actually reflect a shift in cell-type proportions rather than true epigenetic reprogramming. Deconvolution or reference-based approaches may be needed to disentangle these effects.

How EpiNexus handles methylation

EpiNexus wraps the methylation pipeline into the same workflow used for ChIP-seq and other epigenomic assays:

Upload your data

Drag and drop your FASTQ.gz files from WGBS or RRBS experiments, or import published samples from GEO, SRA or ENA by accession. If you have already aligned your reads with Bismark, you can start from its coverage (.cov) files instead. Select your reference genome, put each sample in a group — treatment vs. control, tumor vs. normal, or any comparison — and choose the reference group.

Before the run starts, EpiNexus checks whether your libraries are RRBS or WGBS from the reads themselves: RRBS reads begin at an MspI site (CGG or TGG), WGBS reads begin anywhere. The Review step shows the result, and you can choose the library type yourself instead.

Automated pipeline

Once you click run, EpiNexus executes the full pipeline:

  • Trimming — adapter and quality trimming with Trim Galore; for RRBS, its --rrbs mode removes the filled-in bases at MspI-site ends
  • Bisulfite-aware alignment — Bismark (Bowtie2) against C-to-T and G-to-A converted copies of the genome
  • Deduplication — PCR duplicates are removed for WGBS; RRBS is never deduplicated, as Bismark advises, because every read of a fragment starts at the same MspI site
  • Methylation calling — per-CpG methylation with Bismark's methylation extractor, with the two strands of each CpG merged; M-bias plots show any read-end bias, and for paired-end WGBS the first 2 bp of read 2 are ignored, as the Bismark documentation recommends
  • Differential methylation — differentially methylated regions (DMRs) with dmrseq, which models biological variation across replicates and controls the false discovery rate at the region level; it tests CpGs covered by at least 4 reads in every sample. Without replicates, EpiNexus falls back to a per-CpG Fisher test and labels those results exploratory.
  • Annotation — every DMR is linked to its nearest gene

Explore your results

The Methylation page shows:

An overview — how many CpGs were tested, mean methylation in each group, and the global difference between them. When no region reaches significance, the page says so and shows what was tested.

A volcano plot — methylation difference on the x-axis, statistical significance on the y-axis, hypermethylated and hypomethylated regions in different colours.

Chromosome distribution and a DMR table — where the DMRs fall, with their size, number of CpGs, methylation difference, FDR and nearest gene; the table downloads as a tab-separated file.

Bisulfite QC — for each sample: reads, mapping rate, duplicates removed, mean CpG coverage, methylation in CpG, CHG and CHH context, an estimate of bisulfite conversion (graded against ENCODE's 98% standard), and M-bias plots. A MultiQC report collects the tools' own reports.

Putting methylation in context

Methylation changes are most meaningful when viewed alongside other layers of epigenomic regulation:

  • Methylation and histone marks — a promoter that loses methylation and gains H3K4me3 suggests transcriptional activation.
  • Methylation and chromatin accessibility — demethylation at a region that becomes accessible in ATAC-seq points to active regulatory remodeling.
  • Pathways and motifs — the genes near DMRs, and the transcription-factor motifs within them, suggest which processes and regulators are affected.

You can analyse ChIP-seq and ATAC-seq data from the same system in EpiNexus and compare the genes near your DMRs with those results. A joint methylation–chromatin analysis, and GO or motif enrichment of DMRs, are not part of EpiNexus yet.

Choosing the right approach for your experiment

The right methylation method depends on your question and practical constraints:

WGBS is best for comprehensive, unbiased discovery — when you need to survey the entire methylome, including enhancers, gene bodies, and intergenic regions. Budget 400–800 million reads per sample for mammalian genomes.

RRBS is ideal for focused studies on promoters and CpG islands with many samples. At 30–50 million reads per sample, it's an order of magnitude cheaper than WGBS and well-suited for clinical cohorts.

EM-seq gives WGBS-level coverage with better data quality, especially valuable for limited or degraded input DNA — FFPE tissue, liquid biopsies, rare cell populations.

EpiNexus processes WGBS and RRBS with the same upload-and-click workflow, detecting which one you have from the reads and adapting trimming and deduplication to it. EM-seq isn't supported as its own library type yet.

Getting started

DNA methylation analysis doesn't need to be a months-long computational project. If you have FASTQ files from a WGBS or RRBS experiment, EpiNexus can take you from raw reads to differentially methylated regions and their nearest genes, with bisulfite QC, all from your browser.

Methylation analysis is available in the EpiNexus pilot at app.epinexus.io. Accounts are by invitation — if you'd like one, get in touch; we'd love to hear about your project.


Further reading: For a comprehensive review of methylation analysis methods, see Bock (2012) in Nature Reviews Genetics. For hands-on pipeline tutorials, the Galaxy Training Network and the Computational Genomics with R book provide excellent walkthroughs.