De novo and reference-based genome assembly

Genome analysis

Whole genome sequencing provides information on the entire genetic material of an organism. There are two approaches for assembling high-throughput sequencing reads into longer contiguous genomic sequences:

De novo genome assembly

This approach is used for non-model genomes where no reference genome is available. Sequenced reads are compared to each other, and overlapping reads are used to build longer contiguous sequences. Contigs are oriented and ordered using long reads.

Reference-based genome assembly

This approach involves mapping each read to a reference genome sequence to identify genetic variations like single nucleotide polymorphisms (SNPs), indels, insertions, copy number variants, genome-wide association studies (GWAS), and building haplotypes from genome assemblies.

Eurofins Genomics offers a variety of sequencing platforms such as Illumina MiSeq, NextSeq, NovaSeq, ONT and PacBio, with different read lengths and libraries sizes (paired-end) for whole genome sequencing of humans, animals, plants, and microorganisms like bacteria, viruses, and fungi. Our long read sequencing can handle any genome size, from bacterial genomes to large and complex eukaryotic genomes. Long paired-end reads determine the orientation and relative position of the contigs generated during data assembly.

Genome assembly services 

  • Bacterial/fungal de novo genome assembly
  • PacBio/nanopore bacterial/fungal de novo assembly
  • De novo genome assembly up to 1 Gb
  • Large genome de novo assembly (>1 Gb)
  • Fungal hybrid de novo assembly and analysis
  • Large genome hybrid de novo assembly and analysis
  • Reference-guided genome analysis
  • PacBio/nanopore reference-guided analysis

Bioinformatics workflow and deliverables for genome assemble requests

Quality check of raw reads

  • Quality filtration and adapter trimming
  • Removal of primer sequences, poly(A) tails, and reads from ribosomal DNA templates
  • High-quality data used for downstream analysis

De novo assembly

  • Multiple kmer assembly runs to optimise the assembly
  • PE data assembled using various parameters like kmer length, coverage cut-off, insert length, and standard deviation
  • Best assembly selected based on scaffold N50 and max scaffold length
  • Final assembly evaluated on metrics like scaffold N50, assembly coverage, GC content, completeness, and accuracy

Reference-based analysis

  • Downloading reference genome and gene information from public databases
  • Aligning high-quality reads against the reference genome with optimised parameters

Gene prediction

  • Using statistical models to find gene features like start and stop codons, CDS of the genes
  • Predicting coding regions in the given sample

Annotation

  • Annotating predicted coding regions against databases like NCBI Nr, swissprot, KEGG, and COG using BlastX
  • Mapping coding regions to reference canonical pathways in KEGG
  • Assigning GO terms for functional categorisation

Deliverables

De novo genome assembly

  • Quality filtering of reads
  • De novo assembly to generate scaffolds and contigs
  • Assembly statistics
  • In silico validation using RNA-seq data (for large, complex plant genomes)
  • GC content analysis
  • Repeat identification
  • Gene prediction
  • Gene annotation
  • Gene ontology (GO) analysis
  • Simple sequence repeat (SSR) discovery
  • Phylogenetic analysis
  • KEGG pathway analysis
  • Comparative genomics with closely related genomes
  • Circos plot generation
  • COG orthologous group analysis
  • Antimicrobial resistance (AMR) and virulence factor analysis
  • Comprehensive report with publication-standard methodology, graphs, and table

Reference-based genome analysis

  • Quality filtration of reads
  • Mapping of high-quality reads to the reference genome
  • Alignment summary (including, reads mapped, uniquely mapped reads, reads unmapped, genome coverage)
  • Consensus sequence in fasta format
  • Gene prediction using gtf/gff
  • SNP/Indels identification
  • SNP/Indels annotation
  • Core gene analysis
  • Comparative genomics
  • Phylogenetic analysis
  • Comprehensive report with publication-standard methodology, graphs, and tables

Find a site near you!

Clear Filters
Reset Filters
Getting the map ready...