MOC - Bioinformatics in Microbiology

Computational analysis of microbial sequence and omics data — from raw reads to clinical and epidemiological interpretation.

Parent: Home · Map: Encyclopedia Map
Companion: MOC - AI in Microbiology (learning models on these data)
Practical: Bioinformatics Toolkit for Microbiology · Genomics Command-Line Cheatsheet

Overview

Bioinformatics turns raw reads into actionable microbial knowledge: species ID, resistance and virulence genes, plasmids, community composition, and outbreak relatedness. Clinical microbiology increasingly depends on these pipelines downstream of Whole-Genome Sequencing and amplicon PCR.

flowchart TB
  Raw[Raw reads FASTQ] --> QC[[Read QC and Preprocessing]]
  QC --> Assembly[[Genome Assembly]]
  QC --> Map[Mapping]
  Map --> Var[[Variant Calling in Bacteria]]
  Assembly --> Annot[[Genome Annotation]]
  Annot --> AMR[[AMR Gene Databases]]
  Annot --> Vir[[Virulence Factor Databases]]
  Annot --> Pan[[Pangenome Analysis]]
  Assembly --> Typ[[MLST and cgMLST]]
  Assembly --> Plas[[Plasmid and Mobile Element Analysis]]
  Var --> Tree[[Phylogenetic Tree Building]]
  Typ --> Tree
  Tree --> Dyn[[Phylodynamics]]
  Raw --> Meta[[Metagenomics]]
  AMR --> Report[Clinical / epi report]
  Tree --> Report

1. Data and Foundations

2. Genome Reconstruction and Interpretation

3. Comparative and Population Genomics

4. Typing, Phylogeny, Epidemiology

5. Beyond the Genome (multi-omics)

6. Culture-Independent Analysis

7. Clinical AMR Genomics

8. Practice, Data Stewardship, Reproducibility

9. Bridge to AI

Core Principles

  • Reference and database versions are part of the result (Reproducible Bioinformatics Workflows)
  • Genotype ≠ phenotype — correlate with Antimicrobial Susceptibility Testing when therapy depends on it
  • Contamination, mixed cultures, and low coverage invalidate everything downstream (Contaminant and Mixed-Culture Detection)
  • Metadata quality limits epidemiological value more often than sequence quality
  • Every clinical result must be traceable from report back to raw reads
  • Recombination and population structure must be modeled before outbreak or GWAS claims

Tool Reference Card

CategoryExamplesQuestion answered
QCFastQC, MultiQC, fastpAre reads usable?
Species screenKraken2, Mash, GTDB-TkWhat organism(s)?
AssemblySPAdes, Unicycler, Flye, ShovillWhat is the genome?
Assembly QCQUAST, CheckM, BUSCOIs it complete/clean?
AnnotationProkka, Bakta, PGAPWhich genes?
VariantsBWA/minimap2, bcftools, SnippyWhich SNPs?
AMRAMRFinderPlus, ResFinder, CARD-RGIWhich resistance determinants?
Typingmlst, chewBBACA, KleborateWhich lineage/cluster?
PlasmidsPlasmidFinder, MOB-suite, geNomadMobile context?
PangenomeRoary, Panaroo, PPanGGOLiNCore vs accessory?
PhylogenyMAFFT, IQ-TREE, Gubbins, BEASTHow related, and when?
MetagenomicsMetaPhlAn, metaSPAdes, MetaBAT2What is in the community?
AmpliconQIIME 2, DADA2Taxa from 16S?
VisualizationiTOL, Microreact, BandageHow do I show it?
OrchestrationNextflow/nf-core, Snakemake, DockerHow do I rerun it exactly?

Important Papers

Important Book Chapters

Research Questions

  1. How should labs report “gene present, MIC susceptible”?
  2. What minimum metadata makes AMR genomic surveillance interoperable?
  3. When do plasmids demand long reads for clinical conclusions?
  4. Can pangenome-aware references replace single-reference SNP calling in routine surveillance?
  5. What is the acceptable failure mode when a pipeline meets a novel species?

Review Article Opportunities

  • Practical WGS pipeline for clinical microbiology laboratories
  • Database discordance → reporting standards
  • Metagenomic diagnostics: sensitivity, contamination, regulation
  • From MAGs to clinical relevance: what is missing

Learning Aids

Build Status

ClusterStatus
Data foundations, assembly, annotation, variantsdone
Assembly QC, long-read/hybrid, contamination gates✅ 2026-08-02
Comparative / pangenome / plasmids / typingdone
ANI/GTDB, bacterial GWAS, recombination-aware trees, PopPUNK
Clinical WGS pipelines + prophage annotation
Phylogenetics and phylodynamicsdone
Multi-omics (RNA, protein, structure)done
Metagenomics, MAGs, microbiome statisticsdone
Reproducibility, databases, FAIR, cheatsheetdone
Worked examples with real datasetsbacklog