Home
Quickstart¶
If you're looking to get up and running with in silico serotyping of genome assemblies as fast as possible, you're in the right place. We'll skip the heavy algorithmic details here and just focus on getting you your results.
1. Install Kaptive¶
2. Download a Database¶
Kaptive uses beautifully curated databases containing reference surface antigen loci for different species. To see what's available for download, just run:
To pull down a specific database, use the install command with the database name. For example, if you're working with the Acinetobacter baumannii K-locus, you'd want ab_k:
3. Serotype Your Genomes!¶
Once Kaptive and your database are installed, you're ready to roll. Your input genome assemblies should be in FASTA format (and don't worry, gzip-compressed files work perfectly too).
Let's run Kaptive and save the output to a file called results.tsv:
4. Understanding the Output¶
Kaptive spits out a tab-separated values (TSV) report, which you can easily open up in Excel, Numbers, or any text editor to browse through.
Here are the most critical columns to keep an eye on in your results.tsv file:
- Assembly: The name of your input genome file.
- Best match locus: The best-matching serotype or locus found in the database (e.g.,
KL1). - Confidence: How confident Kaptive is in this call - this is either "Typeable" or "Untypeable"
For a deeper dive into all the other columns and more advanced outputs, check out the full Outputs documentation. Happy serotyping!
Tutorial¶
Step-by-step video and documented tutorials are available, covering:
- Kaptive's features and their scientific rationale
- How to run Kaptive
- Examples, illustrating how to run and interpret results
- Further investigations (e.g. exploring novel loci, IS insertions)
Kaptive theory with Kelly!
Kaptive usage with Tom!
Note
The tutorials are based on Kaptive 2.0, but the principles are similar for Kaptive 3.0.
Dependencies¶
As of v3.3.0, Kaptive does not rely on any external binaries and is 100% self-contained.
We do rely on the following dependencies, as defined in the pyproject.toml:
- python: The core language of the library.
- numpy: For vectorised numerical operations.
- numba: For speeding up heavy algebra.
- rammappy: For fast, sensitive, nucleotide alignment.
- gb-io: For parsing databases.
Optional Dependencies¶
Kaptive comes with optional dependency groups to keep installation light for
most users. These can be easily installed with (uv|pip) install kaptive[group1,group2,...].
Citation¶
If you use Kaptive in your work, please cite:
@article{mbs:/content/journal/mgen/10.1099/mgen.0.001428,
author = "Stanton, Thomas David and Hetland, Marit A.K. and Löhr, Iren H. and Holt, Kathryn E. and Wyres, Kelly L.",
title = "Fast and accurate in silico antigen typing with Kaptive 3",
journal= "Microbial Genomics",
year = "2025",
volume = "11",
number = "6",
pages = "",
doi = "https://doi.org/10.1099/mgen.0.001428",
url = "https://www.microbiologyresearch.org/content/journal/mgen/10.1099/mgen.0.001428",
publisher = "Microbiology Society",
issn = "2057-5858",
type = "Journal Article",
keywords = "serotyping",
keywords = "Klebsiella",
keywords = "tools",
keywords = "capsule",
keywords = "sero-epidemiology",
keywords = "antigen",
eid = "001428",
abstract = "Surface polysaccharides are common antigens in priority pathogens and therefore attractive targets for novel control strategies such as vaccines, monoclonal antibody and phage therapies. Distinct serotypes correspond to diverse polysaccharide structures that are encoded by distinct biosynthesis gene clusters; e.g. the Klebsiella pneumoniae species complex (KpSC) K- and O-loci encode the synthesis machinery for the capsule (K) and outer-lipopolysaccharides (O), respectively. We previously presented Kaptive and Kaptive 2, programmes to identify K- and O-loci directly from KpSC genome assemblies (later adapted for Acinetobacter baumannii), enabling sero-epidemiological analyses to guide vaccine and phage therapy development. However, for some KpSC genome collections, Kaptive (v≤2) was unable to type a high proportion of K-loci. Here, we identify the cause of this issue as assembly fragmentation and present a new version of Kaptive (v3) to circumvent this problem, reduce processing times and simplify output interpretation. We compared the performance of Kaptive v2 and Kaptive v3 for typing genome assemblies generated from subsampled Illumina read sets (decrements of 10× depth), for which a corresponding high-quality completed genome was also available to determine the ‘true’ loci (n=549 KpSC, n=198 A. baumannii). Both versions of Kaptive showed high rates of agreement to the matched true locus amongst ‘typeable’ locus calls (≥96% for ≥20× read depth), but Kaptive v3 was more sensitive, particularly for low-depth assemblies (at <40× depth, v3 ranged 0.85–1 vs v2 0.09–0.94) and/or typing KpSC K-loci (e.g. 0.97 vs 0.82 for non-subsampled assemblies). Overall, Kaptive v3 was also associated with a higher rate of optimal outcomes; i.e. loci matching those in the reference database were correctly typed, and genuine novel loci were reported as untypeable (73–98% for v3 vs 7–77% for v2 for KpSC K-loci). Kaptive v3 was >1 order of magnitude faster than Kaptive v2, making it easy to analyse thousands of assemblies on a desktop computer, facilitating broadly accessible in silico serotyping that is both accurate and sensitive. The Kaptive v3 source code is freely available on GitHub (https://github.com/klebgenomics/Kaptive), and has been implemented in Kaptive Web (https://kaptive-web.erc.monash.edu/).",
}
