How to run
Kaptive's main function is in silico serotyping via kaptive type.
Basic usage is outlined below and full instructions, including all command line options, are detailed on the CLI Usage pages.
Inputs¶
Kaptive performs in silico serotyping on bacterial whole genome assemblies. Your input genome assemblies should be in FASTA format (they can be gzip-compressed), with one assembly per file.
Go serotyping!¶
Once you have installed Kaptive and your chosen database, you are ready to serotype.
To run Kaptive on your assemblies you can run:
| Bash | |
|---|---|
This assumes your input files have the .fasta file extenion and are avialable in the current directory.
Kaptive will run on each assembly and print the output to a tab-delimted file called results.tsv.
kpsc_k is the database keyword, and points to the Klebsiella pneumoniae Species Complex K database. For a full list of supported key words see here. You can also run Kaptive on a custom database by following these instructions.
Outputs¶
Tabular¶
The main output of the assembly typing mode is a tab-delimited table of the results. See here for tips on interpreting these results. For full explanation of the column content see here.
The default is to print this table to stdout. You can use UNIX
redirection operators (> or >>) or the -o/--out flag to write to
a file.
If the summary table already exists and is not empty, Kaptive will append to it (not overwrite it) and suppress the header line. This allows you to run Kaptive in succession on sets of assemblies, all outputting to the same table file.
To disable the tabular output, simply redirect the output to
/dev/null.
Locus sequences¶
The -l/--loci flag produces a fasta file of the region(s) of the
assembly which correspond to the best locus match. This may be a single
piece (in cases of a good assembly and a strong match) or it may be in
multiple pieces (in cases of poor assembly and/or a novel locus).
You can specify either a directory, which will write one file per
assembly named as {assembly}_kaptive_results.fna, or a single file
("-" for stdout), which will write all the sequences to that file.
For example:
| Bash | |
|---|---|
This results in default behaviour which will produce one file per assembly in the current directory. However, to specify a directory:
| Bash | |
|---|---|
| Bash | |
|---|---|
Note
This is the same as the --fna flag in kaptive convert.
Locus plots¶
Kaptive can generate interative plots showing the locus pieces and genes present in your input assemblies.
| Bash | |
|---|---|
If no directory is specified, Kaptive will generate the plots in the current directory.
Note
Make sure you have installed the correct dependencies to support plotting.
PHA4GE genotyping spec¶
The Public Health Alliance for Genomic Epidemiology has developed the PHA4GE Microbial Genotyping Data Specification, which represents a standardised format for communicating genotyping methods and results. Kaptive can optionally output results in this format:
| Bash | |
|---|---|
JSON¶
The -j/--json flag produces a JSON file of the results which allows
Kaptive to reconstruct the TypingResult objects after a run which can
be used with kaptive-convert. Unlike
previous version (2 and below), this is a JSON lines file (or "-" for
stdout), where each line is a JSON object representing the results for
a single assembly. If the file already exists, Kaptive will append to it
(not overwrite it).
The default is to write this file to: kaptive_results.json, however
the path can be specified after the flag, for example:
| Text Only | |
|---|---|
1 | |
Warning
It is possible to write all text formats (TSV, JSON and FASTA) to the same file (including stdout), however this is not recommended for downstream analysis.