Bacteriophages are viruses that infect bacteria and shape microbial communities. Growing antibiotic resistance and advances in biotechnology have renewed interest in understanding how phages infect, manipulate, and lyse their hosts.

Despite the rapid growth of phage genome data, current annotation approaches still rely largely on sequence homology, and many predicted proteins remain hypothetical or poorly characterized. This phage unknome limits mechanistic insight into host takeover and the discovery of proteins with biotechnological potential.

Image: Graham Beards / Wikimedia Commons , CC BY-SA 3.0 , modified and colorized.

Graphical abstract of PhageGAP: The model features processing of protein sequences with a protein language model followed by function classification using a CNN. Processed sequences are placed in a t-SNE landscape of reference phage, bacterial, and viral proteins. Nearest neighbors are identified and sequence comparison, genomic context, and structural information can be explored, together with predicted function, to support hypothesis generation.

PhageGAP predicts curated functional classes for phage proteins using protein language model embeddings.

Proteins are placed within an interactive embedding landscape of reference phage, bacterial, and viral proteins and identifies their nearest neighbors. Sequence comparison, genomic context, and structural information can then be explored together to support interpretation and hypothesis generation.


Provide Protein Sequence(s) in FASTA Format via Text and/or File Input
and/or
  • Input must be in FASTA format. Line length is unrestricted; ambiguous amino acids are converted to X. The first header field after > is used as the protein ID; remaining header text is treated as description. Text and file input are combined.
  • Optional Genomic Context can be added from the toolbar by providing a GFF file. FASTA and GFF protein IDs must match.
  • Function Prediction uploads data for classification. Results are added to the Embedding Landscape as User Data (○). Additional submissions are appended to existing results.
  • Results table/session data can be downloaded from the toolbar and restored using Reload Session.

Embedding Landscape
Selected Protein:
Min. Class Prediction Probability:
Function Prediction

Click a User Data point in the embedding landscape to view nearest neighbor information.

Nearest-Neighbor Structure
Sequence Alignment
 Match   Mismatch   InDel   Gap
Genomic Context
PhageGAP logo

PhageGAP provides:

  • protein function prediction using protein language model embeddings and a trained classifier.
  • interactive exploration of predictions, embedding similarity, sequence alignments, structures, and genomic context.
  • confidence-aware analysis of individual proteins and complete phage proteomes.

How to use PhageGAP

Quick Start

On the Function Prediction page, follow these steps to analyze your protein sequences:

  1. Paste protein sequences or upload a FASTA file.
  2. Click Function Prediction.
  3. Explore the submitted proteins in the Embedding Landscape.
  4. Select a protein to inspect its prediction and nearest-neighbor sequence, structure, and similarity.
  5. Optionally add genomic context from a GFF3 file.
  6. Download the results or save the complete session for later use.

An example session with CDS of the phiKZ phage genome (NC_004629.1) is available via Download Example Session function.

1) Provide Protein Sequences

Enter protein sequences directly or upload a FASTA file. If both are provided, all sequences are combined into one submission.

  • Sequences must be provided in FASTA format.
  • The protein ID is derived from the FASTA header; remaining text is used as its description.
  • Ambiguous amino acids are accepted and converted to X.
  • Protein IDs should be unique within a submission.
>prot_001 putative capsid protein
MNNR...
>prot_002 tail fiber protein
MDSK...

2) Run Function Prediction

Click Function Prediction to analyze the submitted proteins. Predictions are added to the current session and shown as User Data in the Embedding Landscape.

  • Cookie consent must be accepted before classification is enabled.
  • Submitted protein sequences are processed for analysis but are not persistently stored by PhageGAP.
  • Classification requests are rate-limited. If a request is rejected, wait before submitting again.

3) Explore the Results

The Embedding Landscape places submitted proteins within the embedding landscape of the PhageGAP reference dataset.

  • Reference proteins are colored by their functional class; submitted proteins are shown as User Data.
  • Use the legend to show or hide individual functional classes.
  • Hover over proteins to inspect their annotation, prediction, and nearest-neighbor information.
  • Use the toolbox to zoom, reset the view, or hide reference-identical submitted proteins.
  • Set a minimum prediction probability to hide submitted proteins below the selected threshold.

When no protein is selected, the Function Prediction panel summarizes the highest predicted class probabilities across the submitted proteins, separated into reference-identical and non-identical proteins.

Select a User Data protein by clicking its point or using the searchable protein-selection menu. The detail views are synchronized with the current selection.

  • The retained nearest reference proteins are connected to the selected protein in the Embedding Landscape.
  • Hover over a connection to inspect the corresponding nearest neighbor and its embedding distance.
  • Click a connection to inspect a different retained nearest neighbor; the closest neighbor is selected by default.
  • Function Prediction shows the three highest-ranked functional predictions and their probabilities.
  • Nearest-Neighbor Structure shows the available structure and metadata for the selected nearest neighbor.
  • Nearest-Neighbor Sequence Alignment shows its pairwise alignment with the submitted protein and reports sequence identity.
  • Alignment colors distinguish matches, mismatches, insertions, and deletions and are linked to the structure view.
  • Use the alignment zoom control and interactive structure controls for detailed inspection.

A structure may not be available for every reference protein.

The sizes of the main visualization panels can be adjusted by dragging the vertical divider.

4) Add Genomic Context

Use Add Genomic Context to upload a GFF3 file (.gff or .gff3).

  • Only CDS features are displayed.
  • The GFF protein identifier must match the corresponding FASTA protein ID.
  • Matched CDS features are colored according to their predicted functional class.
  • Selecting a protein centers the genomic view on its local neighborhood.
  • Selecting a CDS in the genomic view selects the corresponding protein in the Embedding Landscape.
  • The Genomic Context panel enables you to zoom in and out using scrolling and panning operations, allowing you to inspect individual loci or the complete annotated region.

5) Save, Restore, and Export

  • Download Results exports prediction results as TSV.
  • Download Session saves the complete current analysis as compressed JSON (.json.gz).
  • Reload Session restores a previously saved session for continued exploration or additional predictions.
  • Screenshot exports the current analysis panels as PNG.