WAVEBench is a benchmark for evaluating evolutionary generalization in genome language models through viral-versus-cellular classification of wastewater metagenomic reads. It combines taxonomic cross-validation, fixed evaluation reads, and composition controls to measure performance across levels of taxonomic novelty.
The datasets are currently private and available to authorized collaborators. Full raw and deduplicated corpora and original per-read Kraken2 outputs are not distributed. Data provenance and preprocessing methods will be described in the accompanying paper (link forthcoming).