Aug 20, 2026
PLOS (Public Library of Science)
Author summary Identifying genetic variants in human and non-human species is important across agriculture, biotechnology, ecology, and evolution. However, sequencing machines are not perfect, and they often produce errors that look like real mutations. A key challenge is to develop computational workflows that reliably filter out these errors and find true genetic variations. Here, we report that standard workflows are largely optimized for human genomes and introduce systematic biases when applied to non-human species. To solve this, we developed a simple and portable workflow called “pseudo-database” (pseudoDB). Instead of relying on external information, this approach uses the raw sequencing data to build its own internal benchmark, allowing it to reduce technical noise without needing any prior genomic knowledge. We find that the pseudoDB workflow outperforms existing approaches across a diverse range of species, including brown bear, swan goose, and stevia. Notably, we uncovered tens of thousands of variants in overlooked regions that control how genes are regulated. This work levels the playing field for non-human research and enables the optimal use of genomic technology for biodiversity conservation and improving global food security.