Abstract:Molecular species identification and environmental DNA (eDNA) technologies have developed rapidly in recent years. They are widely used in biodiversity surveys and environmental monitoring because they are efficient and non-invasive. However, their performance and accuracy depend strongly on the completeness and reliability of reference databases, especially DNA barcode libraries. In China, most river basins still lack systematic, standardized, and comprehensive barcode resources. This limitation restricts the application of molecular methods in ecological assessment and biodiversity conservation. The Yarlung Tsangpo River is one of the highest-altitude large transboundary rivers in the world. Its unique climate and complex hydrological and geomorphological conditions support rich and distinctive freshwater fish diversity. However, a basin-wide DNA barcode reference library for fish is still lacking. Systematic genetic data remain insufficient, which limits the use of molecular monitoring approaches in this region. In this study, we compiled DNA sequences from fish specimens collected by our research team across the Yarlung Tsangpo River basin from 1998 to 2024. We also incorporated sequences from public databases. All data were processed using standardized methods, followed by strict quality control and reliability assessment based on genetic distances. Based on these steps, we established the first comprehensive DNA barcode reference dataset covering the entire basin. The dataset contains 3,174 high-quality DNA sequences. Among them, 2,890 sequences (91.1%) were newly generated, and 284 sequences (8.9%) were obtained from public databases. Samples were collected from 82 sites, including the main stem, major tributaries, and associated lakes and wetlands. The elevation range spans from 155 to 4,600 m. The dataset includes 78 species from 49 genera, 21 families, and 8 orders. These species comprise 62 native species and 16 non-native species. The dataset covers 69% of endemic fish species in the basin and 100% of fish species in the reach upstream of the Lhagu River confluence. The average sequence length is approximately 813 bp. The main molecular markers are cytochrome c oxidase subunit I (COI) and cytochrome b (Cyt b). Their sequence length ranges are 461-1779 bp and 798-1390 bp, respectively. Most species show clear DNA barcode gaps, which support reliable species-level identification. The dataset adopts a “metadata-sequence data” separation framework. It follows the FAIR (Findable, Accessible, Interoperable, Reusable) and CARE (Collective Benefit, Authority to Control, Responsibility, Ethics) principles. It includes standardized information on taxonomy, voucher specimens, distribution, sampling time, collectors, and sequence sources. Color photographs are available for some species. The dataset is openly accessible through the Science Data Bank(ScienceDB; DOI:10.57760/sciencedb.36688). This dataset fills a key gap in DNA barcode reference resources for fish in the Yarlung Tsangpo River basin. It supports species identification, biodiversity inventory, non-native species monitoring, and eDNA metabarcoding studies. It also provides essential data support for the conservation and management of plateau river ecosystems.