雅鲁藏布江流域鱼类DNA条形码数据集
CSTR:
作者:
作者单位:

1.西藏大学;2.中国科学院水生生物研究所;3.中国科学院大学;4.生态环境学院,青藏高原生物多样性与生态环境保护教育部重点实验室;5.拉萨城市湿地生态系统西藏自治区野外科学观测研究站

作者简介:

通讯作者:

中图分类号:

基金项目:

第二次青藏高原综合考察研究(2024QZKK0200; 2019QZKK05010102); 西藏自治区科技计划基地与人才计划项目(XZ202501JD0019)


A basin-scale fish DNA barcode dataset from the Yarlung Tsangpo River
Author:
Affiliation:

1.Xizang University;2.Institute of Hydrobiology, Chinese Academy of Sciences;3.University of Chinese Academy of Sciences;4.Key Laboratory of Biodiversity and Environment on the Qinghai-Tibetan Plateau, Ministry of Education, School of Ecology and Environment;5.Lhasa, Urban Wetland Ecosystem, Observation and Research Station of Tibet Autonomous Region;6.University of Chinese Academy of Sciences, Beijing

Fund Project:

the Second Tibetan Plateau Scientific Expedition and Research Program (2024QZKK0200; 2019QZKK05010102); Base and Talent Program Projects of Science and Technology Program of Tibet Autonomous Region(XZ202501JD0019)

  • 摘要
  • |
  • 图/表
  • |
  • 访问统计
  • |
  • 参考文献
  • |
  • 相似文献
  • |
  • 引证文献
  • |
  • 附件
  • |
  • 文章评论
    摘要:

    分子物种鉴定与环境DNA(environmental DNA, eDNA)技术近年来在生物多样性调查与环境监测中迅速发展,因其高效、非侵入式等优势而受到广泛关注,但其应用效果和结果准确性在很大程度上依赖于DNA条形码等参考数据库的完整性与可靠性。当前,我国多数流域尚缺乏系统、规范且覆盖全的 DNA 条形码数据资源,已成为制约分子监测技术在生态评估与生物多样性保护中应用的瓶颈。雅鲁藏布江作为世界上海拔最高的大型跨境河流之一,独特的气候条件和水文地貌孕育了丰富多样、极具特色的淡水鱼类区系,然而其流域尺度的鱼类DNA条形码数据库建设仍滞后,缺乏系统性数据支撑,限制了分子监测手段在该区域生物多样性调查与保护中的应用。本研究集成了研究团队1998—2024年采自雅鲁藏布江流域鱼类样品DNA序列,并整合公共数据库中的雅鲁藏布江流域鱼类序列,经统一数据处理、严格质量控制和基于遗传距离的可靠性评估,汇编形成了首个覆盖雅鲁藏布江全流域的鱼类 DNA 条形码参考数据集。数据集包含3,174条高质量 DNA序列,其中自主测序序列2,890条(占总序列数的91.1%),公共数据库序列284条(占8.9%)。数据样点覆盖雅鲁藏布江干流、主要支流及附属湖泊、湿地,共计82个样点;海拔跨度为155—4,600 m;物种涵盖8目21科49属78种,包括土著鱼类62种,外来鱼类16种;覆盖了69%的雅鲁藏布江特有鱼类物种和100%的泸公河汇口及其上游江段鱼类;序列平均长度约为813 bp,核心分子标记为 COI和Cyt b,序列长度范围分别为461-1779 bp和798-1390 bp;多数物种表现出清晰的 DNA 条形码间隙,可支持分子水平的物种鉴定。数据集采用“元数据—序列数据”分离管理模式,遵循 FAIR(可发现、可访问、可操作、可重用)和 CARE(集体利益、控制权、责任、伦理)国际数据治理原则,包含物种分类地位、标本凭证号、分布信息、采集时间、采集人及序列来源等标准化字段,并附有部分测序物种的彩色照片。数据集通过科学数据银行(ScienceDB;DOI:10.57760/sciencedb.36688)向用户开放访问。本数据集弥补了雅鲁藏布江流域鱼类分子鉴定参考数据的关键缺口,可为物种鉴定、生物多样性编目、外来物种监测及eDNA 宏条形码研究提供参考依据,并为高原河流生态系统的科学保护与系统管理提供数据支撑。

    Abstract:

    Molecular species identification and environmental DNA (eDNA) technologies have developed rapidly in recent years. They are widely used in biodiversity surveys and environmental monitoring because they are efficient and non-invasive. However, their performance and accuracy depend strongly on the completeness and reliability of reference databases, especially DNA barcode libraries. In China, most river basins still lack systematic, standardized, and comprehensive barcode resources. This limitation restricts the application of molecular methods in ecological assessment and biodiversity conservation. The Yarlung Tsangpo River is one of the highest-altitude large transboundary rivers in the world. Its unique climate and complex hydrological and geomorphological conditions support rich and distinctive freshwater fish diversity. However, a basin-wide DNA barcode reference library for fish is still lacking. Systematic genetic data remain insufficient, which limits the use of molecular monitoring approaches in this region. In this study, we compiled DNA sequences from fish specimens collected by our research team across the Yarlung Tsangpo River basin from 1998 to 2024. We also incorporated sequences from public databases. All data were processed using standardized methods, followed by strict quality control and reliability assessment based on genetic distances. Based on these steps, we established the first comprehensive DNA barcode reference dataset covering the entire basin. The dataset contains 3,174 high-quality DNA sequences. Among them, 2,890 sequences (91.1%) were newly generated, and 284 sequences (8.9%) were obtained from public databases. Samples were collected from 82 sites, including the main stem, major tributaries, and associated lakes and wetlands. The elevation range spans from 155 to 4,600 m. The dataset includes 78 species from 49 genera, 21 families, and 8 orders. These species comprise 62 native species and 16 non-native species. The dataset covers 69% of endemic fish species in the basin and 100% of fish species in the reach upstream of the Lhagu River confluence. The average sequence length is approximately 813 bp. The main molecular markers are cytochrome c oxidase subunit I (COI) and cytochrome b (Cyt b). Their sequence length ranges are 461-1779 bp and 798-1390 bp, respectively. Most species show clear DNA barcode gaps, which support reliable species-level identification. The dataset adopts a “metadata-sequence data” separation framework. It follows the FAIR (Findable, Accessible, Interoperable, Reusable) and CARE (Collective Benefit, Authority to Control, Responsibility, Ethics) principles. It includes standardized information on taxonomy, voucher specimens, distribution, sampling time, collectors, and sequence sources. Color photographs are available for some species. The dataset is openly accessible through the Science Data Bank(ScienceDB; DOI:10.57760/sciencedb.36688). This dataset fills a key gap in DNA barcode reference resources for fish in the Yarlung Tsangpo River basin. It supports species identification, biodiversity inventory, non-native species monitoring, and eDNA metabarcoding studies. It also provides essential data support for the conservation and management of plateau river ecosystems.

    参考文献
    相似文献
    引证文献
引用本文
相关视频

分享
文章指标
  • 点击次数:
  • 下载次数:
  • HTML阅读次数:
  • 引用次数:
历史
  • 收稿日期:2026-02-03
  • 最后修改日期:2026-04-17
  • 录用日期:2026-04-20
  • 在线发布日期: 2026-06-22
  • 出版日期:
文章二维码
您是第    位访问者
地址:南京市江宁区麒麟街道创展路299号    邮政编码:211135
电话:025-86882041;86882040     传真:025-57714759     Email:jlakes@niglas.ac.cn
Copyright:中国科学院南京地理与湖泊研究所《湖泊科学》 版权所有:All Rights Reserved
技术支持:北京勤云科技发展有限公司

苏公网安备 32010202010073号

     苏ICP备09024011号-2