GB/T 47347-2026 Information technology—Biometrics—Annotation format for genotyping data based on high-throughput sequencing English, Anglais, Englisch, Inglés, えいご
This is a draft translation for reference among interesting stakeholders. The finalized translation (passing through draft translation, self-check, revision and verification) will be delivered upon being ordered.
ICS
CCS
National Standard of the People's Republic of China
GB/T 47347-2026
Information technology - Biometrics - Annotation format for genotyping data based on high-throughput sequencing
信息技术 生物特征识别 高通量测序基因分型数据注释格式
Issue date: 2026-03-31 Implementation date: 2026-10-01
Issued by the General Administration of Quality Supervision, Inspection and Quarantine of the People's Republic of China
the Standardization Administration of the People's Republic of China
Contents
Foreword
1 Scope
2 Normative References
3 Terms and Definitions
4 Abbreviations
5 Requirements for Annotation Format
Annex A (Informative) Example of Genotyping Data Annotation
Bibliography
Information technology — Biometrics — Annotation format for highthroughput sequencing genotyping data
1 Scope
This document specifies the requirements for the annotation format of human genotyping data generated by highthroughput sequencing.
This document applies to the storage, exchange and comparison of human genotyping data generated by highthroughput sequencing.
2 Normative References
This document has no normative references.
3 Terms and Definitions
The following terms and definitions apply to this document.
3.1 highthroughput sequencing
A technology that enables the parallel sequencing of a large number of nucleic acid molecules in a single run.
NOTE: Typically, a single sequencing run can produce no less than 100 Mb of sequencing data.
[Source: GB/T 33767.14-2023, 3.1, modified]
3.2 reference sequence of human genome
A digital nucleotide sequence database assembled from the genome sequences of multiple human individuals.
NOTE: The human reference genome sequence is typically used as a benchmark genomic reference sequence. The most commonly used version of the human reference genome sequence is currently GRCh38, stored in FASTA format.
3.3 short tandem repeat; STR
Tandem repeats on a chromosome with a repeat unit of 2 bp to 6 bp, exhibiting a high degree of individual variation.
[Source: GB/T 33767.14-2023, 3.14]
3.4 single nucleotide polymorphism; SNP
A DNA sequence polymorphism resulting from a single nucleotide change.
[Source: GB/T 33767.14-2023, 3.15]
3.5 insertion deletion; InDel
A class of polymorphic genetic markers formed by the insertion or deletion of deoxyribonucleic acid (DNA) fragments of varying lengths in the genome.
3.6 mitochondrial DNA; MtDNA
The genetic material found in mitochondria.
3.7 microhaplotype; MH
A combination of multiple SNPs that are coinherited on the same chromosome, typically with a length not exceeding 300 bp.
3.8 metainformation
The basic information for the annotation of highthroughput sequencing genotyping data.
NOTE: Includes information such as the annotation file format, file generation date, highthroughput sequencing platform, genotyping software and reference genome.
3.9 keyinformation
Identifiers defined in metainformation lines or data fields and their associated sets of attributes, used for standardised parsing of variant data and genotype information.
NOTE: Includes the key name (ID), data type (Type), number of values (Number) and description (Description).
3.10 quality score
Q
A quality assessment value for the genotype inference of a genetic marker in a sample.
NOTE: Q = –10 lg(p), where p represents the probability that the genotype of the genetic marker is incorrect.
Standard
GB/T 47347-2026 Information technology—Biometrics—Annotation format for genotyping data based on high-throughput sequencing (English)
Standard No.
GB/T 47347-2026
Status
to be valid
Language
English
File Format
PDF
Word Count
9000 words
Translation Price(USD)
270.0
Implemented on
2026-10-1
Delivery
via email in 1~3 business day
Detail of GB/T 47347-2026
Standard No.
GB/T 47347-2026
English Name
Information technology—Biometrics—Annotation format for genotyping data based on high-throughput sequencing
GB/T 47347-2026 Information technology—Biometrics—Annotation format for genotyping data based on high-throughput sequencing English, Anglais, Englisch, Inglés, えいご
This is a draft translation for reference among interesting stakeholders. The finalized translation (passing through draft translation, self-check, revision and verification) will be delivered upon being ordered.
ICS
CCS
National Standard of the People's Republic of China
GB/T 47347-2026
Information technology - Biometrics - Annotation format for genotyping data based on high-throughput sequencing
信息技术 生物特征识别 高通量测序基因分型数据注释格式
Issue date: 2026-03-31 Implementation date: 2026-10-01
Issued by the General Administration of Quality Supervision, Inspection and Quarantine of the People's Republic of China
the Standardization Administration of the People's Republic of China
Contents
Foreword
1 Scope
2 Normative References
3 Terms and Definitions
4 Abbreviations
5 Requirements for Annotation Format
Annex A (Informative) Example of Genotyping Data Annotation
Bibliography
Information technology — Biometrics — Annotation format for highthroughput sequencing genotyping data
1 Scope
This document specifies the requirements for the annotation format of human genotyping data generated by highthroughput sequencing.
This document applies to the storage, exchange and comparison of human genotyping data generated by highthroughput sequencing.
2 Normative References
This document has no normative references.
3 Terms and Definitions
The following terms and definitions apply to this document.
3.1 highthroughput sequencing
A technology that enables the parallel sequencing of a large number of nucleic acid molecules in a single run.
NOTE: Typically, a single sequencing run can produce no less than 100 Mb of sequencing data.
[Source: GB/T 33767.14-2023, 3.1, modified]
3.2 reference sequence of human genome
A digital nucleotide sequence database assembled from the genome sequences of multiple human individuals.
NOTE: The human reference genome sequence is typically used as a benchmark genomic reference sequence. The most commonly used version of the human reference genome sequence is currently GRCh38, stored in FASTA format.
3.3 short tandem repeat; STR
Tandem repeats on a chromosome with a repeat unit of 2 bp to 6 bp, exhibiting a high degree of individual variation.
[Source: GB/T 33767.14-2023, 3.14]
3.4 single nucleotide polymorphism; SNP
A DNA sequence polymorphism resulting from a single nucleotide change.
[Source: GB/T 33767.14-2023, 3.15]
3.5 insertion deletion; InDel
A class of polymorphic genetic markers formed by the insertion or deletion of deoxyribonucleic acid (DNA) fragments of varying lengths in the genome.
3.6 mitochondrial DNA; MtDNA
The genetic material found in mitochondria.
3.7 microhaplotype; MH
A combination of multiple SNPs that are coinherited on the same chromosome, typically with a length not exceeding 300 bp.
3.8 metainformation
The basic information for the annotation of highthroughput sequencing genotyping data.
NOTE: Includes information such as the annotation file format, file generation date, highthroughput sequencing platform, genotyping software and reference genome.
3.9 keyinformation
Identifiers defined in metainformation lines or data fields and their associated sets of attributes, used for standardised parsing of variant data and genotype information.
NOTE: Includes the key name (ID), data type (Type), number of values (Number) and description (Description).
3.10 quality score
Q
A quality assessment value for the genotype inference of a genetic marker in a sample.
NOTE: Q = –10 lg(p), where p represents the probability that the genotype of the genetic marker is incorrect.