Genome annotation of Anopheles gambiae using mass spectometry-derived data

Dário E. Kalume; Suraj Peri; Raghunath Reddy; Jun Zhong; Mobolaji Okulate; Nirbhay Kumar; Akhilesh Pandey

doi:10.1186/1471-2164-6-128

Genome annotation of Anopheles gambiae using mass spectometry-derived data

Dário E. Kalume, Suraj Peri, Raghunath Reddy, Jun Zhong, Mobolaji Okulate, Nirbhay Kumar, Akhilesh Pandey

Laboratory Medicine and Pathology

Research output: Contribution to journal › Article › peer-review

54 Scopus citations

Abstract

Background: A large number of animal and plant genomes have been completely sequenced over the last decade and are now publicly available. Although genomes can be rapidly sequenced, identifying protein-coding genes still remains a problematic task. Availability of protein sequence data allows direct confirmation of protein-coding genes. Mass spectrometry has recently emerged as a powerful tool for proteomic studies. Protein identification using mass spectrometry is usually carried out by searching against databases of known proteins or transcripts. This approach generally does not allow identification of proteins that have not yet been predicted or whose transcripts have not been identified. Results: We searched 3,967 mass spectra from 16 LC-MS/MS runs of Anopheles gambiae salivary gland homogenates against the Anopheles gambiae genome database. This allowed us to validate 23 known transcripts and 50 novel transcripts. In addition, a novel gene was identified on the basis of peptides that matched a genomic region where no gene was known and no transcript had been predicted. The amino termini of proteins encoded by two predicted transcripts were confirmed based on N-terminally acetylated peptides sequenced by tandem mass spectrometry. Finally, six sequence polymorphisms could be annotated based on experimentally obtained peptide sequences. Conclusion: The peptide sequences from this study were mapped onto the genomic sequence using the distributed annotation system available at Ensembl and can be visualized in the context of all other existing annotations. The strategy described in this paper can be used to correct and confirm genome annotations and permit discovery of novel proteins in a high-throughput manner by mass spectrometry.

Original language	English (US)
Article number	128
Journal	BMC genomics
Volume	6
DOIs	https://doi.org/10.1186/1471-2164-6-128
State	Published - Sep 19 2005

ASJC Scopus subject areas

Biotechnology
Genetics

Access to Document

10.1186/1471-2164-6-128

Cite this

@article{83c37744ced94d93a61279d2b8601287,

title = "Genome annotation of Anopheles gambiae using mass spectometry-derived data",

abstract = "Background: A large number of animal and plant genomes have been completely sequenced over the last decade and are now publicly available. Although genomes can be rapidly sequenced, identifying protein-coding genes still remains a problematic task. Availability of protein sequence data allows direct confirmation of protein-coding genes. Mass spectrometry has recently emerged as a powerful tool for proteomic studies. Protein identification using mass spectrometry is usually carried out by searching against databases of known proteins or transcripts. This approach generally does not allow identification of proteins that have not yet been predicted or whose transcripts have not been identified. Results: We searched 3,967 mass spectra from 16 LC-MS/MS runs of Anopheles gambiae salivary gland homogenates against the Anopheles gambiae genome database. This allowed us to validate 23 known transcripts and 50 novel transcripts. In addition, a novel gene was identified on the basis of peptides that matched a genomic region where no gene was known and no transcript had been predicted. The amino termini of proteins encoded by two predicted transcripts were confirmed based on N-terminally acetylated peptides sequenced by tandem mass spectrometry. Finally, six sequence polymorphisms could be annotated based on experimentally obtained peptide sequences. Conclusion: The peptide sequences from this study were mapped onto the genomic sequence using the distributed annotation system available at Ensembl and can be visualized in the context of all other existing annotations. The strategy described in this paper can be used to correct and confirm genome annotations and permit discovery of novel proteins in a high-throughput manner by mass spectrometry.",

author = "Kalume, {D{\'a}rio E.} and Suraj Peri and Raghunath Reddy and Jun Zhong and Mobolaji Okulate and Nirbhay Kumar and Akhilesh Pandey",

year = "2005",

month = sep,

day = "19",

doi = "10.1186/1471-2164-6-128",

language = "English (US)",

volume = "6",

journal = "BMC genomics",

issn = "1471-2164",

publisher = "BioMed Central",

}

TY - JOUR

T1 - Genome annotation of Anopheles gambiae using mass spectometry-derived data

AU - Kalume, Dário E.

AU - Peri, Suraj

AU - Reddy, Raghunath

AU - Zhong, Jun

AU - Okulate, Mobolaji

AU - Kumar, Nirbhay

AU - Pandey, Akhilesh

PY - 2005/9/19

Y1 - 2005/9/19

N2 - Background: A large number of animal and plant genomes have been completely sequenced over the last decade and are now publicly available. Although genomes can be rapidly sequenced, identifying protein-coding genes still remains a problematic task. Availability of protein sequence data allows direct confirmation of protein-coding genes. Mass spectrometry has recently emerged as a powerful tool for proteomic studies. Protein identification using mass spectrometry is usually carried out by searching against databases of known proteins or transcripts. This approach generally does not allow identification of proteins that have not yet been predicted or whose transcripts have not been identified. Results: We searched 3,967 mass spectra from 16 LC-MS/MS runs of Anopheles gambiae salivary gland homogenates against the Anopheles gambiae genome database. This allowed us to validate 23 known transcripts and 50 novel transcripts. In addition, a novel gene was identified on the basis of peptides that matched a genomic region where no gene was known and no transcript had been predicted. The amino termini of proteins encoded by two predicted transcripts were confirmed based on N-terminally acetylated peptides sequenced by tandem mass spectrometry. Finally, six sequence polymorphisms could be annotated based on experimentally obtained peptide sequences. Conclusion: The peptide sequences from this study were mapped onto the genomic sequence using the distributed annotation system available at Ensembl and can be visualized in the context of all other existing annotations. The strategy described in this paper can be used to correct and confirm genome annotations and permit discovery of novel proteins in a high-throughput manner by mass spectrometry.

AB - Background: A large number of animal and plant genomes have been completely sequenced over the last decade and are now publicly available. Although genomes can be rapidly sequenced, identifying protein-coding genes still remains a problematic task. Availability of protein sequence data allows direct confirmation of protein-coding genes. Mass spectrometry has recently emerged as a powerful tool for proteomic studies. Protein identification using mass spectrometry is usually carried out by searching against databases of known proteins or transcripts. This approach generally does not allow identification of proteins that have not yet been predicted or whose transcripts have not been identified. Results: We searched 3,967 mass spectra from 16 LC-MS/MS runs of Anopheles gambiae salivary gland homogenates against the Anopheles gambiae genome database. This allowed us to validate 23 known transcripts and 50 novel transcripts. In addition, a novel gene was identified on the basis of peptides that matched a genomic region where no gene was known and no transcript had been predicted. The amino termini of proteins encoded by two predicted transcripts were confirmed based on N-terminally acetylated peptides sequenced by tandem mass spectrometry. Finally, six sequence polymorphisms could be annotated based on experimentally obtained peptide sequences. Conclusion: The peptide sequences from this study were mapped onto the genomic sequence using the distributed annotation system available at Ensembl and can be visualized in the context of all other existing annotations. The strategy described in this paper can be used to correct and confirm genome annotations and permit discovery of novel proteins in a high-throughput manner by mass spectrometry.

UR - http://www.scopus.com/inward/record.url?scp=25444475024&partnerID=8YFLogxK

UR - http://www.scopus.com/inward/citedby.url?scp=25444475024&partnerID=8YFLogxK

U2 - 10.1186/1471-2164-6-128

DO - 10.1186/1471-2164-6-128

M3 - Article

C2 - 16171517

AN - SCOPUS:25444475024

SN - 1471-2164

VL - 6

JO - BMC genomics

JF - BMC genomics

M1 - 128

ER -

Genome annotation of Anopheles gambiae using mass spectometry-derived data

Abstract

ASJC Scopus subject areas

Access to Document

Other files and links

Fingerprint

Cite this