Post-Search Validation and Curation of Site-Resolved N-Glycoproteomics Data

Abstract Given the significant role of glycosylation in modulating protein structure and activity, glycoproteomics is gaining increased interest from the broad scientific community. Large-scale site-resolved N-glycoproteomics relies on automated MS2-based searches, but candidate glycopeptide assignments can remain ambiguous when isomeric or isobaric glycan structures, adducts, chemical modifications, in-source fragments, or incomplete MS2 evidence support more than one plausible interpretation. Here, we systematically categorize common challenges and misassignments in glycoproteomics and present a post-search validation workflow using Skyline software to identify and correct these. The workflow matches search-engine-derived candidate assignments to LC–MS/MS evidence for correct precursor monoisotope assignment, retention time behavior, and glycosite context. The workflow is demonstrated with Byonic-derived glycopeptide candidate lists and converts automated search results into curated, verifiable, site-resolved N-glycopeptide features for downstream quantification and reporting. We applied the workflow to data from 52 human serum samples, and reviewed 3,071 candidate N-glycopeptide IDs. From these, 1,722 MS2 candidate IDs were refuted as inconsistent with chromatographic and/or precursor-level evidence. Curation added 320 glycopeptide features, comprising 152 MS1-supported composition-level assignments and 168 additional LC-resolved isomer features, yielding a final curated feature set of 1,436 N-glycopeptides across the serum N-glycoproteome. Together, these results show that reviewing the raw LC-MS/MS data associated with search-engine results improves both the accuracy and comprehensiveness of detectable N-glycopeptides, supporting more transparent and reliable reporting. The curated dataset provides a resource for future method development, benchmarking and machine–learning efforts directed at automated glycopeptide validation.

Authors

Institutions

Publication Details

Journal
JACS Au
Published
2026-09-09
DOI
https://doi.org/10.1021/jacsau.6c00875
Primary Topic
Glycosylation and Glycoproteins Research
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Post-Search Validation and Curation of Site-Resolved N-Glycoproteomics Data

Lenka Hernychová, Juan Camilo Rojas Echeverri, Noortje de Haan, Adam P. Urminsky
JACS Au
Glycosylation and Glycoproteins Research
article

Post-Search Validation and Curation of Site-Resolved N-Glycoproteomics Data

Lenka Hernychová, Juan Camilo Rojas Echeverri, Noortje de Haan, Adam P. Urminsky
article en

Abstract

Abstract Given the significant role of glycosylation in modulating protein structure and activity, glycoproteomics is gaining increased interest from the broad scientific community. Large-scale site-resolved N-glycoproteomics relies on automated MS2-based searches, but candidate glycopeptide assignments can remain ambiguous when isomeric or isobaric glycan structures, adducts, chemical modifications, in-source fragments, or incomplete MS2 evidence support more than one plausible interpretation. Here, we systematically categorize common challenges and misassignments in glycoproteomics and present a post-search validation workflow using Skyline software to identify and correct these. The workflow matches search-engine-derived candidate assignments to LC–MS/MS evidence for correct precursor monoisotope assignment, retention time behavior, and glycosite context. The workflow is demonstrated with Byonic-derived glycopeptide candidate lists and converts automated search results into curated, verifiable, site-resolved N-glycopeptide features for downstream quantification and reporting. We applied the workflow to data from 52 human serum samples, and reviewed 3,071 candidate N-glycopeptide IDs. From these, 1,722 MS2 candidate IDs were refuted as inconsistent with chromatographic and/or precursor-level evidence. Curation added 320 glycopeptide features, comprising 152 MS1-supported composition-level assignments and 168 additional LC-resolved isomer features, yielding a final curated feature set of 1,436 N-glycopeptides across the serum N-glycoproteome. Together, these results show that reviewing the raw LC-MS/MS data associated with search-engine results improves both the accuracy and comprehensiveness of detectable N-glycopeptides, supporting more transparent and reliable reporting. The curated dataset provides a resource for future method development, benchmarking and machine–learning efforts directed at automated glycopeptide validation.

JACS Au
Masaryk University (CZ), Leiden University Medical Center (NL), Masaryk Memorial Cancer Institute (CZ)
Openalex Percentile: Top 17%
Glycosylation and Glycoproteins Research
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.