Machine learning to the rescue : enabling novel proteomics workflows with data-driven bioinformatics methods
(2022)
- Author
- Ralf Gabriels (UGent)
- Promoter
- Lennart Martens (UGent) and Sven Degroeve (UGent)
- Organization
- Abstract
- Proteins are the molecular work horse of the cell and carry out many functional and structural tasks. In depth knowledge of the complement of all proteins in a cell or tissue, called the proteome, can provide valuable insights into cellular biology in health and disease. To study the proteome in high throughput, liquid-chromatography – tandem mass spectrometry (LC-MS/MS) is most-often the platform of choice. The result of an LC-MS/MS experiment is a large set of peptide spectra that require specific bioinformatics software to be identified. However, the confident identification of peptide spectra is not always straightforward, especially when novel challenging proteomics workflows are used. Examples of such workflows are data independent acquisition (DIA), proteogenomics, metaproteomics, biopeptidomics, and immunopeptidomics. Fortunately, machine learning (ML) can provide accurate predictions of peptide behavior in LC-MS/MS, allowing more LC-MS/MS data to be used in the identification process, resulting in a higher sensitivity. In this PhD research, the use of ML to enable novel proteomics workflows is investigated in-depth. First, the peptide spectrum predictor MS²PIP is significantly improved and extended to more use-cases. Second, a novel paradigm for the proteome-wide identification of DIA data is proposed and developed. Third, a perspective of the current state of peptide LC-MS/MS behavior predictors is given. Fourth, the MS²PIP spectrum predictor is integrated in a fully data-driven post-processing pipeline, which is subsequently applied on the various challenging proteomics workflows mentioned above. Fifth, preliminary results are shown on a novel modification-aware spectrum predictor. Each of the detailed applications of spectrum prediction for improved identification performance resulted in a more sensitive scoring function leading to more confident peptide identifications. In conclusion, ML proved to be a valuable tool for the identification of peptide mass spectra in challenging proteomics workflows. In the future, where proteomics experiments will become increasingly demanding, ML is expected to take up a central role in proteomics data analysis workflows.
- Eiwitten zijn de moleculaire werkpaarden van de cel en voeren verscheidene functionele en structurele taken uit. Diepgaande kennis van het proteoom, het totaal aan eiwitten in een cel of weefsel, kan waardevolle inzichten brengen in mechanismen van gezondheid en ziekte. Om het proteoom in high-throughput te kunnen analyseren is liquid chromatography – tandem mass spectrometry (LC-MS/MS) vaak het platform naar keuze. De resultaten van een LC-MS/MS experiment bestaat uit een grote hoeveelheid peptide spectra die geïdentificeerd moeten worden met specifieke bioinformatica software. Het gevoelig identificeren van peptide spectra is helaas niet altijd even makkelijk, zeker wanneer veeleisende proteomics workflows gebruikt werden. Voorbeelden van zulke workflows zijn data-independent acquisition (DIA), proteogenomics, metaproteomics, biopeptidomics, en immunopeptidomics. Gelukkig kan machine learning (ML) het gedrag van peptiden in LC-MS/MS accuraat voorspellen, waardoor meer informatie gebruikt kan worden in het identificatieproces, wat uiteindelijk leidt tot een hogere identificatiegevoeligheid. In dit doctoraatsonderzoek wordt het gebruik van ML voor het mogelijk maken van nieuw-uitgevonden, veeleisende proteomics workflows diepgaand bestudeerd. Eerst wordt de peptide spectrum predictor MS²PIP significant verbeterd en uitgebreid. Ten tweede wordt een nieuw paradigma voor het proteoom-breed identificeren van DIA-data voorgesteld en ontwikkeld. Ten derde wordt de huidige staat van voorspellingstools voor het gedrag van peptiden in LC-MS/MS beschreven. Ten vierde wordt MS²PIP geïntegreerd in een volledig data-gedreven proteomics post-processing workflow, wat vervolgens wordt toegepast op de verscheidene veeleisende proteomics workflows die hierboven vermeld werden. Ten vijfde worden preliminaire resultaten gedeeld over een nieuw uitgevonden peptide spectrum predictor voor gemodificeerde peptiden. Elk van de beschreven toepassingen van spectrumvoorspelling voor een verbeterde identificatieperformantie resulteerde in een meer gevoelige scoringfunctie, wat op zich dan weer resulteerde in meer peptide identificaties. In conclusie, ML heeft zich bewezen als waardevolle tool in de identificatie van peptide massa spectra in veeleisende proteomics workflows. In de toekomst, waar proteomics meer en meer uitdagend zal worden, wordt ML verwacht een centrale rol op te nemen in de bioinformatica analyse van proteomics data.
- Keywords
- Proteomics, Mass spectrometry, Bioinformatics, Machine learning, Deep learning, Computational biology
Downloads
-
phd-dissertation-gabriels-ralf-2022.pdf
- full text (Published version)
- |
- open access
- |
- |
- 4.85 MB
Citation
Please use this url to cite or link to this publication: http://hdl.handle.net/1854/LU-8754400
- MLA
- Gabriels, Ralf. Machine Learning to the Rescue : Enabling Novel Proteomics Workflows with Data-Driven Bioinformatics Methods. Ghent University. Faculty of Medicine and Health Sciences, 2022, doi:10.5281/zenodo.6580035.
- APA
- Gabriels, R. (2022). Machine learning to the rescue : enabling novel proteomics workflows with data-driven bioinformatics methods (Ghent University. Faculty of Medicine and Health Sciences). https://doi.org/10.5281/zenodo.6580035
- Chicago author-date
- Gabriels, Ralf. 2022. “Machine Learning to the Rescue : Enabling Novel Proteomics Workflows with Data-Driven Bioinformatics Methods.” Ghent, Belgium: Ghent University. Faculty of Medicine and Health Sciences. https://doi.org/10.5281/zenodo.6580035.
- Chicago author-date (all authors)
- Gabriels, Ralf. 2022. “Machine Learning to the Rescue : Enabling Novel Proteomics Workflows with Data-Driven Bioinformatics Methods.” Ghent, Belgium: Ghent University. Faculty of Medicine and Health Sciences. doi:10.5281/zenodo.6580035.
- Vancouver
- 1.Gabriels R. Machine learning to the rescue : enabling novel proteomics workflows with data-driven bioinformatics methods. [Ghent, Belgium]: Ghent University. Faculty of Medicine and Health Sciences; 2022.
- IEEE
- [1]R. Gabriels, “Machine learning to the rescue : enabling novel proteomics workflows with data-driven bioinformatics methods,” Ghent University. Faculty of Medicine and Health Sciences, Ghent, Belgium, 2022.
@phdthesis{8754400,
abstract = {{Proteins are the molecular work horse of the cell and carry out many functional and structural tasks. In depth knowledge of the complement of all proteins in a cell or tissue, called the proteome, can provide valuable insights into cellular biology in health and disease. To study the proteome in high throughput, liquid-chromatography – tandem mass spectrometry (LC-MS/MS) is most-often the platform of choice. The result of an LC-MS/MS experiment is a large set of peptide spectra that require specific bioinformatics software to be identified. However, the confident identification of peptide spectra is not always straightforward, especially when novel challenging proteomics workflows are used. Examples of such workflows are data independent acquisition (DIA), proteogenomics, metaproteomics, biopeptidomics, and immunopeptidomics. Fortunately, machine learning (ML) can provide accurate predictions of peptide behavior in LC-MS/MS, allowing more LC-MS/MS data to be used in the identification process, resulting in a higher sensitivity. In this PhD research, the use of ML to enable novel proteomics workflows is investigated in-depth. First, the peptide spectrum predictor MS²PIP is significantly improved and extended to more use-cases. Second, a novel paradigm for the proteome-wide identification of DIA data is proposed and developed. Third, a perspective of the current state of peptide LC-MS/MS behavior predictors is given. Fourth, the MS²PIP spectrum predictor is integrated in a fully data-driven post-processing pipeline, which is subsequently applied on the various challenging proteomics workflows mentioned above. Fifth, preliminary results are shown on a novel modification-aware spectrum predictor. Each of the detailed applications of spectrum prediction for improved identification performance resulted in a more sensitive scoring function leading to more confident peptide identifications. In conclusion, ML proved to be a valuable tool for the identification of peptide mass spectra in challenging proteomics workflows. In the future, where proteomics experiments will become increasingly demanding, ML is expected to take up a central role in proteomics data analysis workflows.}},
author = {{Gabriels, Ralf}},
keywords = {{Proteomics,Mass spectrometry,Bioinformatics,Machine learning,Deep learning,Computational biology}},
language = {{eng}},
pages = {{128}},
publisher = {{Ghent University. Faculty of Medicine and Health Sciences}},
school = {{Ghent University}},
title = {{Machine learning to the rescue : enabling novel proteomics workflows with data-driven bioinformatics methods}},
url = {{http://doi.org/10.5281/zenodo.6580035}},
year = {{2022}},
}
- Altmetric
- View in Altmetric