PROTrEIN https://protrein.eu Training of computational proteomics researchers Wed, 08 May 2024 14:23:21 +0000 es hourly 1 https://wordpress.org/?v=5.5.1 https://protrein.eu/wp-content/uploads/2020/12/cropped-favicon-32x32.png PROTrEIN https://protrein.eu 32 32 Cross-Border Collaboration: Enhancing Peptide Identification with MS2Rescore and MS Amanda https://protrein.eu/blog/cross-border-collaboration-enhancing-peptide-identification-with-ms2rescore-and-ms-amanda/ Wed, 08 May 2024 14:12:42 +0000 https://protrein.eu/?p=1937 Back in June of 2022 I left Austria to go on my first secondment in Ghent, Belgium to visit Alireza and the rest of the Compomics group. Besides getting to know the people,  enjoying the summer in the beautiful city and eating numerous waffles, I also did a small project together with members of the […]

The post Cross-Border Collaboration: Enhancing Peptide Identification with MS2Rescore and MS Amanda appeared first on PROTrEIN.

]]>
Back in June of 2022 I left Austria to go on my first secondment in Ghent, Belgium to visit Alireza and the rest of the Compomics group. Besides getting to know the people,  enjoying the summer in the beautiful city and eating numerous waffles, I also did a small project together with members of the group. Specifically Arthur Declercq and Ralf Gabriels, who are the main developers of the well known rescoring platform MS2Rescore (https://github.com/compomics/ms2rescore) [1]. At that time I was working on making MS2Rescore compatible with MS Amanda [2], a well established database search engine for peptide identification, developed by my supervisor Viktoria Dorfer. 

This small project had the main purpose of me getting to know the tools developed at CompOmics and this eventually lead to a great collaboration between our two groups that resulted in this recent publication:

“MS2Rescore 3.0 is a modular, flexible, and user-friendly platform to boost peptide identifications, as showcased with MS Amanda 3.0” 

[Caption: Peptide identification result files from many different search engines, including MS Amanda, are parsed to MS2Rescore which will then perform data-driven rescoring, utilizing several different feature generators and rescoring engines. This leads to a higher number of confident identifications at the same false discovery rate (FDR) threshold or a similar number of confident identifications at a more stringent FDR threshold. MS2Rescore results can be easily inspected in the newly added HTML quality control reports.]

We’re excited to introduce the latest, enhanced versions of both MS2Rescore and MS Amanda. MS2Rescore 3.0 is highly modularized and flexible, making it easy to add new input formats through the python package psm_utils (https://github.com/compomics/psm_utils) [3], new feature generation modules and new rescoring modules. This version of MS2Rescore is available both as a command line interface, a graphical user interface and as a python API, making implementation into already existing workflows for peptide identification very easy.  Lastly, MS2Rescore 3.0 will output a HTML quality control report after the rescoring which allows users to assess the effect of data-driven rescoring on their identification workflow and the performance of individual features. With MS Amanda 3.0 come seven new columns in the output CSV file that can be used for rescoring search results as well as automatic integration with Percolator [4]. 

We demonstrate the flexibility of MS2Rescore 3.0 by connecting it with MS Amanda 3.0 and show how using MS2Rescore 3.0 in combination with MS Amanda 3.0 can increase the number of identified spectra in a challenging single-cell data set. 

It feels great knowing that what started as a small project during my secondment eventually ended up with me having my first publication. Since this was a highly collaborative process between our two research groups it is important to mention that Arthur Declercq and I are shared first authors and that Viktoria Dorfer and Ralf Gabriels are shared last authors. Thanks again to Arthur, Ralf, Viki and the rest of my co-authors! 

You can read the publication here:

https://pubs.acs.org/doi/10.1021/acs.jproteome.3c00785

Or the preprint, free of charge, here:

https://chemrxiv.org/engage/chemrxiv/article-details/65b614e2e9ebbb4db9237db9

References

  1. Declercq A, Bouwmeester R, Hirschler A, Carapito C, Degroeve S, Martens L, Gabriels R. MS2Rescore: Data-Driven Rescoring Dramatically Boosts Immunopeptide Identification Rates. Mol Cell Proteomics. 2022 Aug;21(8):100266. doi: 10.1016/j.mcpro.2022.100266. Epub 2022 Jul 6. PMID: 35803561; PMCID: PMC9411678.
  2. Dorfer V, Pichler P, Stranzl T, Stadlmann J, Taus T, Winkler S, Mechtler K. MS Amanda, a universal identification algorithm optimized for high accuracy tandem mass spectra. J Proteome Res. 2014 Aug 1;13(8):3679-84. doi: 10.1021/pr500202e. Epub 2014 Jun 26. PMID: 24909410; PMCID: PMC4119474. 
  3. Gabriels R, Declercq A, Bouwmeester R, Degroeve S, Martens L. psm_utils: A High-Level Python API for Parsing and Handling Peptide-Spectrum Matches and Proteomics Search Results. J Proteome Res. 2023 Feb 3;22(2):557-560. doi: 10.1021/acs.jproteome.2c00609. Epub 2022 Dec 12. PMID: 36508242.
  4. Käll L, Canterbury JD, Weston J, Noble WS, MacCoss MJ. Semi-supervised learning for peptide identification from shotgun proteomics datasets. Nat Methods. 2007 Nov;4(11):923-5. doi: 10.1038/nmeth1113. Epub 2007 Oct 21. PMID: 17952086. 

The post Cross-Border Collaboration: Enhancing Peptide Identification with MS2Rescore and MS Amanda appeared first on PROTrEIN.

]]>
Exploring Cellular Complexity: Unveiling Single-Cell Proteomics https://protrein.eu/blog/exploring-cellular-complexity-unveiling-single-cell-proteomics/ Fri, 08 Sep 2023 15:48:03 +0000 https://protrein.eu/?p=1910 There are enormous amounts of biological cascades in every cell – the smallest functional compartment of our body [1]. Understanding the mechanisms underlying the vast array of phenomena is not only the key element to finding any clues about fatal diseases such as Alzheimer’s and cancer but also to progress in developing treatment for such […]

The post Exploring Cellular Complexity: Unveiling Single-Cell Proteomics appeared first on PROTrEIN.

]]>
There are enormous amounts of biological cascades in every cell – the smallest functional compartment of our body [1]. Understanding the mechanisms underlying the vast array of phenomena is not only the key element to finding any clues about fatal diseases such as Alzheimer’s and cancer but also to progress in developing treatment for such diseases. To address this need, scientists across diverse biological disciplines have embraced a multi-omic analysis approach, deciphering meaningful codes enciphered by cells, such as the genome, transcriptome, and proteome [1,2]. Recently, the discovery of new developments in the omics approaches allows us to conduct these analyses at single-cell resolution [3,4].

In this month of the journal club, we would like to touch upon single-cell technologies from the mass spectrometry-based proteomics perspective. The title of the selected article is “Ultra-high sensitivity mass spectrometry quantifies single-cell proteome changes upon perturbation” published by Andreas-David Brunner et al [4]. Before going into details about the article, we would like to give brief information about single-cell technology.

One of the most well-known analogies for the single-cell method is the smoothie example [3]. Imagine that you take a sip from a smoothie, and you will sense different flavors of fruits. However, this feeling will vary depending on the amount and characteristic tastes of fruits. For example, if there are lots of oranges in it, the high acidity due to the number of oranges might mask a less noticeable taste such as blueberry. Things become more complicated when you want to predict the ratio of each fruit that was blended for the smoothie. What if you have the opportunity to taste each fruit separately without making a smoothie? In the context of this analogy, the conventional bottom-up proteomic approach refers to analyzing protein levels by mixing all cells like a smoothie. Thus, the measurement of proteins would be the average of all cells that were involved in the study. Distinguishing the taste of blueberries in this case low-abundant species, rare cell types, and sub-populations becomes more complicated. However, single-cell proteomics enables us comprehensive understanding of cellular heterogeneity such as the immune system and cancer formation phases [2]. Since the combination of single-cell methods with MS-based proteomic studies was implemented less than 10 years ago, some challenges still need to be solved in terms of depth of coverage, sensitivity, robustness, and cost [1]. In this article, they proposed a novel true single-cell proteome (T-SCP) pipeline by enhancing sensitivity and optimizing the mass spectrometry (MS) setup.

They introduced the PASEF acquisition scheme for noise-reduced quantitative mass spectra, enabling highly sensitive and complete proteome measurements. Through a dilution series of HeLa cell lysate, they identified over 550 proteins from low amounts and achieved excellent quantitative reproducibility. To enable true single-cell proteomics, they significantly increased MS sensitivity and adapted their workflow. The researchers sorted and analyzed individual single HeLa cells, achieving protein identification and quantification. Sensitivity improvements allowed the identification of more than 1,890 proteins from as few as six single cells. They established a «core-proteome» subset of stable proteins that could serve as normalization factors and identified distinct protein regulation mechanisms.

Further sensitivity enhancements, including reduced flow rates, enabled a ten-fold increase in sensitivity, leading to the identification and quantification of over 3,900 HeLa proteins from just 1 ng of material. They adopted the diaPASEF acquisition mode for increased reproducibility.

Applying this technology, the study explored the cell cycle’s impact on single-cell proteomes. They treated HeLa cells with thymidine and nocodazole, quantifying up to 2,501 proteins per single cell across different cell cycle stages. The data revealed high quantitative precision and allowed the differentiation of cell cycle stages based on proteomic profiles. Using marker proteins, the researchers successfully assigned cellular states and distinguished cell cycle phases.

The single-cell proteome analysis also revealed differential expression of known cell cycle regulators and highlighted novel proteins associated with the G2/M transition.

Figure 1. A novel mass spectrometer allows the analysis of true single-cell proteomes. (Second figure in the article). A: Raw signal increase from standard versus modified TIMS-qTOF instrument (left) and at the evidence level (quantified peptide features in MaxQuant) (right). B: Proteins quantified from one to six single HeLa cells, either with MBR in MaxQuant (orange) or without MBR (blue). The outlier in the three-cell measurement in gray (no MBR) or white (with MBR) is likely due to failure of FACS sorting as it identified a similar number of proteins as blank runs (Horizontal lines within each respective cell count indicate median values). C: Quantitative reproducibility in a rank order plot of a six-cell replicate experiment. D: Same as C for two independent single cells. E: Rank order of protein signals in the six-cell experiment (blue) with proteins quantified in a single cell colored in orange. F: Raw MS1-level spectrum of one precursor isotope pattern of the indicated sequence and shared between the single-cell (top) and six-cell experiments (bottom).

SC proteomes compared to transcriptomes:

The study evaluated over 430 single-cell proteomes and compared them with Drop-seq and SMART-Seq2 scRNA-seq data to gain technology-independent insights. Proteome measurements exhibited higher average correlations among cells compared to scRNA-seq methods. Protein completeness per cell followed a normal distribution, with proteomic 

capturing about 49% of observed proteins, whereas SMART-Seq2 captured only 27% and Drop-seq captured 8%. Detecting limitations in protein measurements, bimodality in lower protein abundance range indicated potential benefits of imputation or tailored parameter estimation methods. While bulk transcript and protein levels showed moderate correlation, single-cell levels diverged, underscoring distinct regulatory mechanisms.

Examining shared gene CVs, single-cell transcriptomes displayed consistent quantitative variation, unlike proteomes. This emphasized different single-cell regulation for protein and RNA abundance. The research concluded that protein and RNA measurements offer complementary information, uncovering unique regulatory mechanisms. A stable core proteome subset of top 200 proteins with low CVs was identified, representing normalization factors and essential cellular processes. This subset’s distribution across the proteome’s dynamic range indicated stability even during remodeling. Overall, the study deepens insights into single-cell proteome dynamics and its interplay with gene expression.

References 

[1]       K. Vandereyken, A. Sifrim, B. Thienpont, and T. Voet, “Methods and applications for single-cell and spatial multi-omics,” Nat Rev Genet, p. 1, Aug. 2023, doi: 10.1038/S41576-023-00580-2.

[2]       E. Flynn, A. Almonte-Loya, and G. K. Fragiadakis, “Single-Cell Multiomics,” https://doi.org/10.1146/annurev-biodatasci-020422-050645, vol. 6, no. 1, Aug. 2023, doi: 10.1146/ANNUREV-BIODATASCI-020422-050645.

[3]       “What is single-cell sequencing? – Single Cell Discoveries.” https://www.scdiscoveries.com/blog/knowledge/what-is-single-cell-sequencing/ (accessed Aug. 21, 2023).

[4]       A. Brunner et al., “Ultra-high sensitivity mass spectrometry quantifies single-cell proteome changes upon perturbation,” Mol Syst Biol, vol. 18, no. 3, Mar. 2022, doi: 10.15252/MSB.202110798.

The post Exploring Cellular Complexity: Unveiling Single-Cell Proteomics appeared first on PROTrEIN.

]]>
Modeling Lower-Order Statistics to Enable Decoy-Free FDR Estimation in Proteomics https://protrein.eu/blog/modeling-lower-order-statistics-to-enable-decoy-free-fdr-estimation-in-proteomics/ Wed, 23 Aug 2023 11:59:13 +0000 https://protrein.eu/?p=1902 In our previous Journal Club blog post, we discussed «nanopore profiling,» a method for identifying proteins within the field of proteomics. Today we proceed with «how to validate the identification statistically other than the traditional target-decoy method with an improved version of the decoy-free approach» based on Dominik et al.’s article «Modeling lower-order statistics to […]

The post Modeling Lower-Order Statistics to Enable Decoy-Free FDR Estimation in Proteomics appeared first on PROTrEIN.

]]>
In our previous Journal Club blog post, we discussed «nanopore profiling,» a method for identifying proteins within the field of proteomics. Today we proceed with «how to validate the identification statistically other than the traditional target-decoy method with an improved version of the decoy-free approach» based on Dominik et al.’s article «Modeling lower-order statistics to enable decoy-free FDR estimation in proteomics«, we present in this blog post a new method for validating identification in proteomics.

There are two methodologies for calculating the false discovery rate (FDR) in proteomics: decoy-based and decoy-free. In decoy-based approaches, spectra are compared to the sequences of naturally occurring (target) proteins and decoys that are in-silico generated, based on peptide sequences of the target database. Using decoy peptide spectrum matches (PSMs), the conventional target-decoy FDR calculation takes into consideration the characteristics of incorrect target PSMs. Depending on the decoy generation method, this increases the cost of computation and decreases the likelihood that accurate PSMs will be considered. In addition, decoy-based methods may under- or overestimate the FDR if the scoring function used in the step of searching the target-decoy database has some bias or if there are insufficient decoys in the region where the models of the correct and incorrect PSMs overlap. Decoy-free statistical validation tools that employ only target PSMs constitute another category of peptide identification validation methods. While the majority of decoy-free statistical validation tools model the score distribution of the highest-scoring PSMs, some also exploit the PSMs with lower scores. The current decoy-free methods typically lack a solid theoretical foundation and rely significantly on the empirical characteristics of the data, or they rely on theoretical assumptions that are not always well-justified.

In this article, the new idea of decoy-free FDR estimation is propose with a semiempirical framework based on lower-scoring target PSMs where as a relationship between the parameters of the distributions of low-order statistics of the log transformed e-value (TEV) score and a necessary empirical optimization to fit a single parameter to real data. The theoretical sharing of parameters μ and β across different orders of Target E-Value (TEV) distributions in proteomics. However, empirical estimation of these parameters reveals slight deviations from the theoretical values. To address this, the article proposes two semiempirical optimization methods for estimating adjusted μ and β values for the Top Null Model (TNM). The methods include a linear regression-based approach and a mean β-based approach. These approaches are applied to data sets, and the best TNM variant is selected based on the Bayesian information criterion (BIC). 

The performance of different Top Null Models (TNMs) estimated using lower-order models generated by the Tide and Comet search engines was evaluated. The evaluation focused on FDR estimation and compared the TNM approaches against Couté’s method, the Gumbel TEV model, and the common decoy distribution (CDD) method. The comparison considered metrics such as false discovery proportion (FDP) and the number of correctly identified spectra at different FDR thresholds. FDR control was performed using the Benjamini-Hochberg (BH) procedure. A ground truth data set was created by searching files against a target-decoy database and generating incorrect PSMs. The TNMs were generated based on the top-scoring PSMs, and the optimal μ and β estimates were determined using the proposed semiempirical estimation methods. FDR control was then applied using the BH procedure, and the results were evaluated using the ground truth labels. The process was repeated on bootstrapped samples, and mean FDP and correct identification values with confidence intervals were calculated.

To test the quality of lower-order models and top null models, data sets of natural peptides from five different species (H. sapiens, M. musculus, A. thaliana, S. cerevisiae, and E. coli) were taken from project repositories in the PRIDE archive. Synthetic human peptide data sets from the ProteomeTools project were used in the validation study. Only MS2 spectra with charge states of 2+, 3+, and 4+ were taken into consideration for the evaluation. The lower-order models were found to fit the empirical data well, with some discrepancies for lower order indices. The maximum likelihood estimation (MLE) method performed better than the method of moments (MM) for parameter estimation in the lower-order models. The proposed models accurately fit the empirical distributions and were not significantly affected by the size difference between the analyzed data sets. The study also compared the performance of the lower-order models with decoy-based models and Couté’s method, and found that the lower-order models estimated FDRs better, particularly for Tide results. However, for Comet results, Couté’s method and decoy-based models overestimated FDRs, while the common decoy distribution method provided second-best estimates. The differences in performance between Tide and Comet can be attributed to the less rigorous e-values produced by Comet. The proposed approach, combining theoretical foundations with empirical optimization, showed resistance to issues associated with statistical scoring in shotgun proteomics.

The proposed approach eliminates the need for decoy sequences and offers improved accuracy compared to alternative methods. While further evaluation and tuning may be necessary for different identification tools, this work highlights the untapped potential of lower-scoring PSMs in enhancing statistical validation methods in proteomics research.

Finally, thanks for reading our post and keep tuned for further content!

The post Modeling Lower-Order Statistics to Enable Decoy-Free FDR Estimation in Proteomics appeared first on PROTrEIN.

]]>
Peptide De Novo Sequencing What are the ingredients of that delicious pizza? https://protrein.eu/blog/peptide-de-novo-sequencing-what-are-the-ingredients-of-that-delicious-pizza/ Tue, 08 Aug 2023 10:07:26 +0000 https://protrein.eu/?p=1870 Proteins, the mighty microscopic marvels present in what we eat, hence, what we are. In milkshakes 🥤, ice creams 🍦… your hair, your skin… you 👤. These molecular machines are present in all living things, from viruses to the dog 🐶 that barked at you the other day. They are so essential that they keep […]

The post Peptide De Novo Sequencing What are the ingredients of that delicious pizza? appeared first on PROTrEIN.

]]>
Proteins, the mighty microscopic marvels present in what we eat, hence, what we are. In milkshakes 🥤, ice creams 🍦… your hair, your skin… you 👤. These molecular machines are present in all living things, from viruses to the dog 🐶 that barked at you the other day. They are so essential that they keep the grand show of life going. Proteins are like the ultimate multi-taskers of the cellular world. They’re the tiny construction workers 🏗, dutiful soldiers, speedy messengers, vigilant guards and the skilled repair crew 🔧 that keep our bodies running smoothly day in, day out.

They’re crafted by our cells using the blueprints written in our DNA. Heard of the genetic code in the DNA? This code is like a huge library 📚, stuffed with cookbooks. These cookbooks contain genes, which are like recipes for whipping up every protein your body needs. And there are at least 10000 different proteins keeping you alive.

These proteins are built inside your cells using tiny molecules called amino acids 🧪, which are like the ingredients in your recipe. These cells are guided by the instructions in the DNA on how to put together particular proteins. Picture a cake 🍰, it needs flour, eggs, sugar, and butter, all in specific quantities and added in a certain order. Similarly, proteins need specific amino acids in the right order to be cooked up right. DNA has the recipe, cookbooks 📚. Amino acids are the ingredients. Proteins are the dishes 🍽. Hungry yet? 😋 

As an example, the following genetic code in your DNA, could be the cookbook instruction 📖 for creating the dish (protein) we want. 

TGT – TAC – ATT – CAA – AAT – TGT – CCT – CTC – GGT 

These instructions would lead the cells to place amino acids (ingredients) one after another leading to a molecular chain of amino acids placed in the following order. 

Cysteine – Tyrosine – Isoleucine – Glutamine – Asparagine – Cysteine – Proline – Leucine – Glycine 


CYIQNCPLG 

In this case, the instruction earlier in the DNA was for a small protein (technically peptide, but we will get to that.) called Oxytocin 😍, often referred to as the «love hormone» or the «cuddle chemical.» It plays a crucial role in social bonding, fostering feelings of trust, empathy, and connection. The amino acid sequence for this oxytocin is CYIQNCPLG.
There are over 20 amino acids, and only using 20 of these, our bodies create 10,000 different types of proteins, with each protein playing a different role. For example, hemoglobin protein 💉 is made by the cells in your bone marrow, because the cells are instructed by your DNA to do so, which is then released into your bloodstream. Tyrosinase is another protein, which is responsible for producing melanin, which gives your hair the color 🌈. Loss of this protein, or cells choosing not to produce it can lead to loss of hair color. 

While oxytocin is just 9 amino acids, hemoglobin is about 300 and Tyrosinase is over 500 amino acids long.

So, What is de novo sequencing?

Here’s the big challenge. Let’s call it «de novo sequencing» which is a fancy way of saying «figuring out the ingredients of a protein by tasting it» 🕵️‍♀️🔍. It’s like heading to a restaurant, eating a lip-smacking dish 🍲, and then trying to identify the ingredients just based on your taste! Sounds fun but tough, doesn’t it? The goal is to identify the composition and sequence of the protein. That is, to identify the amino acid pattern. 

At this point, allow me to introduce the superstars of our story — peptides. Peptides are like mini proteins. Think about a giant pizza 🍕(that’s your protein), peptides are like those cute mini pizza bites, smaller in size but oh-so-crucial. Peptides are just shorter chains of these all-important amino acids. Remember oxytocin, that’s a peptide, a mini protein. 💖 

From now on, we’re shining the spotlight 🎯 on these little strands of amino acids, the peptides. Let’s think of them as the underdogs, small but mighty, and loaded with info about our bodies. 🏋️‍♀️ 

So, why the heck do we want to figure out the sequence of the peptide? 🤷‍♀️ Well, if we crack the sequence of amino acids – the ‘secret ingredients’ 🧪 – in our peptide, we unlock a whole treasure chest 🗝🔓 of information about what it is and what it does. It’s like decoding a secret message written in an alien language! 👽📜 

By identifying a peptide in a sample, we could potentially develop a new wonder drug 💊 that uses the peptide, or a treatment that blocks it. It’s like having a secret weapon in the battle against diseases! 🛡 We could even learn something groundbreaking about how the body works! 🧠⚡ It’s like discovering a new law of physics… but inside our bodies! How cool is that? 🤓🚀

So, how do we do it? 

Well, we could just squint really hard at them under a microscope 🦠🔬. But trust me, even if you have the eyes of an eagle, you wouldn’t see much. They’re way too small. Sure, there have been some incredible advances recently with techniques that read each amino acid, but it’s still not always possible.

How about just weighing it? 🏋️‍♀️ If my peptide weighs, let’s say, 200 (in super-simplified units), then I could whip out a catalog 📚 and see that 200 is the mass of oxytocin. So, it must be oxytocin, right? Ehhh… not so fast. There could be a whole bunch of other peptides out there that also weigh 200. Worse, what if this peptide is from a snake venom 🐍 that’s never been observed before and hence, doesn’t even make it to the catalog?

Aha! Scientists early on had a light bulb 💡 moment. They thought, «Let’s break the peptide into pieces, then weigh the pieces.» It’s like disassembling a Lego tower to understand how it was built.

But then, a new challenge raises its ugly head 🐲. How do we weigh so many tiny pieces at once? It’s like trying to weigh a bunch of confetti 🎉 – all at the same time!


Enter the unsung heroes of our story – mass spectrometers 🌠. These are like high-tech super scales 🧱⚖ that can weigh a ton of things all at once. 

But wait, there’s more! We usually measure a lot of things at once, and mass spectrometers don’t just stop at weighing. They also measure quantities 🔢. It’s like a super-smart scale saying… «Oh, your sample had more M&Ms 🍫 than Coffee beans ☕ in your Mocha coffee bean cookie 🍪.» How cool is that!? 

Let’s say we’ve got our hands on a peptide called CAT (not the furry creature 🐱 but a peptide whose amino acid sequence is C, A, and then T). In fact, let’s say we have a whole pile of them, like a giant CAT party 🎉. But before we can start the weighing party, we’ve got to break them into smaller parts. Kind of like smashing a piñata 🎊. And we have some pretty fancy ways to do this (imagine sophisticated scientific pinata smashers like CID, ETD etc.). 

When we bash our CAT peptide, it can break into C, CA pieces. 

ADITI can break into A, AD, ADI and ADIT and so on. 

So, let’s say, we’ve got our CAT peptide broken into C and CA parts. Now, when we feed these pieces into our trusty mass spectrometer to weigh, we should see two distinct signals 📈📉. One corresponds to the smaller C piece and another to the larger CA piece. 

But hold on, what about the whole, unbroken CAT peptide? Well, it can still hang around in our sample, giving us a third signal 📊. 

So, in the end, we’ve got three peaks on our mass spectrometer’s readout, one for each of the C, CA, and the intact CAT. 

For example, talking about CAT, we know that amino acid C weighs about 103 giving us a peak at 103. Similarly, we know that A weighs about 70, so the mass of CA must be 103 + 70 = 173, where we see the second peak. The mass of amino acid T is about 100, giving us another peak from CAT, at 103 + 70 + 100 = 273. 

The same process happens with a peptide called ADITI. The peak at 300 is due to ADI, and as T is about 100, we get a peak at 400 corresponding to ADIT. 

The values can be generated for any sequence. Here’s a tool to try it out yourself. (use b ions checkbox only in the tool, this will be explained in detail later) 

Fragment Ion Calculator – systemsbiology.net 

But let’s go back to CAT.. Let’s say, you didn’t know this was CAT peptide and we told you that this peptide had three amino acids. Can you figure out the amino acid sequence? Are you ready to perform De Novo Sequencing?

Let’s do De Novo Sequencing 

So, you have taken an unknown peptide, smashed it and then sent it into the mass spectrometer. The mass spectrometer has spit out a spectra with three humps. Time to figure out what was in the sample. 
To begin, let’s just call the unknown peptide X1 X2 X3.

The furthest peak must be from the largest mass, X1 X2 X3 itself, the unbroken peptide.  

So, X1 X2 X3 weighs 273. This cannot directly tell us anything, as we know the masses of amino acids only. And yes, we are approximating here for simplicity.

One can see that the peak for the smallest mass is 103 which should be the smallest fragment. So in our case, X1 is the smallest fragment and must have a mass of 103. This is a single amino acid. But which is it? This is when you go through the amino acid list. C is an amino acid with a mass of 103, so X1 must be C. 

The next peak is at 173 and it must be from X1 X2. So, X1 X2 weighs 173. 

This tells us that X2 must be an amino acid with a weight 70. Go through the amino acid list, there is an amino acid corresponding to 70, it’s A, Alanine. So, X2 must be A. And finally, the peak at 273 must be from the whole peptide X1 X2 X3, so since X1 X2 is 173, X3 must be an amino acid with mass 100. Which our list tells us, X3 must be T.

So.. we now know X1 X2 X3 is CAT. Congrats. That’s your first de novo sequencing. 

Now that you know, wanna try out another? Here’s some help, this one is 5 amino acids long peptides.

And the solution is…

TEAM.

CEAM? That’s possible too.. If we consider 102 to be associated with C. This is a very simplified version. Because, isn’t T, 101.04?  
Yep, you caught us! 🎣 We did some rounding up, sort of like rounding up π to 3 (although not quite as dramatic!). However, it’s important to note there’s another simplification here that matters a lot. The ends of these peptides – imagine them like the caps on a tube of toothpaste 🦷 – can have some effects on the values we see in the spectra. 

We hope you enjoyed this read and got a grasp of what De Novo Sequencing is. Before we dive deeper into the topic, let’s take a breath. Stay tuned for the second part of the post, in which we will also play some fun puzzles 🧩.

References :
1) Main Paper Reference : Medzihradszky, K.F. and Chalkley, R.J. (2013) ‘Lessons inde novopeptide sequencing by Tandem Mass Spectrometry’, Mass Spectrometry Reviews, 34(1), pp. 43–63. doi:10.1002/mas.21406.
2) Secondary Paper : CHONG, K.F. and LEONG, H.W. (2012) ‘Tutorial on de novo peptide sequencing using MS/ms mass spectrometry’, Journal of Bioinformatics and Computational Biology, 10(06), p. 1231002. doi:10.1142/s0219720012310026.
3) Blog Post Referred : https://www.ionsource.com/tutorial/DeNovo/DeNovoTOC.htmI must add. 4) Tutorial «Manual spectra annotation and automatic database/library search» Viktoria Dorfer, FHOOE, PROTrEIN summer School 2021:
https://www.youtube.com/watch?v=ztyglWkY1iI

The post Peptide De Novo Sequencing What are the ingredients of that delicious pizza? appeared first on PROTrEIN.

]]>
Mass spectrometry-based proteomics imputation using self-supervised deep learning https://protrein.eu/blog/mass-spectrometry-based-proteomics-imputation-using-self-supervised-deep-learning/ Thu, 03 Aug 2023 13:51:21 +0000 https://protrein.eu/?p=1858 Hello and welcome back to another issue of the PROTrEIN Journal Club! This occasion we will cover important topics in proteomics, missing values and imputation. We will try to shed light to some of the challenges regarding these matters with the aid of the article titled: “Mass spectrometry-based proteomics imputation using self supervised deep learning” […]

The post Mass spectrometry-based proteomics imputation using self-supervised deep learning appeared first on PROTrEIN.

]]>
Hello and welcome back to another issue of the PROTrEIN Journal Club! This occasion we will cover important topics in proteomics, missing values and imputation. We will try to shed light to some of the challenges regarding these matters with the aid of the article titled: “Mass spectrometry-based proteomics imputation using self supervised deep learning” from Henry Webel et al.1 Also after a few weeks of hiatus, we are back with some machine learning too.

The search for biomarkers and the identification of new drug targets are important use cases of mass spectrometry (MS) based label-free proteomics. However, the downstream analysis of acquired data is largely impacted by missing values. There can be many roots of missing values, but the main contributing factors can be divided into two categories: biological factors such as proteins not existing in the sample or their abundance is below instrument detection limit and analytical factors, for instance poor ionisation efficiency, bad peptides-spectrum matches, stochasticity of precursor selection for fragmentation.2

To address the problem of missing values, different imputation methods were developed. These methods can impute quantification values on various levels such as precursor, aggregated peptides and protein group levels. One common approach is median imputation per feature across samples and another is interpolation of missing features by close replicates. A more sophisticated approach exists that imputes data at protein group level using random draws from down-shifted normal (RSN) distribution. The assumption here is that the values are missing due to absence or lower abundance in the sample than the detection limit. However, that can create biases and skew the downstream analysis. The authors of the paper are presenting three machine learning models to predict missing quantification values. 

These  alternative deep learning (DL) models use different strategies —collaborative filtering (CF), denoising autoencoder (DAE), and variational autoencoder (VAE)— to impute missing values in proteomics data sets. The training objectives, complexity, and therefore capabilities of the models are different which led authors to evaluate their performance in comparison to each other. The CF and autoencoder objective only focuses on reconstruction, whereas the VAE adds a constraint on the latent representation. Furthermore, the first two modeling approaches use a mean-squared error (MSE) reconstruction loss, whereas the VAE uses a probabilistic loss to assess the reconstruction error.

These models were applied to large (N≈450) and small (N≈50) MS-based proteomics data sets of HeLa cell line tryptic lysates acquired over two years during continuous quality control in two different labs at Novo Nordisk Foundation Center for Protein Research (NNF CPR) and Max Planck Institute of Biochemistry. The effectiveness of the models was assessed in comparison to two heuristic-based methods: median imputation and interpolation of missing features. The results show that the self-supervised models, i.e. CF, DAE, and VAE, outperform the heuristic-based approaches, with half of the median imputation mean absolute error (MAE). By identifying (+23.6%) more significantly differentially abundant protein groups, the VAE model in particular is demonstrated to be useful in illness prediction.

The two autoencoder architectures represented a sample in a low-dimensional space using all of the data. The CF model, in contrast, needed to learn a latent embedding space for both the samples and the features. Overall, While the DL methods and median imputation can impute all missing values, interpolation does not replace missing values in case a value is missing in all replicates. The study also discovers that the models’ performance changes based on how frequently a protein group is observed, with better performance for groups observed in more than 80% of the samples. For protein-level data, the models’ overall performance is shown to be the poorest, for aggregated peptides, it is better, and for precursors, it is best.

In the development datasets, while the DL techniques outperform interpolation and have around half the median imputation MAE, The three DL approaches perform about the same. Consequently, when compared to the self-supervised models, the median imputation and interpolation models performed about 1.8–2.4 times worse.

Performance of imputation methods at the level of protein groups, aggregated peptides, and precursors for MaxQuant outputs

The authors were testing the impact of their developed imputation techniques on a real-world dataset of 455 blood plasma proteomics samples from a cohort of alcohol-related liver disease (ALD) and healthy controls. The study3 where the real-world data originated from, was looking for biomarkers of ALD in the proteomics samples that could enable MS-based liver disease testing. One of the key pathological features of alcohol-related liver disease is fibrosis, therefore proteins related to fibrosis were monitored. In addition then they were training machine learning models to predict fibrosis and inflammation from the MS plasma protein groups. In the referenced article3, they used the RSN imputation approach, hence the authors compared their methods to the results of that publication. From the PIMMS methods the authors selected the variational encoder model for imputation and found 23.6% more differentially expressed proteins. They then investigated whether the differentially regulated proteins can be associated with disease using the DISEASE database, to find that 20 of these proteins had an association entry to fibrosis. With the newly found proteins the authors retrained the predictive model from the original study for liver condition development. They found that the retrained model performed as good or slightly better than the original model, concluding that these proteins can have predictive power.

Just like in the previous editions of PROTrEIN Journal Club we have sent our questions to the authors to conduct a short interview with them and gain more insight in their work. Please read our short interview below:

Blog post team: How self-supervised deep learning models that you chose contribute to the imputation of missing values in label-free quantification (LFQ) proteomics data? What advantages do they offer compared to other models?

Henry Webel: The models are in the category of machine learning models. In comparison to let’s say a random forest, you will additionally get embeddings of the features and samples, either separately (collaborative filtering) or joined (Autoencoder based architectures). Clustering in the embedding – also called latent – space could be used to compare it to e.g. hierarchical clustering of the original data. 

Blog post team: The paper suggests assessing if machine learning models can be trained on lower-level data. Could you elaborate on this suggestion and discuss the potential benefits and challenges associated with training machine learning models using lower-level data in the context of MS-based proteomics?

Henry Webel: In mass spectrometry- based bottom-up proteomics the unit of measurement are ions of peptides. The aggregation to protein groups therefore normally implies an implicit imputation using a neutral element as described by Lazar et al.4 Therefore, the elements of interest should be rather measured peptides, but currently the dominant approach is to use aggregated protein groups. 

Blog post team: It is mentioned that performance of the self-supervised models was better than heuristic approaches, which included median, interpolation or shifted normal distribution imputation, what was the main reason for that and could it be affected somehow by features of the data or size of the data? Do you think it is applicable to use supervised models instead of self-supervised models?

Henry Webel: The performance comparison always depends on the design of the comparison. We therefore extended the down-stream comparison in a revised version of the article. Additionally, we added supervised models such as random forests published as R packages. The main idea is to allow users to compare how well the imputation approaches perform on missing completely at random data (MCAR).

Blog post team: What are the future outlooks for PIMMS? Are you planning further developments?

Henry Webel: We added now many R methods to the comparison. Otherwise it would be great to extend the comparison by other methods of creating the validation and test data splits to get an overview of the methods used in the field.

References:

1. Webel, H. et al. Mass spectrometry-based proteomics imputation using self supervised deep learning. bioRxiv 2023.01.12.523792.

2. Jin, L. et al. A comparative study of evaluating missing value imputation methods in label-free proteomics. Sci. Rep. 11, 1760 (2021).

3. Niu, L. et al. Noninvasive proteomic biomarkers for alcohol-related liver disease. Nat. Med. 28, 1277–1287 (2022).

4. Lazar, C., Gatto, L., Ferro, M., Bruley, C. & Burger, T. Accounting for the Multiple Natures of Missing Values in Label-Free Quantitative Proteomics Data Sets to Compare Imputation Strategies. J. Proteome Res. 15, 1116–1125 (2016).

The post Mass spectrometry-based proteomics imputation using self-supervised deep learning appeared first on PROTrEIN.

]]>
Nanopore profiling: a scalable approach to protein identification https://protrein.eu/blog/nanopore-profiling-a-scalable-approach-to-protein-identification/ Mon, 22 May 2023 10:24:17 +0000 https://protrein.eu/?p=1838 Proteomics research relies heavily on mass spectrometry, which has emerged as the most prominent and widely utilized approach for the identification of proteins. However, progress is being made in other analytical methods for peptides and proteins. In this month PROTrEIN ITN’s Journal Club, we are taking a glance at the characterization of proteins with nanopores […]

The post Nanopore profiling: a scalable approach to protein identification appeared first on PROTrEIN.

]]>
Proteomics research relies heavily on mass spectrometry, which has emerged as the most prominent and widely utilized approach for the identification of proteins. However, progress is being made in other analytical methods for peptides and proteins. In this month PROTrEIN ITN’s Journal Club, we are taking a glance at the characterization of proteins with nanopores through a publication from the University of Groningen entitled “Protein identification by nanopore peptide profiling”1.

Nanopores are naturally occurring protein channels in cell membranes through which molecules can pass. Nanopore sequencing and profiling approaches make use of these structures to identify molecules. The nanopore is embedded in a membrane that separates two chambers filled with an electrolyte solution. An electrical potential is then applied across the membrane, creating a current that flows through the nanopore. The analytes that travel through the pore disrupt the ion current, which can then be measured (Fig. 1). The principle behind nanopore sequencing is based on the fact that different molecules (nucleotides in the case of nanopore sequencing) have different sizes and shapes, and therefore induce distinct changes in the ion current. Currently, nanopore technologies are the most commonly used for the sequencing of DNA, but progress is being made in their utilization for the identification of peptides and proteins. 

Fig. 1: Graphical overview of the nanopore protein fingerprinting approach. Peptides are pre-hydrolyzed by a specific protease (e.g. trypsin) and the resulting peptides are measured as they translocate the nanopore. Each peptide entering the nanopore reduces the open pore current (Io) to the blocked pore current (IB). The resulting excluded current (ΔIB = Io − IB) relates to the volume of the peptide. The subsequent histogram of the percent of excluded currents (Iex %= ΔIB/ IO %) is used to identify the protein. Source: Nat Commun (2021) 12, 5795

In the presented article the authors use nanopores to characterize individual proteins following their digestion by trypsin. After digestion, peptides go into the nanopore analyser and generate a change in ion current when they pass through the pore. The protein information is collected in an excluded current spectrum, which is a histogram that summarizes the translocation events of all the peptides. Since each change in the ion current relates mainly to the analyte volumes, the resulting spectrum is a representation of the peptide volumes after protease digestion and can be used for the identification of the original protein.

The identifications obtained with the nanopore analyses are compared to the results collected with ESI-MS (electrospray ionization mass spectrometry). For the comparison, the peptides identified through the mass spectrometry analysis are mapped into an inferred excluded current spectrum. This spectrum is generated through a computational calibration algorithm and compared to the nanopore-obtained one through spectral matching techniques. 

Based on the performed comparisons, the authors observe that the reproducibility of the obtained spectra is quite high. Although for specific proteins the correlation between the nanopore-observed spectrum and the inferred one is lower, this could be due to the computational inference of the spectra based on the peptide mass rather than the volume, which is key in nanopore analyses.In conclusion, the authors observe that although the nanopore analyzers might require improvements in terms of resolution, they could constitute low-cost solutions for performing high-throughput analyses. They especially point out that this technique could offer considerable advantages when having small amounts of material, as is the case of low-abundant proteins or heterogeneous ones. For this reason, mass spectrometry and nanopore techniques could even be used in parallel, as they show different and complementary strength when it comes to protein identifications.

Q&A with Florian Lucas, the first author of the publication

Since nanopore technologies can identify single molecules and do not require large amounts of material, do you foresee a wider use of nanopore peptide profiling in single-cell proteomics?

Nanopore profiling gains its strength from the low sample volumes and high sensitivity able to pick up, previously undetected, unidentified contaminants in our filtered water supply2. They may therefore find their way into some single-cell proteomic pipeline. However, their development is some years from mainstream application due to several engineering challenges remaining for the sensory technology. Nonetheless, these are inevitably solved in due time, as exemplified by commercial nanopore DNA sequencing devices.

More specifically, while single prokaryotic cells may provide currently unachievably small numbers of analytes, it is undeniable that nanopores can be used to detect proteolytic peptides from single eukaryotic cells. These can remain largely undiluted as the total measurement volumes of state-of-the-art nanopore chambers is less than 150 nanolitres, which is still amenable to down-scaling. However, the theoretical resolution of peptide profiling using nanopore will not be sufficient to extract single-protein information in such complex samples. Rather there are two future developments I foresee for the technology. Firstly, the nanopore can be coupled with upstream separation techniques such as liquid chromatography to reduce sample complexity. Secondly, motor proteins can (in theory) be attached on top of the nanopore to directly sequence full proteins (one-by-one), similar to nanopore DNA sequencing. This concept has been shown by the groups of G. Maglia and C. Dekker in two independent ways3, 4.

The current resolution of nanopore technology does not seem to allow for the detection of ‘small’ modifications (e.g methylation). Do you think this approach can one day reach that level of accuracy?

This is an interesting question, as the answer cannot be expressed with a simple yes or no. It is important to acknowledge that nanopores do not detect the mass, rather, the displacement of ions. This makes them ideal for the detection of modifications that alter the ionic current and will therefore be able to detect some, but not all, modifications. Methylation in particular is a modification we can detect using nanopore DNA sequencing, and there is no doubt that this modification on peptides can be detected. Moreover, our recent work shows that glycans can be observed on peptides5. We have also previously shown that conformational differences between leucine and isoleucine can be discriminated6, and our colleagues displayed the discrimination of phosphorylated peptides7.

In the article, only one protein digest is analyzed at a time. How feasible would it be to scale the technique to more complex samples, and what would be the main challenges?

The field is still in the proof-of-concept phase, which is best compared to mass spectrometry in the 1980s. Even mass spectrometry as a stand-alone technique is unable to extract all information from highly complex samples. Instead, it builds on the synergy between analytical separation, e.g. liquid chromatography (LC) or capillary electrophoreses. These methods amplify the separation power to allow complex analysis. Unpublished results show that we can have a steady flow of several milliliters per minute across the nanopore. This is more than compatible with downstream flows of analytical LCs. However, we cannot exclude the possibility of nanopores sequencing proteins one by one when coupled with a motor protein, but this is still in its theoretical and very early proof-of-concept phase3, 4.

References

1. Protein identification by nanopore peptide profiling. Florian Leonardus Rudolfus Lucas, Roderick Corstiaan Abraham Versloot, Liubov Yakovlieva, Marthe T. C. Walvoort and Giovanni Maglia. Nat Commun (2021) 12, 5795. DOI: 10.1038/s41467-021-26046-9

2. The Manipulation of the Internal Hydrophobicity of FraC Nanopores Augments Peptide Capture and Recognition. Florian Leonardus Rudolfus Lucas, Kumar Sarthak, Erica Mariska Lenting, David Coltan, Nieck Jordy van der Heide, Roderick Corstiaan Abraham Versloot, Aleksei Aksimentiev, and Giovanni Maglia. ACS Nano 2021 15 (6), 9600-9613. DOI: 10.1021/acsnano.0c09958

3. Bottom-up fabrication of a proteasome–nanopore that unravels and processes single proteins. Shengli Zhang, Gang Huang, Roderick Corstiaan Abraham Versloot, Bart Marlon Herwig Bruininks, Paulo Cesar Telles de Souza, Siewert-Jan Marrink, and Giovanni Maglia Nat. Chem. 2021 13, 1192–1199. DOI: 10.1038/s41557-021-00824-w

4. Multiple rereads of single proteins at single–amino acid resolution using nanopores. Henry Brinkerhoff, Albert S. W. Kang, Jingqian Liu, Aleksei Aksimentiev, and Cees Dekker. Science 2021 374, 1509-1513(2021). DOI:10.1126/science.abl4381

5. Quantification of Protein Glycosylation Using Nanopores. Roderick Corstiaan Abraham Versloot, Florian Leonardus Rudolfus Lucas, Liubov Yakovlieva, Matthijs Jonathan Tadema, Yurui Zhang, Thomas M. Wood, Nathaniel I. Martin, Siewert J. Marrink, Marthe T. C. Walvoort, and Giovanni Maglia. Nano Letters 2022 22 (13), 5357-5364. DOI: 10.1021/acs.nanolett.2c01338

6. In silico assessment of a novel single-molecule protein fingerprinting method employing fragmentation and nanopore detection. Carlos de Lannoy, Florian Leonardus Rudolfus Lucas, Giovanni Maglia, Dick de Ridder. iScience 2021 24 (10), 103202. DOI: 10.1016/j.isci.2021.103202.

7. Label-Free Detection of Post-translational Modifications with a Nanopore. Laura Restrepo-Pérez, Chun Heung Wong, Giovanni Maglia, Cees Dekker, and Chirlmin Joo. Nano Letters 2019 19 (11), 7957-7964. DOI: 10.1021/acs.nanolett.9b03134

The post Nanopore profiling: a scalable approach to protein identification appeared first on PROTrEIN.

]]>
Papers and patents are becoming less disruptive over time https://protrein.eu/blog/papers-and-patents-are-becoming-less-disruptive-over-time/ Mon, 03 Apr 2023 14:15:46 +0000 https://protrein.eu/?p=1785 The fields of science and technology have been the engines of progress for many years, driving innovation and shaping the world we live in. Contributing to this scientific knowledge and progress is one of the main aspirations most of us researchers have. However, there is a growing number of studies suggesting that scientific progress is […]

The post Papers and patents are becoming less disruptive over time appeared first on PROTrEIN.

]]>
The fields of science and technology have been the engines of progress for many years, driving innovation and shaping the world we live in. Contributing to this scientific knowledge and progress is one of the main aspirations most of us researchers have. However, there is a growing number of studies suggesting that scientific progress is slowing in several fields (1–3). In today’s journal club, we present the work of Ph.D. candidate Michael Park, prof. dr. Erin Leahy and prof. dr. Russel J. Funk, titled “Papers and patents are becoming less disruptive over time” published in Nature last month, where they explore this phenomenon by conducting a large-scale analysis of the innovative activity in science and technology (4). This study is based on 25 million papers from the years 1945-2010 from the Web of Science and 3.9 million patents of the United States Patent and Trademark Office from the years 1976-2010, to try to understand both the extent of the slowdown in innovation and the reasons behind it.

To begin, the authors based their analysis on the distinction of two types of breakthroughs: consolidating contributions and disruptive contributions. The former are works that improve and contribute to further establishing existing knowledge, while the latter challenge current understanding making it obsolete and therefore driving science and technology towards new frontiers. To quantify this characteristic, the authors relied on a metric called the CD index (5), which is based on the number of citations a paper or patent receives and how these citations relate to previous work in the field. As Park et al. best put it: “[…] if a paper or patent is disruptive, the subsequent work that cites it is less likely to also cite its predecessors […] If a paper or patent is consolidating, subsequent work that cites it is also more likely to cite its predecessors”. Figure 1 illustrates the concept of CD index and shows some examples.

Figure 1: This figure shows a schematic visualization of the CD index. a, CD index value of three Nobel Prize-winning papers and three notable patents in our sample, measured as of five years post-publication (indicated by CD5). b, Distribution of CD5 for papers from WoS (n = 24,659,076) between 1945 and 2010 and patents from Patents View (n = 3,912,353) between 1976 and 2010, where a single dot represents a paper or patent. The vertical (up–down) dimension of each ‘strip’ corresponds to values of the CD index (with axis values shown in orange on the left). The horizontal (left–right) dimension of each strip helps to minimize overlapping points. Darker areas on each strip plot indicate denser regions of the distribution (that is, more commonly observed CD5 values). c, Three hypothetical citation networks, where the CD index is at the maximally disruptive value (CDt = 1), midpoint value (CDt = 0), and maximally consolidating value (CDt = −1). The panel also provides the equation for the CD index and an illustrative calculation.

By measuring the CD index of each paper and patent at 5 years after publication, the authors were able to determine that there is a massive decline in disruptiveness in science and technology across all major fields in the last decades (decline 91.9-100% for papers, 78.7-91.5% for patents) (figure 2). The authors found the same trend also using other indicators, namely the linguistic composition of published titles and abstracts. For example, the type-token ratio (unique words to total words) of paper and patent titles has declined significantly, especially before 1970 for papers and 1990 for patents, and a similar decline was observed in the combinatorial novelty of the words used. Finally, there is a decrease in the novelty of the combinations of previous work cited by papers and patents.

Fig 2. Decline in CD5 over time, separately for papers (a, n = 24,659,076) and patents (b, n = 3,912,353)

The authors found that the decline in disruptive activity was not due to a decrease in the quality of science and technology or an artifact of the CD index itself. They observed similar patterns of decline in disruptiveness when they computed the CD index using other data sources, such as JSTOR and PubMed. They also found that the decline was not due to changing publication or citation practices by conducting additional analyses, such as regression models and Monte Carlo simulations. Park et al. also considered the relationship between the growth of knowledge and the decline in disruptiveness. While they found conflicting results, with a positive effect of the growth of knowledge on disruptiveness for papers and a negative effect for patents, they observed a decline in the use of previous knowledge among scientists and inventors. This suggests that scientists and inventors are increasingly focusing on narrower slices of previous work, which could be limiting the potential for disruptive discoveries and inventions.
In conclusion, this study provides evidence of a decline in disruptive activity in the fields of science and technology and suggests that this decline may be related to a decline in the use of previous knowledge among scientists and inventors. This highlights the importance of fostering an environment in which scientists and inventors can engage with a diverse range of knowledge, which is crucial for driving disruptive discoveries and inventions.

References

  1. B. F. Jones, The Burden of Knowledge and the “Death of the Renaissance Man”: Is Innovation Getting Harder? Rev. Econ. Stud. 76, 283–317 (2009).
  2. N. Bloom, C. I. Jones, J. Van Reenen, M. Webb, Are Ideas Getting Harder to Find? Am. Econ. Rev. 110, 1104–1144 (2020).
  3. J. S. G. Chu, J. A. Evans, Slowed canonical progress in large fields of science. Proc. Natl. Acad. Sci. 118, e2021636118 (2021).
  4. M. Park, E. Leahey, R. J. Funk, Papers and patents are becoming less disruptive over time. Nature. 613, 138–144 (2023).
  5. R. J. Funk, J. Owen-Smith, A Dynamic Network Measure of Technological Change. Manag. Sci. 63, 791–817 (2017).

Q&A with Michael Park

Michael Park, one of the authors of the publication, kindly agreed to answer some questions we had:

Ayesha and Marc: Your work suggests that the observed decline in disruptive scientific findings is possibly related to the “publish or perish” culture. A deviation from this would necessitate deeper reforms in policies and is unlikely to occur in the short term. What would be your advice, to us as a network of Ph.D. students, to mitigate the effect of the “publish or perish” culture and increase our chance to produce truly innovative findings?

Michael: I would just point out that this is beyond the scope of our study and the scientific community would benefit from further research on how researchers can better navigate the potential pitfalls of the «publish or perish» culture. Nevertheless, I am happy to share some of my casual thoughts on this. Although it can be difficult, I think trying to not be too limited in the breadth of your research pursuit is important. For example, staying informed about the latest research in not only your own specific field but adjacent fields as well could be helpful. Actively seeking out collaborations with scientists from other disciplines may also be enriching. Although there may be many reasons why disruption is decreasing across time, our paper does suggest that a narrow research focus limited to specific fields is closely linked to nondisruptive research. 

Ayesha and Marc: We are currently seeing major breakthroughs in the field of artificial intelligence and conversational models. Do you think such AI tools can play a catalyzing role to help scientists digest and combine the increasing stock of knowledge? A scaffold, of sorts, that would help us climb on the shoulders of ever taller giants?

Michael: Again, the question is outside of the scope of the paper as well as my main research area since I don’t specialize in AI research. Nevertheless, I am happy to share my non-expert opinion if you would like. I think AI may help make certain parts of the research process more efficient, such as literature review, data analysis, and writing. However, I think the idea generation aspect of research, which is most influential in determining the extent to which a piece of work is disruptive, will largely remain a human task. Therefore, unless there is some AI technology that is capable of producing disruptive new ideas (there may well be as you point out), I personally don’t think AI adoption in research and education will lead to an increase in disruptive discoveries and inventions.

Ayesha and Marc: A more philosophical question: Some support the idea that because humans are biological organisms, they have a specific scope and limits, and that includes their cognitive capacities (for example “What Kinds of Creatures are We?” by Noam Chomsky). What is your opinion on this? Is scientific and technological progress doomed to halt someday? Could the decline in disruptiveness shown in your study be an early indication of this phenomenon?

Michael: I do agree that a human being does have cognitive limitations. This is part of the reason why when researchers are forced to publish many papers quickly, they limit the scope of their specialty, which seems to be linked to the decline in disruptiveness. However, I don’t think the linkage between cognitive limitations and research scope suggests that we are near the «doomed» day yet. Although individual human beings may be limited in their cognitive capacity, science progresses due to the efforts of a community of people. Researchers feed off of the advances but also the limitations of others’ works. So I don’t necessarily think that the cognitive capacities of human beings are detrimental to the creation of disruptive work. In addition, we show in Figure 4 that the number of highly disruptive works is pretty consistent across time. This likely suggests that we are not «running out» of disruptive things to discover.

Below is Figure 4 of the article mentioned above.

This figure shows the number of disruptive papers (a, n = 5,030,179) and patents (b, n = 1,476,004) across four different ranges of CD5 (papers and patents with CD5 values in the range [−1.0, 0) are not represented in the figure). Lines correspond to different levels of disruptiveness as measured by CD5. Despite substantial increases in the number of papers and patents published each year, there is little change in the number of highly disruptive papers and patents, as evidenced by the relatively flat red, green and orange lines. This pattern helps to account for simultaneous observations of both aggregate evidence of slowing innovative activity and seemingly major breakthroughs in many fields of science and technology. The inset plots show the composition of the most disruptive papers and patents (defined as those with CD5 values >0.25) by field over time. The observed stability in the absolute number of highly disruptive papers and patents holds despite considerable churn in the underlying fields of science and technology responsible for producing those works. ‘Life sciences’ denotes the life sciences and biomedicine research area; ‘electrical’ denotes the electrical and electronic technology category; ‘drugs’ denotes the drugs and medical technology category; and ‘computers’ denotes the computers and communications technology category.

Cover Image source: https://ellipse.prbb.org/the-art-of-publishing-a-scientific-article/

The post Papers and patents are becoming less disruptive over time appeared first on PROTrEIN.

]]>
Changing the proteomics shell towards the DIA-world https://protrein.eu/blog/changing-the-proteomics-shell-towards-the-dia-world/ Mon, 27 Mar 2023 15:07:08 +0000 https://protrein.eu/?p=1769 Data-independent acquisition (DIA) proteomics increasingly becomes the method of choice for researchers since it provides better reproducibility, identification rates, and accuracy compared to data-dependent acquisition (DDA). More and more tools are developed for DIA analysis and even established proteomics data processing software now can analyze DIA data. However, analysis of multiplexed spectra, characteristic of DIA, […]

The post Changing the proteomics shell towards the DIA-world appeared first on PROTrEIN.

]]>
Data-independent acquisition (DIA) proteomics increasingly becomes the method of choice for researchers since it provides better reproducibility, identification rates, and accuracy compared to data-dependent acquisition (DDA). More and more tools are developed for DIA analysis and even established proteomics data processing software now can analyze DIA data. However, analysis of multiplexed spectra, characteristic of DIA, remains challenging. This becomes especially crucial once research has more unknown parameters to consider besides just protein sequences – for example in the case of phosphorylation site identification.

In a paper titled “Rapid and site-specific deep phosphoproteome profiling by data-independent acquisition without the need for spectral libraries” Bekker-Jensen et al. strive to develop a reliable approach for DIA phosphorylation data analysis. They started with their optimized instrument settings for DDA and devised in a similar manner the optimized workflow for DIA Then they compared the quantification accuracy and precision of each workflow with a mixed-species approach showing that DIA was able to identify twice as many phosphopeptides as DDA while still accurately estimating the ratios used in mixtures. Next, they developed a PTM localization algorithm specific to DIA and tested its performance on a set of synthetic peptides with known phosphorylation site localization. Again, with DIA they identified more sites identified with lesser error rates compared to DDA. They also adapted a machine learning-based approach for calculating phosphorylation site stoichiometry achieving better precision and accuracy with DIA in comparison with the standard DDA approach in a mixed-species experiment. Last but not least, they did several kinase inhibitor assays to test the workflow in biological conditions and found that the results are consistent with current knowledge.

Figure 1. Comparison of DDA and different types of DIA using a kinase inhibitor assay. a Experimental workflow. bOverview of identified phosphopeptides, localized phosphosites, and ANOVA (s0 = 0.1, FDR 0.5) regulated sites for the different methods. c Heatmap of unsupervised clustering analysis of ANOVA-regulated phosphosites for DDA workflow (d) and for DIA workflow with project-specific library (e). Linear sequence motif analysis for two major clusters marked in colored boxes on heatmaps.

The main reason for choosing this article is to share the promising method for handling DIA data with the other fellows. They proposed directDIA (dDIA) in the paper as an alternative to classical DIA analysis whose processing procedure is more complicated since it does not need an experiment-specific spectral library. Alternatively, dDIA may eliminate the low abundance issue of phospho-site rather than conventional DIA. The number of identified phospho-sites and the combinations of them across all experiments was approximately similar to the DDA method. From this perspective, we commented that this approach still has to be enhanced in terms of sensitivity. On the other hand, In DIA, obtaining more distinctive spectra would be accessible by increasing the number of windows to extend fractionation (sampling). This can be more achievable with a more recent version of the mass spectrometry instrument.

Taking advantage of the latest developments in Large Language Models (LLM), we include here a fictional interview with the authors of the article. Also, the cover image of this post was created using some keywords of the abstract and a text-to-image AI tool. Would you have guessed it shall reflect phosphoproteomics?

PROTrEIN: We are very excited to have with us today Dr. Dorte Bekker-Jensen, Dr. Oliver Bernhardt, Dr. Alexander Hogrebe, Dr. Ana Martinez-Val, Dr. Lynn Verbeke, Dr. Tejas Gandhi, Dr. Christian Kelstrup, Dr. Lukas Reiter and Dr. Jesper Olsen from the University of Copenhagen and Biognosys AG. Welcome!

LLM: Thank you for having us.

PROTrEIN: We really enjoyed reading your paper and we have some questions for you. First of all, can you tell us what motivated you to develop this DIA-based phosphoproteomics method?

LLM: Sure. As you know, protein phosphorylation is one of the most important post-translational modifications that regulates cellular signaling networks. However, it is also very challenging to analyze phosphoproteomes on a large scale using conventional data-dependent acquisition (DDA) mass spectrometry, which relies on stochastic peptide selection and requires extensive spectral libraries for identification and quantification. We wanted to overcome these limitations and enable high-throughput phosphoproteomics with minimal sample preparation and analysis time.

LLM: That’s right. We decided to use data-independent acquisition (DIA) mass spectrometry, which is a more systematic and reproducible way of acquiring MS/MS spectra for all peptides in a given mass range. DIA has been successfully applied to proteomics and metabolomics, but not so much to phosphoproteomics, mainly because of the complexity and diversity of phosphopeptides and the lack of accurate site localization tools for DIA data.

PROTrEIN: Interesting. So how did you solve these challenges?

LLM: Well, we developed a new workflow that combines efficient phosphopeptide enrichment, fast LC-MS analysis, and advanced computational analysis using Spectronaut software. We optimized the DIA settings to achieve high resolution and sensitivity for phosphopeptides, and we incorporated an accurate site localization score into Spectronaut that uses fragment ion intensities and retention time information to assign phosphorylation sites confidently.

LLM: We also compared different strategies for generating spectral libraries for DIA analysis, such as using public databases, project-specific DDA runs, or direct DIA without any library at all. We found that direct DIA performed close to project-specific libraries, quantifying more than 20,000 phosphopeptides in 15 minutes single-shot LC-MS analysis per condition.
Pinar: Wow, that’s impressive. And how did you validate your method?

LLM: We applied our method to various biological systems, such as HeLa cells stimulated with epidermal growth factor (EGF), mouse liver tissue, and human plasma samples. We showed that our method can identify and quantify thousands of phosphorylation sites with high reproducibility and accuracy across different sample types and instruments.

LLM: We also demonstrated the ability of our method to determine phosphorylation site stoichiometry using a 3D multiple regression model-based algorithm that corrects for peptide abundance variations in DIA data. This allows us to measure the fraction of phosphorylated molecules at each site and compare them across different conditions.

PROTrEIN: That’s very cool. And what did you learn from applying your method to EGF signaling?

LLM: We used our method to systematically analyze the effects of 30 kinase inhibitors on EGF signaling in HeLa cells. We quantified more than 25,000 phosphorylation sites across 900 conditions in less than two weeks of LC-MS analysis time. We identified hundreds of kinase inhibitor targets and their downstream effects on EGF-regulated phosphorylation sites.

LLM: We also discovered some unexpected interactions between kinase inhibitors and EGF signaling pathways, such as the cross-talk between PI3K/AKT/mTOR and MAPK/ERK pathways, or the feedback activation of EGFR by some inhibitors. These findings reveal new insights into the complexity and dynamics of EGF signaling network.

LLM: We think that our method can be applied to other signaling systems and drug discovery projects, as well as other PTMs such as ubiquitination or acetylation. We also plan to further improve the speed and sensitivity of our method by using novel MS instruments and data analysis algorithms.

PROTrEIN: That sounds very exciting. Thank you so much for sharing your work with us today. It was a pleasure to talk to you.
We hope our readers enjoyed this episode and learned something new.

The post Changing the proteomics shell towards the DIA-world appeared first on PROTrEIN.

]]>
Are you a morning person? – Chronotypes, Circadian Rhythms, and Questionnaires https://protrein.eu/blog/are-you-a-morning-person-chronotypes-circadian-rhythms-and-questionnaires/ Wed, 22 Feb 2023 14:41:04 +0000 https://protrein.eu/?p=1753 Life as we know it has been shaped by the constraints of the environment. The aerodynamic shape of leaves to prevent the tree from toppling at high winds, or gravity that affects heights of organisms (extraterrestrial humanoids on the fictional pandora in Avatar), or the color of a polar bear. The environment is a critical […]

The post Are you a morning person? – Chronotypes, Circadian Rhythms, and Questionnaires appeared first on PROTrEIN.

]]>
Life as we know it has been shaped by the constraints of the environment. The aerodynamic shape of leaves to prevent the tree from toppling at high winds, or gravity that affects heights of organisms (extraterrestrial humanoids on the fictional pandora in Avatar), or the color of a polar bear. The environment is a critical factor that shapes life. For us humans, along with our living companions on earth, this includes revolving and rotating around a single star, with a single moon. This has enabled us to develop several rhythms, such as circadian cycle (24 hours), ultradian cycle (less than 24 hours), and infradian rhythms such as the human menstrual cycle have periods longer than a day. There are clocks within us, in every organ, every cell. At the cellular level, these are just chemical clocks that play with protein concentrations.

Circadian rhythm is not only important to understand better human pathologies or to find new potential therapeutic candidates, but may also help in personalized medicine. Knowing that the immune system and hepatic metabolism change throughout the day, it is possible to study at which moment a specific drug can be given to a specific patient in order to maximize its effect or reduce its toxicity. Unfortunately the inner clock’s timing can not be generalized, each individual has a different rhythm that depends both on the lifestyle and genetic predisposition [1]. Moreover the rhythm can be totally disrupted, causing each organ’s clock not to be in phase with the body’s one and/or with the day-night cycle. Several factors, especially now with humans’ modern lifestyle, may induce such disruption: examples are the jet-lag, jobs that require frequent night shifts or any kind of stress that precludes a person from having a constant resting-active state cycle [2].

One possible way to investigate a person’s circadian rhythm is through the use of questionnaires. They are used to categorize patients in chronotypes through questions about normal day habits like at what time the subject wakes up, has meals, goes to sleep and how they affect his/her life [3]. Through such data it is possible to determine if the person is an early bird or a night owl (which are called chronotypes) and with this information it is possible to say that, for example, the last type will have its active phase shifted towards the evening. As it is understandable there are several limitations: this approach must rely on what the subject says and do not allow to say with certainty if there is a circadian rhythm disruption.

In this month’s journal club, we decided to take a dive into the circadian rhythm, chronotypes, and the molecules that affect them. We also took the opportunity to develop a web app that helps you figure out your chronotype and what it means for your health, personality, etc. We turned to literature for connections between chronotypes and health, and we zeroed in on a paper by Fabbian et al., 2016 [5]. 

We decided to create an online questionnaire that with simple questions and a nice interface is able to predict the user’s chronotype. The idea of such a questionnaire is to find a way to make data gathering more interesting for those who have to respond to questions, translating a simple questionnaire into a game that in the end may give some information. Such information about chronotypes is given using simple terms in order to be understandable even without any scientific knowledge.

The questions used are from the Horne and Östberg Questionnaire [4] and information about the chronotypes were obtained from the paper F. Fabbian et al. 2016 [5]. There were many other papers that were referenced for this journal club, but these two were instrumental to informing the application design.

Chronotypes

There are several chronotypes that have been theorized by scientists, but for ease we decided to select only three. The selected ones are:

Early Birds/Lack. People having this chronotype tends to go to sleep early in the evening (between 9/10 PM) and wake up very early in the morning (7 AM or before). The peak of activity and energy is reached at the morning, while during the rest of the day they will become more and more tired.

Eagles. This chronotype is a mixture of the other two: peak of activity is reached after noon and they usually go to sleep between 10 and 11 PM and wake up between 8 and 9 AM.

Night Owls. Such chronotype describes people that usually stay awake at night until 12PM or later and wake up in the morning after 9 AM. The peak of activity is reached at the late afternoon, making these people feel less energetic when they wake up while becoming more active and efficient as the day progresses.

Circadian rhythm regulation

To understand better what a chronotype means it is necessary to explain before how the circadian rhythm is regulated, both at the systemic and molecular level.

At the systemic level the main center of regulation is situated in the brain, in particular is a region called the suprachiasmatic nucleus (SNC) in the amygdala. These neurons receive inputs from the whole brain and in particular from eyes and through them they are able to determine if it is day or night [6]. The SNC uses this information to regulate several body functions like animal’s behavior and psychology, body temperature, metabolism and body temperature. All such functions oscillate during the day and do not need inputs from the SNC to do so, although its main function is to reset the whole body circadian clock to ensure that all functions are in phase together and with the environment [7]. To ensure that each organ and tissue is synchronized, there are several genes embedded in the genome that regulate the rhythm in each cell at the molecular level. CLOCK and BMAL1 are the two core genes coding for the two homonym transcription factors, which form a complex that enhances transcription of several other genes [8].

Simple representation of how BMAL1-CLOCK regulate themselves through other transcriptional factors (Courtesy of Pickel L. et Sung H. K. 2020)

CLOCK-BMAL1 complex promotes CRY-PER factors production, which then acts on CLOCK and BMAL1 promoters suppressing their expression, suppressing then indirectly also their own expression and restarting the cycle again. This negative feedback loop is the main mechanism that allows CLOCK-BMAL1 expression to be cyclical with a period of around 24 hours [9]. Other genes promoted by CLOCK-BMAL1 complex are those coding for the nuclear receptors REV-ERB and ROR, which then regulate the expression of several genes involved in cell metabolism and other vital functions. Moreover, REV-ERBα acts also as a suppressor for BMAL1, while RORα enhances it, showing redundancy in the system that regulates core clock genes [10]. An interesting fact is those core clock genes are the same for each cell type, although they regulate a wide number of genes that are different for each tissue. This is in line with the fact that during specific phases of the day organs and tissues may be more or less active, so CLOCK-BMAL1’s effect can not be the same for each cell population. It is although not clear how this exactly happens: some studies found that CLOCK-BMAL1 activates different transcription factors for each cell type, while others suggest that clock genes can activate different promoters thanks to complexes with tissue-specific transcription factors [11, 12].

Considering what said so far, it is possible to reconsider each chronotype as the moment in which a person has the peak of CLOCK-BMAL1: for early birds it will be between late morning and noon while for late owls will be more shifted toward the afternoon. It is not clear although how much a chronotype is determined by genetic predisposition or environmental factor, probably a mixture of both. 

Several studies found strong correlation between night owls and detrimental behaviors, at the point that the chronotype can be considered a risk factor for several pathologies. One important thing to point out is that such correlation may be caused by other factors, like for example circadian rhythm disruption, since night owls’ life-style may conflict with modern world working hours and everyday life [13]. Moreover psychiatric disorders such as depression commonly cause difficulties in falling asleep and waking up early in the morning. This obviously does not mean that such people have an evening-type chronotype, even though through a questionnaire it may seem so.

References
1. Vitaterna MH, Takahashi JS, Turek FW. Overview of circadian rhythms. Alcohol Res Health. 2001;25(2):85-93. PMID: 11584554; PMCID: PMC6707128.
2. Eastman, C. I., Tomaka, V. A., & Crowley, S. J. (2016). Circadian rhythms of European and African-Americans after a large delay of sleep as in jet lag and night work. Scientific Reports, 6(1), 36716.
3. Zavada, A., Gordijn, M. C. M., Beersma, D. G. M., Daan, S., & Roenneberg, T. (2005). Comparison of the Munich Chronotype Questionnaire with the Horne‐Östberg’s Morningness‐Eveningness score. Chronobiology International, 22(2), 267–278.
4. Horne, J. A., & Östberg, O. (1976). A self-assessment questionnaire to determine morningness-eveningness in human circadian rhythms. In International Journal of Chronobiology (Vol. 4, pp. 97–110). Gordon and Breach Science Pub Ltd.
5. Fabbian, F., Zucchi, B., De Giorgi, A., Tiseo, R., Boari, B., Salmi, R., Cappadona, R., Gianesini, G., Bassi, E., Signani, F., Raparelli, V., Basili, S., & Manfredini, R. (2016). Chronotype, gender and general health. Chronobiology International, 33(7), 863–882.
6. Hastings, M. (1998). The brain, circadian rhythms, and clock genes. BMJ, 317(7174), 1704–1707.
7. Moore, R. Y. (2007). Suprachiasmatic nucleus in sleep–wake regulation. Sleep Medicine, 8, 27–33.
8. Trott, A. J., & Menet, J. S. (2018). Regulation of circadian clock transcriptional output by CLOCK:BMAL1. PLOS Genetics, 14(1), 1–34.
9. Yu, W., Nomura, M., & Ikeda, M. (2002). Interactivating Feedback Loops within the Mammalian Clock: BMAL1 Is Negatively Autoregulated and Upregulated by CRY1, CRY2, and PER2. Biochemical and Biophysical Research Communications, 290(3), 933–941.
10. Guillaumond F, Dardente H, Giguère V, Cermakian N. Differential Control of Bmal1 Circadian Transcription by REV-ERB and ROR Nuclear Receptors. Journal of Biological Rhythms. 2005;20(5):391-403.
11. Kondratov, R. V, Shamanna, R. K., Kondratova, A. A., Gorbacheva, V. Y., & Antoch, M. P. (2006). Dual role of the CLOCK/BMAL1 circadian complex in transcriptional regulation. The FASEB Journal, 20(3), 530–532.
12. Qu, M., Duffy, T., Hirota, T., & Kay, S. A. (2018). Nuclear receptor HNF4A transrepresses CLOCK:BMAL1 and modulates tissue-specific circadian networks. Proceedings of the National Academy of Sciences, 115(52), E12305–E12312.
13. Togo, F., Yoshizaki, T., & Komatsu, T. (2022). Interactive effects of job stressor and chronotype on depressive symptoms in day shift and rotating shift workers. Journal of Affective Disorders Reports, 9, 100352.

The post Are you a morning person? – Chronotypes, Circadian Rhythms, and Questionnaires appeared first on PROTrEIN.

]]>
Casanovo, a transformer model to identify De novo mass spectrometry peptide sequencing https://protrein.eu/blog/casanovo-a-transformer-model-to-identify-de-novo-mass-spectrometry-peptide-sequencing/ Fri, 10 Feb 2023 11:49:22 +0000 https://protrein.eu/?p=1740 In the last Journal club, we present a paper by Yilmaz et al. called «De novo mass spectrometry peptide sequencing with a transformer model» [1] introducing a deep learning model for de novo peptide sequencing. What? You do not know exactly what is de novo peptide sequencing? Let me explain it. Imagine that you do […]

The post Casanovo, a transformer model to identify De novo mass spectrometry peptide sequencing appeared first on PROTrEIN.

]]>
In the last Journal club, we present a paper by Yilmaz et al. called «De novo mass spectrometry peptide sequencing with a transformer model» [1] introducing a deep learning model for de novo peptide sequencing.


What? You do not know exactly what is de novo peptide sequencing? Let me explain it. Imagine that you do not have enough prior knowledge about your sample. How can you use database search methodology? Under this condition, we try to identify peptide sequences directly from experimental spectra. The principle of it is to find the specific fragmentation pattern based on the regular breaks in the mass spectrometric detection of the peptide molecules after protease cleavage and calculate the corresponding amino acid information according to the mass difference between the mass spectrum peaks as well as the post-translational modifications of the amino acid.

Figure 1: Casanovo performs de novo peptide sequencing. Source: Yilmaz et al., 2022 [1]

Early de novo methods used the heuristic search or dynamic programming to score peptide sequences. Recently, some efforts have been made to develop Deep learning models to predict peptides sequence from MS2 including DeepNovo, SMS, and PointNovo. However, these models include complex post-processing steps. In addition, their structures are based on recurrent neural networks, which are slow to train and suffer from long dependency issues.
To resolve mentioned drawbacks, they proposed a transformer-based model called Casanovo for de novo peptide sequencing. Casanovo uses the self-attention mechanism to translate from a variable-length sequence of observed spectrum peaks to a variable-length sequence of amino acids, analogous to the neural machine translation model in the natural language processing setting. o consists of a transformer encoder and decoder, where the encoder takes d-dimensional spectrum peak embeddings as input and outputs d-dimensional latent representation vectors.
Casanovo was trained by 30 million labeled spectra which contain Seven different types of variable modifications (methionine oxidation, asparagine deamidation, glutamine deamidation, N-terminal acetylation, N-terminal carbamylation, N-terminal NH3 loss, and the combination of N-terminal carbamylation and NH3 loss).

Figure 2: Casanovo architecture. Source: Yilmaz et al., 2022 [1]

To evaluate the performance of Casanovo and compare it with other de novo peptide sequencing models, they used the nine-species benchmark data set. This data set combines a total of about 1.5 million mass spectra from nine different experiments, each using the same instrument to analyze peptides from a different species.

Casanovo, leverages the transformer architecture to produce a unified solution to translate mass spectra directly into peptide sequences, without resorting to the discretization of the spectrum m/z axis and without complex post-processing.

We had an interview with Melih Yilmaz and asked him the two following questions about the paper:

Do you think what is the most challenging issue to have a better deep model for de novo sequencing?

Thanks for reaching out and for your interest in Casanovo! A challenge that we tried to overcome with our new preprint was that the original version of Casanovo was trained on MS data from peptides digested with trypsin enzyme which didn’t perform as well for samples that were digested using a different enzyme. To mitigate this, we fine-tuned the existing Casanovo model on a non-enzymatic data set which significantly improves performance on non-tryptic data.

Do you have any plan to improve your model by considering more modifications in your training data?

We don’t have short-term plans to increase the number of post-translational modifications in the current version model. However, it would be straightforward to fine-tune the current model with an extended training set containing the new modifications.

References:

[1] Yilmaz et al., bioRxiv (2022), doi.org/10.1101/2022.02.07.479481

The post Casanovo, a transformer model to identify De novo mass spectrometry peptide sequencing appeared first on PROTrEIN.

]]>
Ad Hoc Learning of Fragmentation https://protrein.eu/blog/ad-hoc-learning-of-fragmentation/ Tue, 10 Jan 2023 13:39:11 +0000 https://protrein.eu/?p=1722 In our last Journal Club blog post, we presented ProteomicsML a web platform with tutorials for machine learning in the field of proteomics. Today, we stay on the spot with machine learning, however, on this occasion, we are presenting a new approach and model from Tom Altenburg et al. based on their article “Ad hoc learning […]

The post Ad Hoc Learning of Fragmentation appeared first on PROTrEIN.

]]>
In our last Journal Club blog post, we presented ProteomicsML a web platform with tutorials for machine learning in the field of proteomics. Today, we stay on the spot with machine learning, however, on this occasion, we are presenting a new approach and model from Tom Altenburg et al. based on their article “Ad hoc learning of peptide fragmentation from mass spectra enables an interpretable detection of phosphorylated and cross-linked peptides“.

But what is ad hoc learning, what does this name stand for? “Ad Hoc” means ‘for this specific purpose‘ with which the authors would like to deliver the message that they developed a model that is learning from fragmentation for a specific purpose, in this case, phosphorylation detection. 

Current deep learning applications in the proteomics field are using peptide sequence information, and often masses of amino acids, ion types, losses, or combinations of the aforementioned parameters. Contrary, the model presented in this article abstracts fragmentation patterns from spectra that are important to recognize phosphorylated peptides based on their fragmentation spectra. The key is that the model is recognizing these essential patterns without being explicitly told about them.

To understand what the model is learning we have to introduce the concept of interpretability. Interpretability is the degree to which a human can understand the cause of a decision that the model made or the degree to which a human can consistently predict the model’s result. The authors used two methods to interpret the knowledge of their model. SHAP values (SHapley Additive exPlanations) were used to prove that the model’s decisions are based on peaks that belong to actual fragment ions rather than noise peaks. Also, PathExplain was used to compute pairwise interactions per spectrum to match those with relevant delta masses.

They proposed a two-vector representation, holding intensity and mass-over-charge (m/z) remainder information to encode spectra directly in the deep learning base model AHLF. The story behind such encoding is to promote the learning of associations between any peaks while respecting their location within a spectrum. Therefore, convolutional layers were used, because they preserve the location of a feature as the outputs from a convolution are equivariant concerning their inputs. Subsequently, the higher layers of deep neural networks can make use of the presence and the location of peaks. To be exact, they use convolutions with gaps, commonly called dilated convolutions. Ultimately, due to the parameter-sharing properties of convolutions, the total number of trainable weights is low compared to a fully connected network. Overall, the model architecture allows AHLF to use the entire two-vector spectrum as-is.

Figure 1. Illustration of how long-range associations can be learned by AHLF via dilated convolutions.

To demonstrate the ability of AHLFp in detecting spectra of phosphorylated peptides, they evaluate the performance of AHLFp on 19.2 million labeled spectra from 112 individual PRIDE repositories. In addition, They demonstrate the broad scope of their approach by applying AHLF to a distinct task, namely the detection of cross-linked peptides (AHLFx). Also to put this detection capability into practice, the model was utilized to rescore peptide matches using Percolator and the results were compared to PhoStar, a random forest model with carefully generated phospho detecting features.

Curious about the results? Check out the original article here!

As you might have seen in the blog post about ProteomicsML, we are now contacting the authors of publications covered in our Journal Club to make short interviews and gain further insights about these scientists’ work. Here we would like to thank Tom Altenburg and Bernhard Y. Renard for taking up the challenge and answering our questions! Below you can read our interview with the authors:

Blog post team: How does this research fit with your research interests and your institute’s research aims?

Tom Altenburg and Bernhard Y. Renard: In our group, we develop statistical and computational methods for high throughput techniques, including next-generation sequencing and MS-based proteomics.

Blog post team: What sparked the idea of using Ad hoc learning from peptide fragmentation?

Tom Altenburg and Bernhard Y. Renard: The data situation in MS-based proteomics is highly favorable for approaches like AHLF. There are two reasons for that: i) the availability of public proteomics data now is enormous and constantly growing and ii) there are methods that label data in an automatized fashion (i.e. peptide and protein identification) in MS-based proteomics. Specifically, proteomics search engines can identify and annotate peptides from MS data automatically. On the one hand, this makes training deep learning methods like AHLF feasible. On the other hand, current predominantly algorithmic approaches can further be improved by the integration of ML-based methods like AHLF.

Blog post team: One main novelty of your work is that the model is learning from fragment spectra without any sequence information and expert knowledge. What other fields of LC-MS could use a similar approach?

Tom Altenburg and Bernhard Y. Renard: Related approaches are emerging in lipidomics and metabolomics. For example, a group at TUM is currently working on detecting lipid species using deep learning in a way similar to AHLF. Furthermore, there are learning-based scoring functions that extend or outperform algorithmic scoring schemes in the case of metabolomics. However, from our perspective having a large pool of training data is key but may be not feasible in some other fields.

Blog post team: You show that pairwise interactions of respective delta masses coincide with expert knowledge about phosphopeptide fragmentation in a significant number of cases. Do you try to analyze cases you see in your pairwise analysis that are not explained with expert knowledge? Do you hope to draw/gain new knowledge from what AHLFp learned?

Tom Altenburg and Bernhard Y. Renard: To be honest, it was a bit of a surprise that for a large fraction of the most prominent pairwise interactions, the respective delta masses are relevant in the context of phosphoproteomics and we could explain them. However, we only considered losses and combinations thereof. One could take this a step further and match specific elemental compositions (i.e. compositions of C, H, N, O, and P). A comparison of compositions with or without phosphor may give good additional insights. This, in turn, can be extended to other elements or signatures of other modifications to gain new knowledge and thus is an interesting future perspective.

Blog post team: One of us (Arslan) works with protein-nucleic acid crosslink data analysis, and has some questions regarding the applicability of AHLF with this type of data:

As mentioned in the manuscript for protein-protein crosslink, AHLFx improves the results. For protein-nucleic acid crosslink, we can find spectra of a peptide with nucleic acid, where the crosslinker binds to nucleic acid (mass adducts). Upon fragmentation, the MS/MS spectra are more challenging e.g. we can find a,b,y ions, precursor ions, marker ions, and sometimes not a very nice shifted ion series. Do you have any suggestions on how we could adapt AHLF in protein-nucleic acid crosslinking protocols?

Tom Altenburg and Bernhard Y. Renard: The identification rate can be improved by using AHLF as an additional feature for rescoring or to pre-filter spectra to run a dedicated downstream analysis. Specifically, spectra that contain fragments from a cross-linker and a peptide may be treated differently from spectra that contain fragments from all three: cross-linker, a peptide, and a nucleic acid strand. For example, if a spectrum does not contain fragments from a nucleic acid strand they may be searched by a classical search engine or existing cross-linking search engines, such as xisearch. This at least gives an idea about which peptides may be cross-linked and only for those peptides and filtered (e.g. predicted by AHLFx) protein-nucleic acid cross-linked containing spectra a dedicated search needs to be performed.

Blog post team: The AHLF model is trained without crosslinking spectra, and as your results show, for protein-protein crosslinking it is fine to use transfer learning. Do you think the transfer learning approach could work for protein-nucleic acid crosslinking?

Tom Altenburg and Bernhard Y. Renard: The elemental composition of nucleotides differs from amino acids. Therefore, it might be a good entry point. At least, I would expect patterns (shifts or groups of peaks) that belong to the peptide, others that point to the DNA, and others that belong to the cross-linker. However, the data is probably still very limited in this area. The prediction performance (e.g., AHLFx) may vary with the instrument type and thus might be an additional constraint. In any case, transfer learning (initializing with a pre-trained AHLFp or AHLFx) is probably a good starting point for protein-nucleic acid crosslinking data.

Blog post team: How does the deep learning model behave if we add traditional information as predefined features (as did in Prosit)? It might affect the sensitivity of the identification of peptides.

Tom Altenburg and Bernhard Y. Renard: If a certain expected feature can be learned by the model, i.e. the architecture has no inherent bottleneck regarding that type of feature and if there is enough data for training – then it should not be necessary to include predefined features. However, if any of these two requirements is not met, it could help add features, e.g. to compensate for the lack of data. Otherwise, transfer learning goes in a similar direction. The model is trained on a domain with lots of data (learning ubiquitous and general features such as losses and delta masses etc). At that point does it make a difference if these features were pre-defined or rather pre-learned? Imagine we define a specific neutral loss, fixed conceptually and fixed numerically. The model must accept it and make it work, for better or worse. In contrast, if the model learned some neutral loss but its value is slightly off (w.r.t. the new domain), then it can adjust the neutral loss by adjusting the respective weight during transfer learning. Another intriguing example is the combinatorial complexity of internal fragments. Internal fragments occur when a fragment undergoes a second (or more) fragmentation event. Inevitably, the number of combinations of possible fragments (fragments outside the typical fragment ladder assumption) explodes. The basic building blocks (i.e. masses of amino acids and neutral losses) may be relatively easy to predefine but precalculating all possible internal fragments would be rather expensive. Luckily, this is where deep learning comes in very handy because if the complexity (i.e. combinatorics of internal fragments) follows some kind of hierarchy (i.e. fragments are subsets of each other) the hierarchical structure of a deep learning model might have a chance to pick this up and help us in this situation – provided that there was enough training data.

Blog post team: What are your plans with the model, if you have any?

Tom Altenburg and Bernhard Y. Renard: There are many interesting future directions and this is a fast pacing field. For example, we had the idea to extend AHLF to further improve phosphosite localization and therefore integrate the SHAP values in a way that helps us to pinpoint localization. In our paper, we could show that the FLR is not inflated by using AHLFx. However, one could take this a step further and integrate the SHAP values and build a dedicated localization tool based on this idea.

Again, we would like to thank Tom Altenburg and Bernhard Y. Renard for elaborating on their paper! Finally, thanks for reading our post, and keep tuned for further content!

The post Ad Hoc Learning of Fragmentation appeared first on PROTrEIN.

]]>
Recalling memories of our first in-person project meeting and SciComm training school in the beautiful city of colours ‘Barcelona’ https://protrein.eu/blog/recalling-memories-of-our-first-in-person-project-meeting-at-the-beautiful-city-of-colors-barcelona/ Mon, 12 Dec 2022 16:38:20 +0000 https://protrein.eu/?p=1689 Consortium meeting: After we ESRs had participated in the 13th International MaxQuant Summer School on Computational Mass Spectrometry-based Proteomics, the opening dinner of the consortium meeting was the first time that all PROTrEIN members, incl. supervisors, actually met in person. The venue was amazingly beautiful and close to the sea. Where we all enjoyed a […]

The post Recalling memories of our first in-person project meeting and SciComm training school in the beautiful city of colours ‘Barcelona’ appeared first on PROTrEIN.

]]>
Consortium meeting:

After we ESRs had participated in the 13th International MaxQuant Summer School on Computational Mass Spectrometry-based Proteomics, the opening dinner of the consortium meeting was the first time that all PROTrEIN members, incl. supervisors, actually met in person. The venue was amazingly beautiful and close to the sea. Where we all enjoyed a delicious dinner, with the added bonus of a beautiful night time sea view.

A stunning sea view with a full moon that ESRs enjoyed was next to our dinner location.

ESRs and supervisors attended the consortium meeting on September 12 and 13. The meeting featured presentations of the various work packages, keynote addresses, reports, and a meeting of the supervisory board. The «Management and Coordination» presentation by the PROTrEIN coordination team opened the work package presentations. An overview of previous sessions, including «scientific project planning,» midterm check, and summer school, was given at the beginning of this presentation. It went on to talk about internal reporting before citing the impending deadlines. Lastly, a little overview of Ghent Winter school. The nicest part of the meeting was the speed dating session when all ESRs had been given a chance to briefly introduce themselves to other project supervisors.
The ESRs were well-prepared the following day to provide an update on the status of their research work. Many of us ESRs were rather anxious, not only because of the supervisors, but also because of the «time machine» that was set up to inform us of the allocated time. Nevertheless, everything went smoothly, and we received many useful ideas from the other supervisors to enhance our direction.


«Novel machine learning predictors» was the first work package presented by five of the ESRs. The next six ESRs presented their works about «New algorithms for mass spectrometry raw data processing». This series of presentations was completed in the afternoon with «Integration and visualization of omics data» presented by the final four ESRs.

Sharing one of the memories of our consortium meeting when Shamil was giving updates on his project and while most of us seemed at ease as they finished, several of us were feeling anxious as we were waiting for our turn.

After a long day, Jonas had planned a surprise cooking workshop for us. It was a great team-building experience where we made tasty and colorful tapas. We were instructed by two very nice chefs, they divided us into two groups of warm and cold tapas, and of course, they didn’t forget the vegetarian options. The chefs first provided us with instructions and recipes, which we then heartily enjoyed. Our favorite tapas were the salmon and avocado tartar with pistachios mayo, and cod fritters with honey allioli. Even thinking about those fantastic tapas right now makes us want to go back and try everything over and over again.

The image serves as proof that the PROTrEIN team is capable of doing anything as a team, including cooking, in addition to being excellent researchers.
Sharing the memorable picture of the dinner in which it is clear  that Arthur was thoroughly enjoying his food and didn’t give a damn about the camera.

The second day of the consortium meeting started with keynote lecturer Juan Antonio’s presentation on open-research and data reusability. The day was followed by presenting the remaining three work packages and discussing the upcoming winter school. The consortium meeting ended with supervisory board and coordination meetings, and all ESRs went to ELISAVA to start their science communication course.

A group photo of the PROTrEIN team following a productive consortium meeting.

Science Communication workshop

Design thinking for scientists was the first module of the workshop on science communication, and when we were asked at the beginning of the course “What we are expecting from this course?”, most of us thought it would be some spoken practice for our project presentations and visuals. However, it turned out to be much more than that.

Blanca Guasch steps in to explain how to apply design techniques to make our research more fruitful and efficient.  Design thinking involves several steps, including Empathize, Define, Ideate, Prototype, Test, and Implement. We created individual collages to illustrate the design thinking module Empathize, describing ourselves imaginatively with the aid of printed images, and describing how we are contributing to the PROTrEIN project.

Metaphors used by the ESRs

Another interesting exercise that we ESRs had done is to put our names on the PROTrEIN WPn map according to the importance and relation of our contributions to each module of the design thinking process. Then, for the Ideate section, we discover the links between various ESRs and how they relate to one another in terms of project goals, secondments, and research methodology.

The ESRs created a WPn map to show the connections between their projects

With the help of different metaphors, ESRs described the work bundles and processes like explaining machine learning, cross-linking concepts, etc.  After expressing our analogies to the group, we took into consideration their responses. After this engaging session 1, we all had some constructive ideas to use in our daily lives.

In the picture Zoltan was trying to find metaphors for crosslinking

On the second day of the course on science communication, Blanca began the session with «Presentation designs for scientists,» another fascinating and crucial module for ESRs. This lesson focused on speaking well, taking the initiative, and being creative in our presentations. To keep things simple, avoid presenting too much information at once, emphasize the importance of consistency and typography, employ various styles according to the situation, and utilize a variety of formats and resources. 

Ane Guerra started a fascinating discussion on science communication storytelling and explaining stories. We completed some exercises as part of the story-telling process by describing the major difficulty we encountered when describing our project and the qualities that stand us apart from the others. Also, most of us ESRs defined our project as if we were a superhero, along with its foes, strengths, and nemesis. Being able to connect your project with some superpower heroes was, in my opinion, not an easy task, but most ESRs were able to accomplish it creatively. Then, we got some tips for developing story-telling skills in our research, including the use of a clear message, consideration of the audience, and the communication context. 

In the explaining stories session, we discovered that we should be aware of the audience when discussing our research and should stick to the core idea the entire time. 

“The best storytellers deliberately listen, watch, and read.” 

The session concludes with some helpful advice on how to improve our public speaking abilities, including knowing your audience, taking breaks, dressing appropriately, using our hands/voices effectively, and soliciting feedback.

ESRs along with their mentors were busy doing their group project

Another relevant and interesting session for us ESRs is data visualization which is to embrace scientific complexity and transform it into a sympathetic visual story that improves all of its best features. 

We then discussed the process of mapping facts to visual structures, known as visual encoding. After that we went through how people in ancient times shared information about data I-e, the 1945 Molecular model of penicillin by Dorothy, and also discussed some good books which are based on data and visualizations.

Visual design lecture by Francesc Ribot ,the coordinator of the Graphic Design Area

In the last section of the course, we ESRs demonstrated our creativity by writing a video script while taking the public and expert audiences into consideration. A script for each audience had to be written according to a set of rules and to last for 60 seconds. Furthermore, we ESRs had been given the opportunity to present our posters and received feedback and discussion from our fellows and mentors. 

Last but not least, Jonas Krebs, all the mentors and the funding body deserves praise for organising such a relevant and informative science communication course for us ESRs. From this course, we have learned that better communication abilities enable researchers to share their discoveries with a wider audience and strengthen links within their scientific groups.

The post Recalling memories of our first in-person project meeting and SciComm training school in the beautiful city of colours ‘Barcelona’ appeared first on PROTrEIN.

]]>
Wikipedia Hackathon Experience https://protrein.eu/blog/wikipedia-hackathon-experience/ Mon, 28 Nov 2022 11:36:35 +0000 http://protrein.eu/?p=1655 Most of us do research because we enjoy finding answers, solving problems and even helping people. But there is also a deep seated human tendency to leave a mark. Contribute to a bigger picture, graffiti on your neighbor’s wall, advance a field. This is probably why we write research papers. (Apart from begging for funding […]

The post Wikipedia Hackathon Experience appeared first on PROTrEIN.

]]>
Most of us do research because we enjoy finding answers, solving problems and even helping people. But there is also a deep seated human tendency to leave a mark. Contribute to a bigger picture, graffiti on your neighbor’s wall, advance a field. This is probably why we write research papers. (Apart from begging for funding and completing a PhD) 😛

But the general public cares less about research papers and is more interested in blog posts such as «Is your cow cheating on you?».

But then, there is a that one moment when you have to win a bet against your friend who thinks that «Cows can hear infrasonic sounds» and you need to open Wikipedia (*coughs* Microsoft Bing) to settle it. Or that time you had to write an assignment on «Cow hybrids”. Yep. One of the authors of this blog is obsessed with cows.

Anyways, Wikipedia to this day remains an important repository of high quality articles on almost any topic you can think of. To this end, we participated in a Wikipedia Hackathon that would facilitate addition of proteomics articles to the growing Wiki database.

So what is that we did on Wikipedia?

All PROTrEIN network early stage researchers participated in the wikipedia hackathon. We didn’t just learn about articles, but also approaches to publishing open access data to Wiki data.

Along with the mentor, Toni Hermoso, we created a web page in wikimedia entitled PROTrEIN Editathon (https://meta.wikimedia.org/wiki/PROTrEIN_Editathon_2022) we learned mainly how to edit into a wikipedia page starting from creating paragraph and subparagraphs into adding images into wiki images and then use it for our articles, we created multiple proposals for related terms and articles to the field of proteomics and mass spectrometry.

One article that we wanted to feature as an example is a Maxquant one. So as you see in the following figure, We created a wikimedia image for the software Maxquant, where we added the article for defining Maxquant alongside a figure for the software logo and viewer interface, we also added links and dates.

We learned how to create links, it works similar to hyperlink on the Microsoft Word processor. Alongside to citing, with automatic citation an editor can simply copy-paste the URL, and it will generate a citation. Whereas, manual citation requires the editor to enter the details of a book, journal, article, or website. After making edits, it is time to leave an edit summary about the changes that have been made to the wiki page.

This Wiki Hackathon took place within the framework of Computational Proteomics MaxQuant Summer School 2022 on September 5-6, 2022.

The post Wikipedia Hackathon Experience appeared first on PROTrEIN.

]]>
Organising a large event as a first year PhD student – behind the scenes of the MaxQuant Summer School 2022 https://protrein.eu/blog/organising-a-large-event-as-a-first-year-phd-student-behind-the-scenes-of-the-maxquant-summer-school-2022/ Mon, 21 Nov 2022 15:46:13 +0000 http://protrein.eu/?p=1601 Starting with the premise that I have never organised such a big event in the past, getting actively involved in the organisation of the MaxQuant summer school (MQSS) has been a great opportunity to learn skills that are usually not directly related to pure research in science. The whole process of organising the summer school […]

The post Organising a large event as a first year PhD student – behind the scenes of the MaxQuant Summer School 2022 appeared first on PROTrEIN.

]]>
Starting with the premise that I have never organised such a big event in the past, getting actively involved in the organisation of the MaxQuant summer school (MQSS) has been a great opportunity to learn skills that are usually not directly related to pure research in science.

The whole process of organising the summer school can be summarised in four different steps:
1) finding the venue to host the event
2) advertising the event
3) registration and booking
4) setting up the event

Finding the venue to host the event

In January this year, I joined the summer school organisation committee, which was composed of my supervisor Jürgen Cox, one post-doc and another more experienced PhD student that has already organised summer schools in previous years. When we discussed possible locations of the event, it was immediately decided to go to Barcelona, not only because of the sun, but also due to the logistic easiness (the MQSS took place there already once in 2018). Besides, Jürgen’s lab has contacts there (Eduard and Jonas from CRG) that were of great help. The initial phase consisted mostly of meetings with people from the agency that helped us with the organisation (Crea Congresos) and to decide when and where to have it. Once a venue that fulfilled the main requisites was found the real organisation from our side started and I got more involved. Such main criteria for the decision were: enough space for 200 participants and their posters, audio and video appliances and technical support and lastly being easy to reach. In this part most of the work, like visiting different venues and being sure of which services were ensured, was done by our counterpart in Barcelona. We were lucky that with their help we could convince the «Centre de Cultura Contemporània de Barcelona» (CCCB) to host the event one more time, same as in 20218.

Advertising the event

The first step was to design a nice logo and then set up the webpage. These were my initial tasks. I took this opportunity also to practice my skills in html and Adobe Illustrator. In particular, designing the new logo was something I enjoyed a lot, since I have always drawn and liked art. In other words, it was a great way to apply art to science.

The final logo of the summer school

Once the drafts were approved and the website was set up, it was time to start advertising the summer school through social media (e.g. Cox lab’s and Max Planck’s twitter account) and through sending emails to all previous years’ participants. In the meanwhile, the five main speakers were found and a first conference program drafted, both done in collaborative efforts. Lastly, Crea Congresos prepared the registration form and opened the bank account where to receive the conference fees from participants.

Registration and booking

As soon as we had set the registration deadline on mid May, the priority went to prepare the budget. All expenses were calculated by Crea Congresos, we checked them and added a buffer for eventual last moment expenses (an issue that indeed happened). On top of that, taxes were added and then 170 was set as the minimum required number of participants in order to cover all expenses. The registration fee for each person was calculated accordingly and the registration form opened. Assisting with the budget report was something totally new to me, but also helpful to get an idea of the costs that such event produces. Honestly, I have never dealt before with such an amount of money in my life. From this point on, we spend most of the time with taking care of public relations, since there was an average number of 8-10 emails per day from people asking for more information. The date to close the registration was at the end postponed to August. After that, we could proceed with booking the restaurant and social activities. Once we reached the second half of August, all participants that registered were contacted again in order to know if they wanted to present a poster and which social activity they wanted to join.

The week before the event, the spreadsheets with all this information were sent to Crea Congresos to let them know about the exact numbers. We organised all necessary documents and data (like MaxQuant, Perseus, tutorials and practical exercises) and transferred them to individual USB pen drives that we planned to give to each participant. Lastly, a new Zoom subscription was purchased in order to allow an online participation of the summer school and a new registration form for online participants was opened.

Checking that everything works before the start

Setting up the event

Just few days before the start of the summer school, it was time to move to Barcelona to set up the venue and test appliances to be sure everything was ready for Monday. From that moment on, luck was not very much on our side. One small issue was that there has been some delays in the preparation of the USB sticks that we planned to give to attendants, but that was solved easily through all collaborative efforts during the weekend before. The bigger issue was a strike that caused flights to Barcelona to be cancelled, all except mine. This has been quite a big deal since on the Friday before the start we were supposed to visit the venue and test all the equipment. At the end we managed to solve everything thanks to the help of the people from Barcelona and the technicians, and all was set and ready for the event to start. The flight issue was exactly one thing we were worried about, especially because it was something we could not do much about it. At the end, it was a good reminder of Murphy’s law: “anything that could go wrong will go wrong”. Lastly during the weekend also the rest of the lab managed to arrive and we all prepared the material and rehearse our talks.

Finally it was time for the summer school to begin and about this I will not write much more since everyone from the PROTrEIN-ITN was there. The final program can be found on the MQSS webpage.

Hamid enjoying not being an organizer for one time

Honestly, there has been also a particular moment during which I cursed being one of the organizers and it was during the joint dinner on Wednesday. That day we kept an easy schedule for participants: social activities like tours or sport in the afternoon and then a dinner all together in a nice restaurant close to the sea. Unfortunately, my talk was exactly the next morning and there were some more data needed by participants that we could not give through the usb sticks. This meant that during the free afternoon I had to prepare my talk and then had to leave the dinner earlier in order to prepare the email that was sent to all participants with the aforementioned data. In other words I missed a nice evening and partying while some of my colleagues did not. Have a look at the picture above to get what I mean, for the record I was writing emails at that moment. Well, at least I did not have a hangover the day after.

In summary, it has been a nice experience, even though doing everything took a lot of time and there has been some stressful moments in particular during the last days. It was also a good opportunity to meet other researchers working in proteomics and to get a feedback from them regarding our general work. Lastly, when the event was over, it was also a great satisfaction to have been part of the organisation and seeing that participants appreciated it and had a good time during the week.

Celebrating the event to be over

The post Organising a large event as a first year PhD student – behind the scenes of the MaxQuant Summer School 2022 appeared first on PROTrEIN.

]]>
New to machine learning in proteomics? Check out the ‘ProteomicsML’ web platform https://protrein.eu/blog/new-to-machine-learning-in-proteomics-check-out-the-proteomicsml-web-platform/ Wed, 16 Nov 2022 15:06:27 +0000 http://protrein.eu/?p=1589 Machine learning approaches have become an established part of the mass spectrometry-based proteomics field in recent years. Several tools capable of predicting different aspects of peptide behavior have been developed and incorporated in data analysis workflows. These tools have proven to be beneficial in peptide and protein identification in proteomics experiments, and it is therefore […]

The post New to machine learning in proteomics? Check out the ‘ProteomicsML’ web platform appeared first on PROTrEIN.

]]>
Machine learning approaches have become an established part of the mass spectrometry-based proteomics field in recent years. Several tools capable of predicting different aspects of peptide behavior have been developed and incorporated in data analysis workflows. These tools have proven to be beneficial in peptide and protein identification in proteomics experiments, and it is therefore of great interest to the people in the community to utilize these tools. It is of high importance to have datasets that are suitable for training and evaluation of machine learning models. However, it is often not a trivial task to prepare the data in a format that is compatible with the different software. At the moment there are different datasets available and often used for training on ProteomeTools, but there is not a formal consensus on which datasets to use for training and evaluation.

This is highlighted by Rehfeldt et al. in their paper “ProteomicsML: An Online Platform for Community-Curated Datasets and Tutorials for Machine Learning in Proteomics”, which is currently available as a preprint on ChemRxiv. As a result of the 2022 Lorentz Center Workshop on Proteomics and Machine learning (Neely et al. 2022 (submitted for review)), the authors have developed a web platform that they hope will bring together the people wanting to use machine learning tools in proteomics with the people developing them.

The authors have tried to overcome machine learning suitable training and evaluation barriers by providing access to datasets that are pre-processed and ready for applying machine learning models. On the ProteomicsML platform you can find tutorials on how to prepare your own data for the different state of the art machine learning models. Tutorials are available for four different data types representing different predictive capabilities in proteomics: retention time, fragment ion intensities, ion mobility and detectability. The platform also contains different datasets in the same four categories with varying complexity that you can download and explore on your own. If you have questions regarding any of the tutorials, you can ask a question in the Tutorials Q&A.

Besides providing tutorials and datasets, the authors also encourage people in the community to contribute themselves either by posting on one of the discussion boards or by contributing with their own dataset. If you want to contribute with datasets or tutorials there is also a very thorough guide that includes their code of conduct.

We think this is a great addition to the proteomics community as many of us PROTrEIN ESRs were newcomers to this field when we started our projects last year. A platform like this would have been a great starting point for us, and we hope that it will benefit many other people in the community! You can check out the ProteomicsML platform here https://proteomicsml.org/

Also, we could contact one of the authors of the paper, Ralf Gabriels, and interview him with the following questions:

Blog post team: In the paper you nicely describe why you think a platform like this would be useful to the proteomics community, but did you get direct feedback from people wanting to use the state-of-the-art machine learning tools or did you identify this challenge with different file formats and file complexity yourselves?

Ralf Gabriels: I think we mostly experienced these hurdles ourselves throughout our PhDs, and still do, of course. Dynamic and accessible educational resources are essential in a fast-growing, but complex field such as proteomics.

Blog post team: We know that the platform is still very new and in its start-up phase, but have you already had discussions about how/in which direction the content on the platform will grow?

Ralf Gabriels: Not too many discussions yet, but we do want to keep it up to date with developments in the field. Other than that,  we are mostly looking towards the community and users to give feedback and feature requests on GitHub.

Blog post team: In which direction do you think that machine learning in proteomics will go over the next years? 

Ralf Gabriels: I am certain that machine learning will continuously be more embedded in how we analyze the complex data that mass spectrometers produce. The better we understand how peptides behave in a mass spectrometer, the better we will be in interpreting the resulting data, and the better we will be at confidently identifying peptides (and thus proteins). Moreover, not only new instrumentation will drive the development of novel machine learning approaches, advancements in machine learning for proteomics will be able to optimize the development of the instruments, essentially opening a positive feedback loop between wet-lab and dry-lab methodological innovation.

Thanks a lot to Ralf Gabriels for his answers and for taking the time to answer our questions! 

References

  1. Rehfeldt T, Gabriels R, Bouwmeester R, Gessulat S, Neely B, Palmblad M, et al. ProteomicsML: An Online Platform for Community-Curated Datasets and Tutorials for Machine Learning in Proteomics. ChemRxiv. Cambridge: Cambridge Open Engage; 2022;  This content is a preprint and has not been peer-reviewed.
  2. Zolg DP, Wilhelm M, Schnatbaum K, Zerweck J, Knaute T, Delanghe B, Bailey DJ, Gessulat S, Ehrlich HC, Weininger M, Yu P, Schlegl J, Kramer K, Schmidt T, Kusebauch U, Deutsch EW, Aebersold R, Moritz RL, Wenschuh H, Moehring T, Aiche S, Huhmer A, Reimer U, Kuster B. Building ProteomeTools based on a complete synthetic human proteome. Nat Methods. 2017 Mar;14(3):259-262. doi: 10.1038/nmeth.4153. Epub 2017 Jan 30. PMID: 28135259; PMCID: PMC5868332.

The post New to machine learning in proteomics? Check out the ‘ProteomicsML’ web platform appeared first on PROTrEIN.

]]>
“Cross-Linking Mass Spectrometry: A sneak peak into its world!! ” https://protrein.eu/blog/cross-linking-mass-spectrometry-a-sneak-peak-into-its-world/ Sun, 26 Jun 2022 15:50:04 +0000 http://protrein.eu/?p=1554 It has been a while since our official PROTrEIN launch. We’ve had monthly ESR meetings during this time, where the majority of us have been able to discuss more details about our projects and present them to the group. Several ESRs are working on developing new tools and data analysis pipelines to improve cross linking […]

The post “Cross-Linking Mass Spectrometry: A sneak peak into its world!! ” appeared first on PROTrEIN.

]]>
It has been a while since our official PROTrEIN launch. We’ve had monthly ESR meetings during this time, where the majority of us have been able to discuss more details about our projects and present them to the group. Several ESRs are working on developing new tools and data analysis pipelines to improve cross linking mass spectrometer data analysis. It is worthy to note that these ESRs are supervised by great supervisors who’ve already made significant contributions to the field of cross linking.
For this month’s journal club, we therefore chose the review paper «Cross-linking mass spectrometry: methods and applications in structural, molecular, and systems biology» by F. O’Reilly and J. Rappsilber[1].
Cross-linking mass spectrometry (CLMS) has emerged as a useful technique in structural biology research, complementing traditional approaches such as x-ray crystallography and electron microscopy. CLMS allows researchers to investigate proteins in solution, capturing them in a dynamic state that is closer to physiological settings. CLMS also allows for the analysis of heterogeneous samples containing compounds in low quantities, showcasing the benefits of CLMS.

Overview of CLMS:

In the cross-linking reaction, covalent bonds are formed between the reactive groups of the cross-linker and surface residues of proteins, peptides and/or nucleic acids (DNA and RNA).
This way, residues that are within a certain reach of each other can be linked, thereby providing information about tertiary structure as well as interactions between protein complexes and/or nucleic acids.
Chemical cross-linkers are molecules with a spacer region flanked by reactive end groups that are extremely specific, while others are not. The chemical properties of the reactive groups and the length of the spacer region establish the limits of a cross-linker and therefore can be designed to accommodate different CLMS workflows. CLMS is not limited to studying the structure and interactions between proteins but can also be performed to gain knowledge about interactions between proteins and nucleic acids, however this blog post focuses specifically on the application and workflows involving proteins and peptides.

The general CLMS workflow is shown in Figure 1.

Figure 1. General cross-linking mass spectrometry (CLMS) workflow. (a) First step is choosing the correct cross-linker for the experiment. Depending on the question you want answered and the workflow, the cross-linker may need to be cleavable in the mass spectrometer, be isotopically labeled or have properties that allow for enrichment. This will be described in further detail in the CLMS workflow section. Once the cross-linker has been added to the sample, inter- and intra-protein crosslinks are formed (b). After the cross-linking reactions, the proteins in the sample are digested by a protease (c) yielding a mix between cross-linked and linear peptides. In some workflows, the cross-linked peptides are enriched (d) before data acquisition by MS/MS (e). The last step is data analysis that aims at identifying cross-linked peptides (f).

CLMS Applications:

The review paper we are discussing focuses on four overall applications of CLMS in protein studies, as illustrated in Figure 2, however CLMS has many more applications.

Modeling of protein complex topology is one of the most common uses of CLMS to investigate how proteins are ordered with respect to one another as they form complexes, a process known as protein complex topology. In circumstances where the complex topology is unknown, information concerning distance limits between surface residues of proteins that are known to be complex can be used with other structural research tools to aid modeling of the complex topology. 

Tertiary protein structure modeling can benefit from High Density (HD) CLMS. HD-CLMS data is acquired by using a cross-linker that is semi specific; one of the reactive groups only binds to specific residues, whereas the other group has no binding restrictions. This type of cross-linker will create a highly dense mapping of distance restraints between residues on the protein surface. Using this information can help exclude certain arrangements of the secondary structure elements of the proteins, thereby supporting modeling of the tertiary structure. Using a semi-specific cross-linker makes the data highly complex due to the large number of crosslinks formed in the reaction. 

Quantitative CLMS based comparative studies can be utilized to study proteins in different conformations. The relative abundance of distinct cross-links generated in each sample can be determined by adding isotopically labeled cross-linkers to samples from different experimental conditions. This information can help detect whether a protein is predominantly in one conformation or another, depending on the conditions of the experiment. These comparative analyses work best for proteins that go through conformational changes that strongly affect the structure since it affects the amount of cross-links that can be produced.

Proteome-wide CLMS studies focus on studying protein-protein interactions (PPIs) in large scale However, due to the enormous number of possible cross-links that might form between peptides, the data from these tests is exceedingly complex, posing some issues. Proteome-wide PPI research can be done in a variety of ways. Targeted pulldown procedures, in which natural protein complexes are identified and examined; cell lysate analysis, in which PPIs in the soluble proteome are explored; and in situ studies of complete cells or organelles are just a few examples.

Figure 2. Four situations where CLMS can be applied and aid modeling of protein-protein interactions and protein structure. (a) studying topology of protein complexes. (b) tertiary structure of single proteins. (c) comparative studies using quantitative CLMS to study protein conformation and (d) proteome-wide studies of protein topology.

CLMS workflows:

Although the overwhelming number of workflows available can be perplexing for newcomers in this field the development of standardized reagents and workflows has significantly boosted the simplicity to use CLMS. For the detection of cross-linked peptides, a number of software solutions are now available. The typical method for gauging confidence, regardless of the search program employed, is to utilize a target-decoy strategy to estimate the false discovery rate (FDR).
Emerging reporting standards and data-visualization tools are facilitating this technique’s accessibility, which are discussed below briefly:

❖ Reporting standards in CLMS: Because this field hasn’t publicly agreed on minimal reporting standards, it’s difficult to evaluate papers and reuse data. “mzIdentML” (http://www.psidev.info/mzidentml/) is an XML-based reporting standard for proteomics data developed by the Human Proteomics Standards Initiative (HUPO-PSI), which includes CLMS.
Raw spectrometric data should be deposited in certain public repository after publication.There is a need for clarification when reporting results when the word ‘cross-link’ is used interchangeably for peptide spectral matching (PSMs), peptide pairings, and residue pairs, because the defined FDR at the PSM or peptide level results in an unknown and typically much larger FDR at the level of residue pair.

❖ Data Visualization and interpretation: Software for visualizing discovered cross-links and the mass spectra that lead to their identification has been developed by numerous laboratories to make CLMS data accessible. Many levels of information is provided by cross-linking studies including:
A. Residue–Residue links
B. 3D structural information
C. Protein–Protein interactions

Figure 3 shows that their combination is one-of-a-kind, necessitating custom visualization.

Figure 3. Visualization solutions for CLMS data. (a) Spectra identified as cross-linked peptides can be manually assessed (b) Cross-linked proteins can be visualized with node and edge graphs to display interconnectivity of proteins (c) Mapping of cross-links on known 3D structures or homology models can score and validate cross-links and show those that violate the distance restraints.

“Notably, the cross-linker spacer’s chemistry can be tweaked, enabling data analysis simpler and boosting confidence in the cross-links found. As a result, before starting a study, it’s important to think about the best cross-linker to be used in conjunction with the analysis pipeline.”

Now we’ll look at some of the most prominent methods for analyzing CLMS.

Universal approach: Most comprehensive method, does not necessitate changing the cross-linker spacer in order to perform downstream analysis, commonly employed in conjunction with commercial cross-linkers, and effective for cross-linkers that can’t be modified in the spacer region , like photo amino acids. Isotope labeling is not crucial for identification and can be employed in quantitative or comparative studies. Using modern mass spectrometers, MS/MS spectra can be recorded at high resolution, which reduces the chance of getting false positive hits in the identification. StavroX[2], Xlink-Identifier[3], and Xcomb[4] generate a database of potentially cross-linked peptide pairs, but as the number of proteins grows, their computational time increases. Modification search combined with experimental heuristics that computationally enrich possible cross-linked peptides, save search time before scoring the spectra in Xi[5], Plink[6], XLSearch[7], Protein Prospector82[8], ECL2[9] , and Kojak[10].

Labeled cross-linker approach: Samples are treated with a mix of a heavy-isotope-labeled cross-linker and its unlabeled equivalent. Cross-linked peptides can then be identified in the MS1 spectra by searching for doublet peaks that are displaced by the mass of the heavy isotopes. MS2 spectra of the ‘light’ (unlabelled cross-linker) and ‘heavy’ (isotopically labeled cross-linker) precursors can reveal which fragment ions can contain the cross-linker, for confident cross-link identification, Hekate[11], StavroX, and the widely used xQuest[12] are just a few examples of search tools that use this approach. Talking about its positive side, this method streamlines data-analysis operations and can even be useful where high-accuracy mass spectrometers are not accessible, but it also increases the complexity of the MS1 spectrum space, potentially lowering recognition rates. Furthermore, requiring both heavy and light precursors for fragmentation can cause problems in complex samples.

MS2-cleavable cross-linker approach: utilizes cross-linkers that are cleavable during MS2 fragmentation, resulting in two peptides per MS2 spectrum that can be observed as unique cross-link-specific fragment ions. As cross-linked peptides are vast and branched, their fragmentation spectra are complex and uneven. The vast number of potential peptide combinations, combined with the frequently poor fragmentation of one of the cross-linked peptides, can make identifying the two peptides challenging, but this can be made easier by separating the two peptides in the mass spectrometer. This technique employs longer duty cycles than MS2-only approaches and requires additionally to execute MS3. Acquisition approaches for these cross-linkers have been designed by several laboratories along with their respective search software, such as ICC-CLASS[13], MeroX[14], X-links/Blinks[15,16] and XlinkX2.0[17,18]. After the review was published in 2018, an additional cross-linking search engine, MS Annika[19] was published in 2021


Figure 4 CLMS data acquisition and analysis workflows.(a) The ‘universal approach’ uses cross-linkers with simple spacers (b) Labeled cross-linker approach using isotopically labeled cross-linkers. (c) Cross-linker approach that uses cleavable cross-linkers in MS2 fragmentation.

Conclusion:

We hope you now understand why CLMS is such a powerful tool for examining protein interactions and topology. CLMS is a blooming field, with new processes and cross-linkers being developed as well as data analysis. Advances in data acquisition should be accompanied by improvements in data analysis. PROTrEIN ESRs, as well as other community members, are working on new tools and analysis pipelines to empower researchers to make new discoveries.

What do you anticipate CLMS will provide next?

References:

  1. O’Reilly, F.J., Rappsilber, J. Cross-linking mass spectrometry: methods and applications in structural, molecular and systems biology. Nat Struct Mol Biol 25, 1000–1008 (2018). https://doi.org/10.1038/s41594-018-0147-0
  2. Götze M, Pettelkau J, Schaks S, Bosse K, Ihling CH, Krauth F, Fritzsche R, Kühn U, Sinz A. StavroX–a software for analyzing crosslinked products in protein interaction studies. J Am Soc Mass Spectrom. 2012 Jan;23(1):76-87. doi: 10.1007/s13361-011-0261-2. Epub 2011 Oct 25. PMID: 22038510.
  3. Du X, Chowdhury SM, Manes NP, Wu S, Mayer MU, Adkins JN, Anderson GA, Smith RD. Xlink-identifier: an automated data analysis platform for confident identifications of chemically cross-linked peptides using tandem mass spectrometry. J Proteome Res. 2011 Mar 4;10(3):923-31. doi: 10.1021/pr100848a. Epub 2011 Feb 16. PMID: 21175198; PMCID: PMC3048902.
  4. Panchaud A, Singh P, Shaffer SA, Goodlett DR. xComb: a cross-linked peptide database approach to protein-protein interaction analysis. J Proteome Res. 2010 May 7;9(5):2508-15. doi: 10.1021/pr9011816. PMID: 20302351; PMCID: PMC2884221.
  5. Giese SH, Fischer L, Rappsilber J. A Study into the Collision-induced Dissociation (CID) Behavior of Cross-Linked Peptides. Mol Cell Proteomics. 2016 Mar;15(3):1094-104. doi: 10.1074/mcp.M115.049296. Epub 2015 Dec 30. PMID: 26719564; PMCID: PMC4813691.
  6. Yang B, Wu YJ, Zhu M, Fan SB, Lin J, Zhang K, Li S, Chi H, Li YX, Chen HF, Luo SK, Ding YH, Wang LH, Hao Z, Xiu LY, Chen S, Ye K, He SM, Dong MQ. Identification of cross-linked peptides from complex samples. Nat Methods. 2012 Sep;9(9):904-6. doi: 10.1038/nmeth.2099. Epub 2012 Jul 8. PMID: 22772728.
  7. Ji C, Li S, Reilly JP, Radivojac P, Tang H. XLSearch: a Probabilistic Database Search Algorithm for Identifying Cross-Linked Peptides. J Proteome Res. 2016 Jun 3;15(6):1830-41. doi: 10.1021/acs.jproteome.6b00004. Epub 2016 May 6. PMID: 27068484; PMCID: PMC5770149.
  8. Trnka MJ, Baker PR, Robinson PJ, Burlingame AL, Chalkley RJ. Matching cross-linked peptide spectra: only as good as the worse identification. Mol Cell Proteomics. 2014 Feb;13(2):420-34. doi: 10.1074/mcp.M113.034009. Epub 2013 Dec 12. PMID: 24335475; PMCID: PMC3916644.
  9. Yu F, Li N, Yu W. Exhaustively Identifying Cross-Linked Peptides with a Linear Computational Complexity. J Proteome Res. 2017 Oct 6;16(10):3942-3952. doi: 10.1021/acs.jproteome.7b00338. Epub 2017 Sep 1. PMID: 28825304. 
  10. Hoopmann MR, Zelter A, Johnson RS, Riffle M, MacCoss MJ, Davis TN, Moritz RL. Kojak: efficient analysis of chemically cross-linked protein complexes. J Proteome Res. 2015 May 1;14(5):2190-8. doi: 10.1021/pr501321h. Epub 2015 Apr 15. PMID: 25812159; PMCID: PMC4428575.
  11. Holding AN, Lamers MH, Stephens E, Skehel JM. Hekate: software suite for the mass spectrometric analysis and three-dimensional visualization of cross-linked protein samples. J Proteome Res. 2013 Dec 6;12(12):5923-33. doi: 10.1021/pr4003867. Epub 2013 Oct 4. PMID: 24010795; PMCID: PMC3859183.
  12. Rinner O, Seebacher J, Walzthoeni T, Mueller LN, Beck M, Schmidt A, Mueller M, Aebersold R. Identification of cross-linked peptides from large sequence databases. Nat Methods. 2008 Apr;5(4):315-8. doi: 10.1038/nmeth.1192. Epub 2008 Mar 9. Erratum in: Nat Methods. 2008 Aug;5(8):748. PMID: 18327264; PMCID: PMC2719781.
  13.  Petrotchenko, E.V., Borchers, C.H. ICC-CLASS: isotopically-coded cleavable crosslinking analysis software suite. BMC Bioinformatics 11, 64 (2010). https://doi.org/10.1186/1471-2105-11-64
  14. Götze M, Pettelkau J, Fritzsche R, Ihling CH, Schäfer M, Sinz A. Automated assignment of MS/MS cleavable cross-links in protein 3D-structure analysis. J Am Soc Mass Spectrom. 2015 Jan;26(1):83-97. doi: 10.1007/s13361-014-1001-1. Epub 2014 Sep 27. PMID: 25261217.
  15. Hoopmann MR, Weisbrod CR, Bruce JE. Improved strategies for rapid identification of chemically cross-linked peptides using protein interaction reporter technology. J Proteome Res. 2010 Dec 3;9(12):6323-33. doi: 10.1021/pr100572u. Epub 2010 Nov 10. PMID: 20886857; PMCID: PMC3018735. 
  16. Anderson GA, Tolic N, Tang X, Zheng C, Bruce JE. Informatics strategies for large-scale novel cross-linking analysis. J Proteome Res. 2007 Sep;6(9):3412-21. doi: 10.1021/pr070035z. Epub 2007 Aug 3. PMID: 17676784; PMCID: PMC2475505.
  17. Liu F, Lössl P, Scheltema R, Viner R, Heck AJR. Optimized fragmentation schemes and data analysis strategies for proteome-wide cross-link identification. Nat Commun. 2017 May 19;8:15473. doi: 10.1038/ncomms15473. PMID: 28524877; PMCID: PMC5454533. 
  18. Liu F, Rijkers DT, Post H, Heck AJ. Proteome-wide profiling of protein assemblies by cross-linking mass spectrometry. Nat Methods. 2015 Dec;12(12):1179-84. doi: 10.1038/nmeth.3603. Epub 2015 Sep 28. PMID: 26414014. 
  19. Pirklbauer GJ, Stieger CE, Matzinger M, Winkler S, Mechtler K, Dorfer V. MS Annika: A New Cross-Linking Search Engine. J Proteome Res. 2021 May 7;20(5):2560-2569. doi: 10.1021/acs.jproteome.0c01000. Epub 2021 Apr 14. PMID: 33852321; PMCID: PMC8155564.

Source of gallery picture of this blog-post: Juan Gaertner/Science Photo Library/Getty Images

The post “Cross-Linking Mass Spectrometry: A sneak peak into its world!! ” appeared first on PROTrEIN.

]]>
A Systematic Review on Bioethics in Proteomics https://protrein.eu/blog/a-systematic-review-on-bioethics-in-proteomics/ https://protrein.eu/blog/a-systematic-review-on-bioethics-in-proteomics/#comments Fri, 20 May 2022 14:03:45 +0000 http://chromdesign.crg.eu/?p=1 The moment to begin exercising control over the rules and regulations that will bind us tomorrow is soon. The time to begin thinking and talking about them is now. In this month’s Journal Club, we have decided to step away from technical topics such the latest advances in proteomics methodologies. Instead, we will approach a more “philosophic” […]

The post A Systematic Review on Bioethics in Proteomics appeared first on PROTrEIN.

]]>
The moment to begin exercising control over the rules and regulations that will bind us tomorrow is soon. The time to begin thinking and talking about them is now.

In this month’s Journal Club, we have decided to step away from technical topics such the latest advances in proteomics methodologies. Instead, we will approach a more “philosophic” topic, that of bioethical principles in the context of (clinical) proteomics. It may not be a topic that is as adrenaline-inducing as the latest and coolest scientific discoveries of our days. None the less important, bioethics is a highly relevant subject that concerns all scientists in their duty to conduct socially responsible research (Resnik & Elliott, 2016). The gallery picture for this post is taken from the cover of the American Journal of Bioethics, December 2021, Volume 21, Number 12.

Introduction

In the article we will present here, entitled “Ethical Principles, Constraints, and Opportunities in Clinical Proteomics” (Mann et al., 2021), the authors introduce the notions of bioethical principles and how they relate to proteomics research. They then present the results of the systematic review of the literature they performed on the principles of bioethics in clinical proteomics and finally, they provide their own perspective on some of these topics. 

Info box: Systematic review
Systematic reviews are reviews that aim to answer a clearly defined question. To this end, they use systematic and reproducible methods to identify, select and appraise all relevant research on the topic of interest. Further, they collect and analyze the data presented in the studies that are included in the review. As systematic reviews strictly adhere to scientific design, there are strict and well-defined guidelines to frame their methodology, such as the PRISMA guidelines (Moher et al., 2009).

Evidently, ethical concerns and responsibilities are not novel in biological sciences. Such issues have often been addressed and taken into account in the past. The field of genomics with its vast expansion during the last two decades constitutes an excellent example where bioethical questions have been extensively addressed and good practices established (for example see Kalia et al., 2017). During the recent advent of proteomics technologies and while the field has focused on developing its foundations, bioethical issues have so far been neglected, but it is now time to consider and discuss them in response to the increased capabilities of proteome analysis. To support this very argument, Mann et al. refer to their work on plasma proteomes highlighting how they can be reindentified based on protein expression levels or variant peptides, and published in an accompanying article (Geyer et al., 2021).

Core Principles

In the first section of their review, the authors introduce the core principles of bioethics. Specifically, they identify and center the rest of the article around four main principles, for which they provide some specific translations to the scientific research context (Fig 1). As argued by the authors, while these issues may seem too theoretical and abstract, they should directly affect the decisions on what data to collect, how to analyze them and how to disseminate the results of proteomic research.

Fig. 1 Examples of specifications and concrete proteomic examples for the bioethical principles. APOE, apolipoprotein E.

Nonmaleficence is, in simple terms, the (moral) duty to not cause harm. A concrete example that the authors provide, is in the context of incidental findings, whereby a result of unknown significance or that indicates a certain predisposition to a condition can emerge from a proteomics analysis.

Beneficence could be easily thought of as the opposite of the aforementioned (non)maleficence, i.e. the idea of benefiting others. Staying within the same example of incidental findings, reporting them to the individuals that are concerned (directly or indirectly) could prove beneficial if such findings are found to contain information that can improve diagnosis or treatment.

The principle of justice is related to fairness and equality, and a prime example most of us are probably familiar with is the application of FAIR principles to our data, but other related topics mentioned here by the authors are the disclosures of conflicts of interests and the need for representative databases. 

Being autonomous is the ability to choose laws (nomos in greek) for oneself (auto in greek). The authors cite the philosopher’s Immanuel Kant work on the matter, who argued that autonomous agents must also respect the autonomy of others and that this forms the basis of human dignity. In terms of biological research, a concrete manifestation of autonomy is the implementation of informed consents filled and signed by study participants.

Experimental procedure

The authors aimed to collect all existing literature that one way or another deals with these bioethical principles in the context of proteomics. They used specific methods for literature search that are described in detail in the publication, which resulted in 365 unique articles. From there, filtering was applied to remove articles if (a) the mention to the four bioethical principles was considered too limited or peripheral, (b) proteomics was not distinguished from genomics or (c) it was written in a language other than the ones with which the authors have native competencies. As a final result, 16 studies were included for further analysis. These were processed by a trained bioethicist in consultation with a clinical proteomic scientist, who, using a mixed-model approach, aimed to extract data (in this case bioethical issues) and carry out a thematic analysis. Ultimately the authors grouped the identified bioethical issues into 10 themes.

Systematic Review Results

As we walk through the informative blog, now we understand what a systematic review is (summary of all the research on a specific topic based on certain criteria). The review includes 16 total articles. Further, we leap to the results. Talking about ethical issues, we remark that they are relevant and required in the field of clinical proteomics. Based on the systematic review, ten unique categories entitled ‘theme’ are introduced to emphasize the importance of bioethical issues in clinical proteomics (Fig 2). The categories are mentioned in detail below.

Fig.2 The figure reveals the ten themes recognized by the author, mentioned in the text briefly.

Theme 1 – Standards and Quality Control: The considerable effects of variation in all stages of the proteomic workflows, offers an urgent need to initiate and establish conventional operating proceduresinclusive of consistent quality checks. The proteomic workflows mentioned in the literature comprises study design, preanalytical factors, sample collection, storage, and shipping condition. Each workflow mentioned by three authors is implemented by them. With the help of these checks, reproducibility, interoperability, and cooperation among proteomics research labs and allied fields are expected to flourish. Talking about the literature review performed by authors, they present an additional check, an ethical perspective to the technical issues evolved. Therefore, ethical view comes into play while contemplating trust and actual expectation than reality by participants, patients, and the research community at large. 

Theme 2 – Integration of New Technologies and Related Fields: Author mentioned about the need to be updated regarding technology and scientific advances. Such instance is an article proposed the usage of blockchain technologies to be cautious about transparent and secure data access management (Boonen et al., 2019). Additionally, emphasized on the benefit and utilization of linking proteomics with the increasing metadata.

Theme 3 – Identifiability/Privacy: The author proceeds with briefly communicating about proteomic profile,the possibility of uniquely identifying individuals. The early studies remarked this proteomic profile as a hypothetical possibility of human tissue or the databank studies. For example, with the help of hair samples, prominent keratin proteins can assist in differentiating individual profiles, implying to utilize the research in forensic use as well (Laatsch et al., 2014). Also, it was a confirmed information that identifying individuals and the ethnic background can be derived from hair proteome. The similar illustration can be performed for plasma studies with the help of plasma proteome.

Theme 4 – Sensitive Data/Discrimination: The most important ethical concern with respect to usage ofindividual personal sensitive data at times can be used to humiliate individuals. Such occurrences have been reported by Geyer et al, revealing the potential of proteomic profiles, which can be utilized to reveal pregnancy, weight, ethnicity, gender, and allele status of the individuals. But looking to an optimistic outcome, such information can be used for knowing about the family members of a particular patient.

Theme 5 – Conflicting Rights and Duties: This category reveals the important aspect of introducing rights and duties, such as duties expected by the patients may be a conflict with the duties expected by the scientific community. For example, the author talks about the data anonymization, which means right to remove or encrypt the dataset from the data which reveals the personally identified information. This could be used to facilitate the privacy concerns of the patients but can also be a serious concern for the research that includes data correlation. The author mentions about the seven studies emphasizing on the importance of possible conflict of interest between several stakeholders in proteomics (clinical), emerging from intellectual property (IP) protections. IP protections serves best to motivate, execute scientific and innovative ideas, but due to present high level of protection, researchers may find it difficult to access required data or materials for the ongoing research (Lennart & Vizcaíno, 2017).

Theme 6 – Beneficence and Justice-Barrier: The author presents the problems faced by researchers in low/middle-income countries urged by IP protections. Nine studies revealed profit sharing or data sharing as an ethical issue, specifically talking about the databases that are not presenting the global human diversity. Moreover, key highlights regarding the inequitable distribution of benefits in proteomics, concluding to inclusion of financial support from organizations for financial expenses and urgent need of scientific experts.

Theme 7 – Incidental Findings: The author introduces four articles pointing out to the incidental discovery in case of plasma proteomics, functional for medical or social services. In addition, findings, and incidental discoveries from re-evaluation of authenticated pre-existing datasets can be advantageous for advancing research (Boonen et al., 2019). A question arises more on how to maintain, store databases, or look for updated information regarding ongoing research in proteomics and make it available to peer.

Theme 8 – Regulation: The author focuses on the issue of regulation for important national and regional rules & regulations. Mostly talking about informed consent, presented as actual solution. But due to several issues emerging pertaining to research in proteomics such as existing regulation is not enough to acknowledge the issues related to interoperability, efficiency and tasks expected in case of patients and scientific advancements. There is an urgent need to address the issues and discuss about it.

Theme 9 – International Guidance: The author highlights the lack of guidance and simultaneously need for international guidance on the ethical issues. The issue is so prominent that it was mentioned by most of the reviewed articles. Additionally, as a solution stresses on the importance of international collaborations on ethical and scientific concerns.

Theme 10 – Aims and Goals of Clinical Proteomics: The author concludes with researchers talking about guidance in the field of proteomics (clinical) should go beyond the rules & regulations to acknowledge fundamental issues focusing more on goals and funding the projects. In addition to it, how clinical proteomics help us, the human race to overcome health challenges.

Additional perspectives

With this section of the blog, the authors represent their own perspective with respect to bioethical possibilities of clinical proteomics. It starts with talking at length about the standards and study designs in clinical proteomics.

Standards and Study Designs

This section of the blog covers three major concerns from authors own outlook.

1) Medical potential: First point outlines the importance of medical potential of clinical proteomics, which depends upon determination of two groups based on differing health or disease states. To find the inference from these techniques, it mostly depends upon study design, statistical power, thus depicts the differentially express or regulated proteins. Talking about plasma proteomics in this regard, the prototype is to analyze the small number of samples thoroughly. But in case if the data is not enough to perform an analysis, each step performed is futile. For such rationale, the author suggests rectangularly shaped study blueprint than triangularly shaped study blueprint, meaning various large samples are analyzed parallelly to enhance statistical power and significant research.

2) Bias: The author introduces an important viewpoint related to biasness. The recent analysis regarding psychological literature recognizes 34 researcher’s choices in laboratory-based study prototype, where conscious or unconscious bias may arise.

3) Open and Transparent science: The key highlights from making the data availability to all the proteome researchers is an important aspect in open and transparent science leading to FAIR accessibility to the data. This will allow researchers from all corners to re-evaluate and reuse data sets, therefore improve the outcome and benefit everyone.

Overcoming Challenges from Clinical Genomics

Talking about Challenges from clinical genomics, general privacy and health data regulations are significant. As individuals may be discriminated based on health data or demography, to address such issues several regulations have been designed. For example, General Data Protection Regulation (GDPR) and the US Health Information Portability and Accountability Act are implemented to overcome some challenges in the domain. In addition to it, US Health regulations strongly believes in informed participant consent. But also comprises of broader research exemptions as solution to this, such as data minimization and pseudo randomization approach to process the data without consent. 

Moving ahead and talking about Individual research and incidental findings, acknowledging actionable and unactionable information is considered of utmost importance, from ethical point of view. According to the reviewers, actionable information should be returned to the individual or their health care provider and non-actionable information should not. Here, actionable information implies to specific actions taken to overcome health condition.  And with this the question remains answered what should be done with the incidental findings, should include them in the scientific databases or in health registries.

Another important aspect related to genomics and proteomics rises, when concerned about data reuse or re-evaluate the pre-existing data where authors are not in reach to be contacted. As an answer to this issue, the authors suggest consent for the return of the data, or its storage could be partially informed. 

Justice is another situation to deal with. For instance, while talking about polygenic risk scores (number that estimates the effect of genetic variations on individual phenotype), it is more precise for European ancestry than any other ethnicity. This is proved with the availability of 79% of reference genomes from caucasian ancestral lines despite only contributing to 16% of the human population. Moreover, producing demographically representing datasets and Caucasian samples leads to bias in the analysis. This could be overcome with one of the factors such as the financial support to low-to-middle income countries so that such experiments can be carried out. Collaborations act as a boon for such research activities. 

Therefore, a strong need of financial assistance somewhere overcomes the access to the benefits from clinical proteomics and additionally in solving issues of injustice. 

Benefits of Clinical Proteomics

Clinical proteomics plays a vital role in collecting the phenotypic information for the researchers and patients at large. This aspect has a potential to advance and fully explore the biomedical domain if bioethical concept is understood and implemented correctly. Proteomic profiling being a significant aspect to clinical proteomics, assists with several advantages such as providing the adequate information regarding the environmental and endogenous effects on health. Such information can be utilized to impart reasonable health information, therefore advancing in the biomedical field, and sharing its relevance in the real-world. 

Clinical Proteomics is inclusive of neglected factors such as lifestyle of the individuals along with sociocultural and biotic & abiotic determinants of the health.

The implementation of such approach is witnessed in one of the studies of scientific wellness. In the study, multiomic (proteomics, dense dynamics personal data) data was profiled in order to recognize putative biomarkers and coming up with actionable health advice thus leading to improvement in the measured biomarkers among the population. 

Conclusion

With this, we come to summarize the discussion in view of characterizing 10 ethical themes from the 16 included studies. Also, we have briefly discussed about the challenges and solution to overcome barriers in the fields of genomics and clinical proteomics. And provide an additional perspective from authors point. Majorly, an important aspect of not only taking proteomics scientists, budding proteomic researchers into account, collaborating physicians but also engage patients, their advocates, or their organizations (as seen in the field of genomics). This broadens the viewpoint and proves an effective measure for the patients undergoing health issues. Additionally, to persevere in the field of clinical proteomics, the author suggests that we should aim for responsible self-regulation. Specifically, self-regulation based on values is generally proved to be more helpful and authorized. Thus, the experience of genomics reveals the relevant discussions about ethical issues existing in clinical proteomics that can be advantageous to different professions like social scientists, lawyers, ethicists, and humanists. To sum up the discussion from where we started “The time to begin thinking and talking about them is now”. 

If you have made it this far, despite the theoretical depth of the subject, we would love to hear your thoughts in general on bioethics, or in particular on a subject that we touched upon in this post. For example, do you think that, as scientists, we should only be concerned with the hard facts of reality, or should we remain attentive to subjective things like the ethical norms discussed here?

Do you think non-actionable findings are insignificant and should not be reported? Do you often talk about such findings in your current research groups?

References

Boonen K., Hens K., Menschaert G., Baggerman G., Valkenborg D. & Ertaylan G., Beyond genes: Re-identifiability of proteomic data and its implications for personalized medicine. Genes 682 (2019).

Geyer, P. E., Mann, S. P., Treit, P. V. & Mann, M. Plasma Proteomes Can Be Reidentifiable and Potentially Contain Personally Sensitive and Incidental Findings. Mol Cell Proteomics 20, 100035 (2021).

Kalia, S. S. et al. Recommendations for reporting of secondary findings in clinical exome and genome sequencing, 2016 update (ACMG SF v2.0): a policy statement of the American College of Medical Genetics and Genomics. Genetics in Medicine 19, 249–255 (2017).

Laatsch, C. N. et al. Human hair shaft proteomic profiling: individual differences, site specificity and cuticle analysis. PeerJ 2, e506 (2014).

Martens L. & Vizcaíno J. A. A golden age for working with public proteomics data. Trends in biochemical sciences 42.5, 333-341 (2017). 

Mann, S. P., Treit, P. V., Geyer, P. E., Omenn, G. S. & Mann, M. Ethical Principles, Constraints and Opportunities in Clinical Proteomics. Mol Cell Proteomics 100046 (2021).

Moher, D., Liberati, A., Tetzlaff, J. & Altman, D. G. Preferred reporting items for systematic reviews and meta-analyses: the PRISMA statement. BMJ 339, b2535 (2009).

The post A Systematic Review on Bioethics in Proteomics appeared first on PROTrEIN.

]]>
https://protrein.eu/blog/a-systematic-review-on-bioethics-in-proteomics/feed/ 1
Adding the third dimension to protein modifications analysis https://protrein.eu/blog/adding-the-third-dimension-to-protein-modifications-analysis/ Fri, 15 Apr 2022 06:13:04 +0000 http://protrein.eu/?p=794 Short description In the study of post-translational modifications from mass spectrometry experiments, proteins are almost always represented as a linear string of amino acids. However, it should be kept in mind that proteins’ function is conferred by their 3-dimensional structure and that understanding the modifications of proteins calls for studying them in their structural context. […]

The post Adding the third dimension to protein modifications analysis appeared first on PROTrEIN.

]]>
Short description

In the study of post-translational modifications from mass spectrometry experiments, proteins are almost always represented as a linear string of amino acids. However, it should be kept in mind that proteins’ function is conferred by their 3-dimensional structure and that understanding the modifications of proteins calls for studying them in their structural context. In this month’s PROTrEIN journal club, we present a method that aims to bridge the gap between protein structure and data from mass-spectrometry experiments in the context of proteins modification studies.The paper we chose is “The structural context of PTMs at a proteome-wide scale” (Bludau et al. 2022). We selected this paper as it demonstrates how the interpretation of data in mass-spectrometry can be taken a step further by leveraging the knowledge provided by the recent advances in structural proteomics.

Introduction

The regulation and fine-tuning of cellular processes are possible because proteins are highly dynamic molecules. Once synthesized proteins undergo many changes that impact their properties and the way they interact with other proteins or molecules. The alteration of proteins’ properties can be induced in many ways such as changes in the environment with variations in pH or temperature for instance. Protein activity can also be modulated by other proteins with the addition or removal of chemical groups to protein. These modifications of proteins that occur after their synthesis are referred to as post-translational modifications (PTMs). Most proteins are highly modified and to this date, more than 150 different types of PTMs have been discovered. The PTMs are at the core of the regulation of many cellular processes. As an example, the modification of a protein by the addition of ubiquitin is a marker for protein degradation as it is recognized by the proteasome (Hochstrasser et, 1996). A not-to-be-missed topic when it comes to PTMs is the modification of histones. Histones are proteins that bind with the DNA to form a complex referred to as chromatin. Chromatin is a dynamic structure that heavily influences the accessibility of transcription factors to the DNA and therefore genes. The histones are known to be heavily modified and deciphering what is the effect of the different histones’ modification on chromatin structure is crucial to the understanding of epigenetic mechanisms (Millán-Zambrano et al, 2022). 

Measuring protein modifications with mass spectrometry

Several techniques are used to study post-translational modification. For example, immunoprecipitation enables to target and measure proteins with a specific modification or set of modifications. Mass-spectrometry techniques allow for the large-scale study of these post-translational modifications. In a mass spectrometry experiment, it is possible to measure the entire protein sequence for many proteins simultaneously. This allows for a global study of changes in PTMs pattern but also raises many challenges in the interpretation of the data.

Linking proteins modification and structure

The objective of a PTMs study is often to identify a specific protein’s PTM or PTM pattern that can be linked to the regulation of a cellular process. Because of the numerous modifications that are present on proteins, it is particularly difficult to determine which modifications are possibly involved in a given cellular process.As proteins are folded into a 3-dimensional structure not all parts of the protein are exposed to the same extent. The function of a specific modification most likely depends on the localization of that modification in the protein conformation. Protein structural information can therefore be included in PTMs studies to better understand and pinpoint PTMs that have a strong regulatory role. Some regions of the protein fold into a stable structure when some other regions do adopt any stable conformation, these unstable amino-acid chains are referred to as intrinsically disordered regions (IDRs). IDRs are regions of proteins that are frequently involved in protein-protein interaction, that are often enriched in post-translational modifications (see animation below).

Animation of the formation of the nucleosome by the binding of histones (in purple) to the DNA (in red). This simulation shows the behaviors of histone tails (in dark purple) that are IDRs. These unstructured and exposed parts of the histone play a crucial role in regulating how this protein interacts with the DNA and other proteins. (From WEHImovies) 

In the paper presented here, the authors set out to explore the structural context in which PTMs take place, with the goal of understanding the functional role of the PTMs based on their position on the 3d structure.

The protein folding problem 

Despite the advances in technology, the experimental definition of a single 3d structure might take up to several years. For this reason, a different paradigm was explored with the aim of determining the 3d structures starting from the sole amino acids sequence of a protein, the so-called ‘protein folding problem’. Over the course of the past years, various computational approaches have been developed for this goal, using methods ranging from molecular simulations to evolutionary approaches. While the approaches based on molecular interactions rely on computationally heavy procedures that become intractable with the growing size of proteins, recent techniques based on machine learning have been proven to be much more effective for predicting the structure of longer peptide sequences.

CASP and AlphaFold2 

The CASP experiment (Critical Assessment of protein Structure Prediction) is recognized as the state-of-the-art method for evaluating the predictive capability of algorithms aimed at solving the protein folding problem. Every two years, the most advanced methods are tested on a set of completely new and previously uncovered protein structures that have been experimentally determined.

During the past CASP14 assessment, held in 2020, the AlphaFold2 approach (Jumper et al. 2021), developed in Google DeepMind’s laboratories, proved to be remarkably effective for solving the protein folding problem. AlphaFold2 is based on an end-to-end deep learning pipeline that incorporates biological knowledge with attention-based methods for graph inference. This architecture obtained considerably low root-mean-squared deviation when predicting the atomic position of the backbone’s carbon atoms. AlphaFold2 enabled the creation of a complete database (https://alphafold.ebi.ac.uk/) containing almost 1 million accurate protein structure predictions (Varaldi et al., 2021).

Although AlphaFold2 constitutes a huge step toward the advancement of structural informatics, it does not provide a way to introduce PTMs information inside the models. We chose therefore to present a paper that leverages the knowledge that AlphaFold2 provided, with the aim of integrating PTM information with the protein structures provided. 

 “The structural context of PTMs at a proteome-wide scale” (Bludau et al. 2022)

Determining IDRs and side-chain exposure from AlphaFold2 dataset

While the experimental information on protein folding is often limited to specific regions of the sequences, relying on the AlphaFold database enables access to extended information on the folding of complete proteins. However, to exploit the provided 3d structures for PTM analyses, some preprocessing steps are needed in order to generate local information on the shape of the proteins. For each amino acid, the most interesting features to be extracted are its side chain exposure and its belonging to either a structured region or an intrinsically disordered one (IDR).

In order to determine the side chain exposure of a specific amino acid, the authors define a metric called prediction-aware part-sphere exposure (pPSE). The pPSE counts the number of alpha carbon atoms that are present in proximity of each amino acid, by considering a half-sphere centered in the amino acid itself and directed towards its side chain. The radius of the spheres is based on the average size of amino acids and on the flexibility of the side chains. However, in order to take into account the uncertainty introduced with the predicted 3d structures, the radius value is adjusted for each amino acid based on the prediction alignment error (PAE) provided by AlphaFold2. Each residue in the sequence can therefore be associated with a value of pPSE: the lower the PSE, the more the other amino acids are distant from the one under consideration, and the more its side chain will be exposed to other interactions. 

The authors also proved that pPSE can be used to predict IDR regions in an accurate way. For the specific goal of annotating IDRs, the whole sphere is considered when computing the pPSE values, in order to disregard the orientation of the side chain, and the pPSE values are smoothed by computing a sliding average along the amino acid sequence. This metric proved to be the most effective in retrieving IDRs when comparing the results to the available information on already annotated proteins.  

Mitogen-activated protein kinase 3 (MAPK3) colored by prediction-aware part-sphere exposure (pPSE) metric (left) and by the predicted IDRs and structured regions (right).
Figure from (Bludau et al., 2022).

Are PTMs localized in IDRs or structured regions? 

Through enrichment analyses, the authors observed that most PTMs, including phosphorylations, seem to appear in IDRs. The same holds true for acetylation when only the sites with known regulatory regions are considered, while they seem to be underrepresented when considering all PTM sites. Ubiquitination appears instead to be enriched in structured regions, especially when these regions are not associated with regulatory functions. The hypothesis presented by the authors, and verified in different experimental conditions, is that ubiquitination in structured regions could be the main driver that causes misfolding of proteins. The misfolding subsequently leads to exposure of normally unexposed regions where the proteasome can bind and proceed with the degradation of the protein. When the PTMs are instead located in IDRs, the authors observe that these modifications are usually enriched in short IDRs that are included between two larger structured sequences. An enrichment analysis of the proteins with modifications in short IDRs  highlighted processes such as ATP binding, transmembrane activity and protein kinase activity. The authors observe a significant overlap between short IDRs and activation loops of kinases, which are known to be subject to structural changes when phosphorylated. Particularly, the majority of the kinases that show overlap between the annotated activation loop and the predicted short IDR, also show regulatory functional phosposites in the same region. These findings show the importance of studying PTM sites in correspondence of short IDRs, as they seem to be directly correlated with functional regulation.

A. Enrichment analysis of proteins with short IDRs. B. Enrichment analysis of regulatory phosphorylation sites in short IDRs compared to other IDRs. C. 1-dimensional visualization of proteins (N- to C- terminus) with structured regions in blue and IDRs in grey. Regulatory sites are represented as red circles, while other phosphosites are presented in salmon. ‘A-loop’ notation represents the annotated kinase activation loops. D&E. AlphaFold2 predicted structures for RIPK2 and MAP4K1, annotated with information on IDRs and phosphosites. Figure from (Bludau et al., 2022).

PTMs proximity in 3D

It has been proven that PTM-modified sites tend to induce other modifications in regions of the sequence near to the originally modified site and that multiple modifications in sequence proximity are key to regulatory functions and binding properties. Thanks to the structures provided by AlphaFold, the proximity between amino acids can now be defined not only in terms of sequence distance, but also in terms of 3D-proximity. The authors calculate the 3D distance between the various annotated modifications in structured regions and short IDRs, while discarding IDRs since they introduce greater structural uncertainty. For each modification type, the distances to the modifications of the same type were measured and compared to the distances that can be obtained from 5 identical structures with randomly permuted PTMs sites. The results show that PTMs indeed tend to cluster together, with preferential modifications of the acceptors situated near an already modified site. The same analysis was repeated to measure the proximity between different PTM types, observing similar preferential modifications. As an example, it is possible to observe that ubiquitinations, acetylations, and methylations seem to preferentially take place near already modified phosphosites.

A. PTMs acceptors are preferentially modified when proximal to sites showing the same modification type. Observed distance values (red circles if statistically significant, salmon circles otherwise) are compared to the mean of five random samples including the same number of modified PTM sites (grey circles). B. Analysis of different types of modifications in the acceptors near modified phosphosites. C. AKR1B1 protein shows phosphorylations (magenta) in proximal sites based on its predicted 3d structure. Figure from (Bludau et al., 2022).

Conclusion

The work presented here illustrates how combining mass spectrometry data with structural information helps to gain further insight into PTMs’ regulatory role. Studying PTMs considering their localization in proteins can unveil patterns that are not visible without structural information. This article is a great example of how the interpretation of mass-spectrometry data can be extended by merging the information with other types of data.

Bibliography:

Hochstrasser, M., 1996. UBIQUITIN-DEPENDENT PROTEIN DEGRADATION. Annual Review of Genetics 30, 405–439.. doi:10.1146/annurev.genet.30.1.405h

Millán-Zambrano, G., Burton, A., Bannister, A.J. et al. Histone post-translational modifications — cause and consequence of genome function. Nat Rev Genet (2022). https://doi.org/10.1038/s41576-022-00468-7

Bludau, I., Willems, S., Zeng, W.-F., Strauss, M. T., Hansen, F. M., Tanzer, M. C., Karayel, O., Schulman, B. A., & Mann, M. (2022). The structural context of PTMs at a proteome wide scale. Cold Spring Harbor Laboratory. https://doi.org/10.1101/2022.02.23.481596

Jumper, J., Evans, R., Pritzel, A. et al. Highly accurate protein structure prediction with AlphaFold. Nature 596, 583–589 (2021). https://doi.org/10.1038/s41586-021-03819-2

Varadi, M. et al., AlphaFold Protein Structure Database: massively expanding the structural coverage of protein-sequence space with high-accuracy models, Nucleic Acids Research, Volume 50, Issue D1, 7 January 2022, Pages D439–D444, https://doi.org/10.1093/nar/gkab1061

The post Adding the third dimension to protein modifications analysis appeared first on PROTrEIN.

]]>
Machine Learning Algorithms Applications in fMRI Data Analysis https://protrein.eu/blog/machine-learning-algorithms-applications-in-fmri-data-analysis/ Mon, 28 Mar 2022 18:22:35 +0000 http://protrein.eu/?p=800 The article Performance of machine learning classification models of autism using resting-state fMRI is contingent on sample heterogeneity (Maya A. Reiter, Afrooz Jahedi, A. R. Jac Fredo, Inna Fishman, Barbara Bailey, Ralph-Axel Müller, Springer Nature 2020) was chosen because it shows potential applications of machine learning in fMRI data analysis. Unsupervised machine learning has been […]

The post Machine Learning Algorithms Applications in fMRI Data Analysis appeared first on PROTrEIN.

]]>
The article Performance of machine learning classification models of autism using resting-state fMRI is contingent on sample heterogeneity (Maya A. Reiter, Afrooz Jahedi, A. R. Jac Fredo, Inna Fishman, Barbara Bailey, Ralph-Axel Müller, Springer Nature 2020) was chosen because it shows potential applications of machine learning in fMRI data analysis. Unsupervised machine learning has been widely used to study the brain connectome (an example is ICA), although in this article supervised learning was used and the authors showed that this approach is capable of producing better results.

The aim of this review is to give a general overview on machine learning and fMRI. Proteomics and fMRI data analysis share many issues like a huge amount of raw data and the necessity of new algorithms able to handle them. To do so machine learning may be a useful approach, both in proteomics and in neuroimaging.

Introduction

MRI and fMRI theory

Magnetic resonance imaging (MRI) is a widely used technique to acquire high quality images of internal tissues. Its main advantages are: it is not invasive, the magnetic field used is harmless and allows to visualize soft tissues, especially those rich in water or fatty acids.
Images are acquired thanks to hydrogen atoms’ magnetic properties: each hydrogen nucleus is positively charged and since it is constantly spinning it creates a small magnetic field around itself. In normal conditions the aforementioned magnetic field is not detectable, since every field has a random direction. But when atoms are immersed in an external magnetic field they will align along its direction [figure A, left]. In these conditions atoms do not remain still and continue to wobble, causing their magnetic field to wobble as well. The frequency of such wobbling is called Larmor frequency and depends on the strength of the external magnetic field. At this point the scanner emits a radiofrequency pulse at the same frequency as Larmor’s one and this cause atoms to spin in phase and their net magnetization (the small magnetic field created by the sum of all atoms’ one) will then turn 90° and will rotate along the X axis [figure A, right]. This is the signal that is recorded by the machine, then when the radiofrequency pulse ceases this signal is gradually lost. The time required to lose this signal will be translated then into darker or brighter areas in the resulting images. Generally speaking, hydrogen atoms that are tightly packed together (like in fatty acids chains) will lose phase faster as well as atoms that are part of molecules that are freely to move in the tissue (like watery tissues).

Figure A: simple representation of hydrogen nuclei’s behavior in the MRI scanner. Below there is the direction of the magnetic field created by all atoms while the scanner applies an external magnetic field (left) and after the RF pulse (right).

Functional magnetic resonance imaging (fMRI) uses the same principle, but instead of acquiring a single high resolution image it is used to obtain several images across a time interval to describe changes of magnetic properties in the studied tissue, in particular neuronal tissue. Haemoglobin when is not oxygenated can modulate the net magnetization, increasing the magnetic signal recorded by the scanner. Such property is used to describe where neuronal activity happens: when a brain region activates it signals to arteries to increase the blood flow to provide more nutrients and oxygen; at the same time some oxygenated haemoglobin spills into the venous system and this reduces the signal recorded and it is translated into brain activity in the specific region. In figure B is possible to visualize the signal recorded in one single voxel. The data is then “cleaned” removing the noise and eventual artefacts and then is analysed to study if oscillations are indeed caused by neural activity.

Figure B: signal from a single voxel, on the X axis the time in seconds and on the Y the signal intensity.

Connectivity analysis

Connectivity analysis is a kind of analysis done on fMRI data, it is based on the assumption that if different brain regions have a similar pattern of activation they may be inter-connected and involved in similar cognitive functions. The aim is to create a map that shows connections between brain regions and how strong they are [figure C]. To do so one way is to use the independent component analysis (ICA): voxels are divided into components based on their activation pattern. ICA although is not the only machine learning approach that can be used, in particular supervised learning approaches seem to be more precise in describing the connectome.

Figure C: connectome results, each colour represent a brain region and on the left there is the left hemisphere and on the right the right one. Lines that connect regions represent which areas co-activated together. (P. Sripad et al. 2016)

Connectome analysis is used also to describe eventual abnormal patterns of brain activity that correlate with psychiatric disorders. Unfortunately it is quite difficult to study how psychiatric disorders affect patient’s brain activity: biological samples can be obtained only with post-mortem analysis and structural changes (that can be studied with MRI or PET scans) generally happen only at a very late stage of the disease. Currently connectome analysis showed some promising results in describing disorders at an early stage, in particular in this paper researchers used the connectome and machine learning approach to describe brain activity in autistic patients.

Machine learning

Since the first days of existence, learning has been an essential skill acquired by any living entity through experience, study, or being taught. With no exception, it is the same process among plants, animals, humans, and nowadays we can also teach machines to perform tasks.

Although machine learning is a new field, it has been improved and explored massively. Machine learning is a type of Artificial Intelligence that will enable machines to predict some outcomes by recognizing patterns in data rather than explicitly instructing them to do so.

Machine learning has a wide variety of applications not only in Proteomics but also in any field that has access to data, as former blog posts discussed the importance of data reproduction.

There are three different categories of Machine Learning, namely: Supervised learning, Unsupervised learning, and Reinforcement learning.

In this paper, researchers chose Random forest as their classifier which is a Supervised machine learning algorithm. Random forest is an Ensemble learning technique, where multiple learning algorithms are used in conjunction. So, instead of solving a problem with only one model, there are a group of models to solve the same problem by using a shuffled subset of the original data. The final result is the majority vote of model predictions.

This paper mentions that random forest is superior to other ensemble methods in terms of accuracy, computational time, overfitting, and user interface to choose tree size. However, some other ensemble methods such as boosting can overcome random forest in case of accuracy and performance, but they also tend to be harder to tune. Random forest is used to build a diagnostic classifier in autism spectrum disorders (ASDs) samples with a focus on gender and symptom severity of 656 children and adolescents. The data was distinguished into four categories: all genders, only males, all genders with high severity range, and only males with high severity range. The classification accuracy of these categories was 62.5%, 65%, 70%, and 73.75%, respectively. Figure D demonstrates the portion of interest regions included in the classifier that achieved the peak classification accuracy from each of the brain networks for each sample group.

Figure D: Figure D: pie charts show the brain connections that helped to achieve peak accuracy in each of the four sample sets, separated by network. a full heterogeneity, b reduced gender heterogeneity, c reduced ASD-symptom heterogeneity, d low heterogeneity

Conclusion

This paper aims to examine the impact of sample heterogeneity on the diagnostic classification of ASD. Reduced heterogeneity concerning gender and range of symptom severity was associated with improved performance of random forest classifiers. Random forest is implemented to build diagnostic classifiers in four ASD samples including a total of 656 participants. In conclusion, stratification by gender, symptom severity, age, cognitive ability, and other factors of variability may be critical in future efforts to pinpoint atypical brain features of ASDs. In particular the algorithm managed to select several brain regions that appears to be more interesting in describing differences between patients with mild and severe autism. Such differences are also in agreement with what has been studied in the past by psychologists and psychiatrists through cognitive tests.

There are although some limitations pointed out by the authors in the study: lack of enough data coming from female patients with ASD (common problem when studying autism since it is more frequent in males) and a relatively small number of samples used to describe the pathology, but they can be easily solved using wider and wider sample cohorts.

The post Machine Learning Algorithms Applications in fMRI Data Analysis appeared first on PROTrEIN.

]]>
How to train your modified peptide MS/MS spectrum predictor? https://protrein.eu/blog/how-to-train-your-modified-peptide-ms-ms-spectrum-predictor/ Thu, 17 Feb 2022 09:20:54 +0000 http://protrein.eu/?p=808 This month’s PROTrEIN journal club covers an article presenting a tool – pDeep2 – capable of predicting MS/MS spectra of modified peptides.1 pDeep2 is built using a machine learning technique that makes it possible to generate prediction even when there are only a few datasets available for training the model. But… What is this machine learning […]

The post How to train your modified peptide MS/MS spectrum predictor? appeared first on PROTrEIN.

]]>
This month’s PROTrEIN journal club covers an article presenting a tool – pDeep2 – capable of predicting MS/MS spectra of modified peptides.1 pDeep2 is built using a machine learning technique that makes it possible to generate prediction even when there are only a few datasets available for training the model. But… What is this machine learning technique? What are modified peptides? And…- What do they have to do with the PROTrEIN international training network? Read along and you’ll get those answers just in a few minutes!

Background

A short intro to PTMs

The number of protein-coding genes is estimated to be ~20 000 but the collection of human proteoforms is remarkably more diverse than that. There are numerous contributors to the variations at DNA, RNA and protein levels. On the DNA level the main source of variations are coding single-nucleotide polymorphisms (cSNPs) and mutations, on the RNA level it’s alternative splicing and on the protein level post-translational modifications (PTMs).2,3

There have been over 400 PTMs discovered already and their number is growing. PTMs can affect protein function, for instance, phosphorylation4 is the reason for many signalling cascades, acetylation shows relation to blood pressure and hormone regulation, neurodegenerative diseases 5,6, glycosylation is involved in many cellular processes, such as the formation of biofilms and the antimicrobial resistance of various pathogens7, ubiquitination of a proteins most commonly results in the degradation of the proteins 8 and the list goes on.

Because of the aforementioned impacts, it’s clear that identifying and gaining more knowledge about PTMs carries a huge relevance in cell biology and it has great biomedical importance. The main challenges regarding their analysis lie in their usual low abundance, their complexity and the fact that not only their presence but their exact location on the backbone peptides is important too.9,10

Spectra prediction & identifying modified peptides

Why to predict MS/MS spectra?

The most common way of peptide identification is done with database search. Database search consists in comparing experimental spectra to theoretical spectra but traditionally it works with uniform intensities. Another method for peptide identification is spectral library search that uses previously acquired and annotated spectra. Spectral library search includes intensity information but libraries are limited to “already seen” peptides. MS/MS spectra prediction can provide the intensity dimension to theoretical spectra using only some basic inputs such as the peptide sequence and collision energy. Using these predictions, database identification scores can be refined, the number of identification can be increased and false identification can be decreased. 

MS/MS spectra prediction is nowadays built on deep learning methods, borrowing a lot of ideas from Natural Language Processing. The most common neural network architectures applied in spectra prediction are built on Long short-term memory (LSTM) and recurrent neural networks. While spectra prediction has its own difficulties, by including PTMs in the process, further challenges emerge. Predictor neural networks are trained using experimental datasets and even though there are numerous annotated datasets for peptides without modifications, the number of such datasets for modified peptides is limited. 

Machine learning strategies for the small dataset problem

Publicly available datasets can be quite scarce when we try to find solutions to our proteomics related problems. Well annotated datasets are not always available therefore different strategies need to be applied to create useful machine learning models. Some of these strategies are choosing simpler models, careful feature selection, combining several models, extending datasets by either creating synthetic samples or pooling data from other possible sources and finally applying transfer learning.

Transfer learning attempts to create methods to transfer knowledge learned in one or more source tasks and use that knowledge to enhance learning in a similar task. Simply put, it means training a universal model on available large data-sets and then reusing and adapting the model to another task with a small available dataset. These pre-trained models are likely to give better predictions than models only trained with the small available datasets and they are working especially well with deep learning methods.11,12

Short summary of the paper

This month’s article of the PROTrEIN Journal club presents a machine learning model developed using transfer learning to predict MS/MS spectra for modified peptides. The methods and results demonstrated in the paper can be a use for the members of the ITN and for anyone who is reading this blog post.

Introduction

            The goal of the authors was to address the problem of training a good machine learning model that is capable of predicting spectra not only for unmodified peptides and peptides with common PTMs but for low-abundant PTMs as well. With their previous model called pDeep they showed it is possible make accurate spectra predictions for unmodified peptides and they used pDeep2 – a faster and more flexible version of pDeep – as basis to develop a model able to predict spectra for modified peptides.

Methods

            Regarding features for the model, the researchers used a one-hot indicator vector with dimension 20 for each amino acid and a chemical composition vector with dimension 8 to represent common PTMs. These features are then concatenated with precursor charge state, instrument type and collision energy. The outputs of the model are the relative intensities of b/y ions with +1 and +2 charge states. Their initial unmodified pDeep2 model had 2 hidden bidirectional dynamic LSTM (Bi-Dy-LSTM) layers with a hidden layer size of the LSTM cell 256. The dropout was set to 0.2 with 100 epochs, mini-batch size of 1024 and a learning rate of 0.001. The selected loss function for the training was mean absolute error.

            After training the first accurate pDeep2 model using data sets of unmodified PSMs, the authors moved on to the transfer learning step. The first Bi-Dy-LSTM layer was fine-tuned to fit the new input PTM features and the output layer was tuned to fit new outputs but the intermediate hidden layers stayed frozen. The researchers developed the PTM version of pDeep2 so it can also consider b/y ions with neutral losses (NLs) of PTMs. The learning parameters for the transfer learning were set to 20 epochs with a mini-batch size of 1024 and learning rate of 0.001 using mean absolute error as loss function.

            The authors used Pearson correlation coefficient (PCC) as a similarity metric to compare the predicted spectrum with the experimental one. P(PCC > x) or PPCC>x was used as a criterion to evaluate the performance of predictions. PPCC>x refers to the proportion of PCCs that is greater than a given value x. For example, PPCC>0.75 = 95% means there are 95% PCCs greater than 0.75.

Results

            The unmodified pDeep2 model was trained and tested on a total of ∼8 000 000 high-quality unmodified peptide-spectrum with the result that PPCC>0.75 was higher than 90% and PPCC>0.90 was higher than 80%. The researchers tested multiple transfer learning methods and found the “tune-first-last” method best suited for their purposes. The data sets they used for fine tuning included sets of ordinary MS runs that included common PTMs such as oxidation on Met and data sets of 21 synthetic PTMs. The pre-trained model was then fine tuned with aforementioned different PTM data sets and it’s prediction performance was compared to the pre-trained, non-transfer trained and combined models – where combined means the model was trained by combining the unmodified and modified data. The transfer trained models outperformed all other models and reached remarkable performance (See figure below).

The authors also presented that the prediction of PTM NL fragment ions of phosphopeptides may provide complementary information for localising the PTM sites. While PCCs of true and false phosphorylation sites for a sequence were almost the same without PTM NL fragment ions, they managed to show difference when PTM NL fragment ions were included in the model. (See figure below)

Conclusions

            As discussed before, the relevance of PTMs in bio-medicine is huge, therefore they have received a lot of attention from researchers in the proteomics field. In this blog, we briefly covered the topics of PTMs, spectra prediction and transfer learning. We presented the article “MS/MS Spectrum Prediction for Modified Peptides Using pDeep2 Trained by Transfer Learning“1 from Zeng et al. shortly with the intent to spark interest towards these topics and give some useful insights. This study is another example showing that, thanks to the great advancements in machine learning, it is now possible to use deep learning to solve proteomics problems. 

The PROTrEIN international training network’s goal is to train bioinformatics researchers in the field of mass spectrometry based proteomics, therefore it is not a surprise that the PhD topics reflect that interest towards PTMs and the use of machine learning. More than one third of the PROTrEIN ITN projects involve modified peptides and/or machine learning. These projects linked to PTMs cover a whole range of different topics such as PTM landscape characterisation, predicting modified peptide behaviour, identifying and quantifying modified peptides. Other programs are focused on using machine learning more broadly to improve and add new features to tools like MaxQuant and PROSIT.

We hope you enjoyed this post and will check back to us next month!

Bibliography 

1. Zeng, W.-F. et al. MS/MS Spectrum Prediction for Modified Peptides Using pDeep2 Trained by Transfer Learning. Anal. Chem. 91, 9724–9731 (2019).

2. Aebersold, R. et al. How many human proteoforms are there? Nat. Chem. Biol. 14, 206–214 (2018).

3. Ramazi, S. & Zahiri, J. Post-translational modifications in proteins: resources, tools and prediction methods. Database J. Biol. Databases Curation 2021, baab012 (2021).

4. Ardito, F., Giuliani, M., Perrone, D., Troiano, G. & Muzio, L. L. The crucial role of protein phosphorylation in cell signaling and its use as targeted therapy (Review). Int. J. Mol. Med. 40, 271–280 (2017).

5. Xia, C., Tao, Y., Li, M., Che, T. & Qu, J. Protein acetylation and deacetylation: An important regulatory modification in gene transcription (Review). Exp. Ther. Med. 20, 2923–2940 (2020).

6. Drazic, A., Myklebust, L. M., Ree, R. & Arnesen, T. The world of protein acetylation. Biochim. Biophys. Acta BBA – Proteins Proteomics 1864, 1372–1401 (2016).

7. Schulze, S. et al. Enhancing Open Modification Searches via a Combined Approach Facilitated by Ursgal. J. Proteome Res. 20, 1986–1996 (2021).

8. Guo, H. J. & Tadi, P. Biochemistry, Ubiquitination. in StatPearls (StatPearls Publishing, 2022).

9. Torres, M. P., Dewhurst, H. & Sundararaman, N. Proteome-wide Structural Analysis of PTM Hotspots Reveals Regulatory Elements Predicted to Impact Biological Function and Disease *. Mol. Cell. Proteomics 15, 3513–3528 (2016).

10.       Su, M.-G. et al. Investigation and identification of functional post-translational modification sites associated with drug binding and protein-protein interactions. BMC Syst. Biol. 11, 132 (2017).

11.       Torrey, L., Shavlik, J., Walker, T. & Maclin, R. Transfer Learning via Advice Taking. in Advances in Machine Learning I (eds. Koronacki, J., Raś, Z. W., Wierzchoń, S. T. & Kacprzyk, J.) vol. 262 147–170 (Springer Berlin Heidelberg, 2010).

12.       7 Effective Ways to Deal With a Small Dataset | HackerNoon. https://hackernoon.com/7-effective-ways-to-deal-with-a-small-dataset-2gyl407s.

The post How to train your modified peptide MS/MS spectrum predictor? appeared first on PROTrEIN.

]]>