News – PROTrEIN https://protrein.eu Training of computational proteomics researchers Wed, 08 May 2024 14:23:21 +0000 es hourly 1 https://wordpress.org/?v=5.5.1 https://protrein.eu/wp-content/uploads/2020/12/cropped-favicon-32x32.png News – PROTrEIN https://protrein.eu 32 32 Cross-Border Collaboration: Enhancing Peptide Identification with MS2Rescore and MS Amanda https://protrein.eu/blog/cross-border-collaboration-enhancing-peptide-identification-with-ms2rescore-and-ms-amanda/ Wed, 08 May 2024 14:12:42 +0000 https://protrein.eu/?p=1937 Back in June of 2022 I left Austria to go on my first secondment in Ghent, Belgium to visit Alireza and the rest of the Compomics group. Besides getting to know the people,  enjoying the summer in the beautiful city and eating numerous waffles, I also did a small project together with members of the […]

The post Cross-Border Collaboration: Enhancing Peptide Identification with MS2Rescore and MS Amanda appeared first on PROTrEIN.

]]>
Back in June of 2022 I left Austria to go on my first secondment in Ghent, Belgium to visit Alireza and the rest of the Compomics group. Besides getting to know the people,  enjoying the summer in the beautiful city and eating numerous waffles, I also did a small project together with members of the group. Specifically Arthur Declercq and Ralf Gabriels, who are the main developers of the well known rescoring platform MS2Rescore (https://github.com/compomics/ms2rescore) [1]. At that time I was working on making MS2Rescore compatible with MS Amanda [2], a well established database search engine for peptide identification, developed by my supervisor Viktoria Dorfer. 

This small project had the main purpose of me getting to know the tools developed at CompOmics and this eventually lead to a great collaboration between our two groups that resulted in this recent publication:

“MS2Rescore 3.0 is a modular, flexible, and user-friendly platform to boost peptide identifications, as showcased with MS Amanda 3.0” 

[Caption: Peptide identification result files from many different search engines, including MS Amanda, are parsed to MS2Rescore which will then perform data-driven rescoring, utilizing several different feature generators and rescoring engines. This leads to a higher number of confident identifications at the same false discovery rate (FDR) threshold or a similar number of confident identifications at a more stringent FDR threshold. MS2Rescore results can be easily inspected in the newly added HTML quality control reports.]

We’re excited to introduce the latest, enhanced versions of both MS2Rescore and MS Amanda. MS2Rescore 3.0 is highly modularized and flexible, making it easy to add new input formats through the python package psm_utils (https://github.com/compomics/psm_utils) [3], new feature generation modules and new rescoring modules. This version of MS2Rescore is available both as a command line interface, a graphical user interface and as a python API, making implementation into already existing workflows for peptide identification very easy.  Lastly, MS2Rescore 3.0 will output a HTML quality control report after the rescoring which allows users to assess the effect of data-driven rescoring on their identification workflow and the performance of individual features. With MS Amanda 3.0 come seven new columns in the output CSV file that can be used for rescoring search results as well as automatic integration with Percolator [4]. 

We demonstrate the flexibility of MS2Rescore 3.0 by connecting it with MS Amanda 3.0 and show how using MS2Rescore 3.0 in combination with MS Amanda 3.0 can increase the number of identified spectra in a challenging single-cell data set. 

It feels great knowing that what started as a small project during my secondment eventually ended up with me having my first publication. Since this was a highly collaborative process between our two research groups it is important to mention that Arthur Declercq and I are shared first authors and that Viktoria Dorfer and Ralf Gabriels are shared last authors. Thanks again to Arthur, Ralf, Viki and the rest of my co-authors! 

You can read the publication here:

https://pubs.acs.org/doi/10.1021/acs.jproteome.3c00785

Or the preprint, free of charge, here:

https://chemrxiv.org/engage/chemrxiv/article-details/65b614e2e9ebbb4db9237db9

References

  1. Declercq A, Bouwmeester R, Hirschler A, Carapito C, Degroeve S, Martens L, Gabriels R. MS2Rescore: Data-Driven Rescoring Dramatically Boosts Immunopeptide Identification Rates. Mol Cell Proteomics. 2022 Aug;21(8):100266. doi: 10.1016/j.mcpro.2022.100266. Epub 2022 Jul 6. PMID: 35803561; PMCID: PMC9411678.
  2. Dorfer V, Pichler P, Stranzl T, Stadlmann J, Taus T, Winkler S, Mechtler K. MS Amanda, a universal identification algorithm optimized for high accuracy tandem mass spectra. J Proteome Res. 2014 Aug 1;13(8):3679-84. doi: 10.1021/pr500202e. Epub 2014 Jun 26. PMID: 24909410; PMCID: PMC4119474. 
  3. Gabriels R, Declercq A, Bouwmeester R, Degroeve S, Martens L. psm_utils: A High-Level Python API for Parsing and Handling Peptide-Spectrum Matches and Proteomics Search Results. J Proteome Res. 2023 Feb 3;22(2):557-560. doi: 10.1021/acs.jproteome.2c00609. Epub 2022 Dec 12. PMID: 36508242.
  4. Käll L, Canterbury JD, Weston J, Noble WS, MacCoss MJ. Semi-supervised learning for peptide identification from shotgun proteomics datasets. Nat Methods. 2007 Nov;4(11):923-5. doi: 10.1038/nmeth1113. Epub 2007 Oct 21. PMID: 17952086. 

The post Cross-Border Collaboration: Enhancing Peptide Identification with MS2Rescore and MS Amanda appeared first on PROTrEIN.

]]>
Exploring Cellular Complexity: Unveiling Single-Cell Proteomics https://protrein.eu/blog/exploring-cellular-complexity-unveiling-single-cell-proteomics/ Fri, 08 Sep 2023 15:48:03 +0000 https://protrein.eu/?p=1910 There are enormous amounts of biological cascades in every cell – the smallest functional compartment of our body [1]. Understanding the mechanisms underlying the vast array of phenomena is not only the key element to finding any clues about fatal diseases such as Alzheimer’s and cancer but also to progress in developing treatment for such […]

The post Exploring Cellular Complexity: Unveiling Single-Cell Proteomics appeared first on PROTrEIN.

]]>
There are enormous amounts of biological cascades in every cell – the smallest functional compartment of our body [1]. Understanding the mechanisms underlying the vast array of phenomena is not only the key element to finding any clues about fatal diseases such as Alzheimer’s and cancer but also to progress in developing treatment for such diseases. To address this need, scientists across diverse biological disciplines have embraced a multi-omic analysis approach, deciphering meaningful codes enciphered by cells, such as the genome, transcriptome, and proteome [1,2]. Recently, the discovery of new developments in the omics approaches allows us to conduct these analyses at single-cell resolution [3,4].

In this month of the journal club, we would like to touch upon single-cell technologies from the mass spectrometry-based proteomics perspective. The title of the selected article is “Ultra-high sensitivity mass spectrometry quantifies single-cell proteome changes upon perturbation” published by Andreas-David Brunner et al [4]. Before going into details about the article, we would like to give brief information about single-cell technology.

One of the most well-known analogies for the single-cell method is the smoothie example [3]. Imagine that you take a sip from a smoothie, and you will sense different flavors of fruits. However, this feeling will vary depending on the amount and characteristic tastes of fruits. For example, if there are lots of oranges in it, the high acidity due to the number of oranges might mask a less noticeable taste such as blueberry. Things become more complicated when you want to predict the ratio of each fruit that was blended for the smoothie. What if you have the opportunity to taste each fruit separately without making a smoothie? In the context of this analogy, the conventional bottom-up proteomic approach refers to analyzing protein levels by mixing all cells like a smoothie. Thus, the measurement of proteins would be the average of all cells that were involved in the study. Distinguishing the taste of blueberries in this case low-abundant species, rare cell types, and sub-populations becomes more complicated. However, single-cell proteomics enables us comprehensive understanding of cellular heterogeneity such as the immune system and cancer formation phases [2]. Since the combination of single-cell methods with MS-based proteomic studies was implemented less than 10 years ago, some challenges still need to be solved in terms of depth of coverage, sensitivity, robustness, and cost [1]. In this article, they proposed a novel true single-cell proteome (T-SCP) pipeline by enhancing sensitivity and optimizing the mass spectrometry (MS) setup.

They introduced the PASEF acquisition scheme for noise-reduced quantitative mass spectra, enabling highly sensitive and complete proteome measurements. Through a dilution series of HeLa cell lysate, they identified over 550 proteins from low amounts and achieved excellent quantitative reproducibility. To enable true single-cell proteomics, they significantly increased MS sensitivity and adapted their workflow. The researchers sorted and analyzed individual single HeLa cells, achieving protein identification and quantification. Sensitivity improvements allowed the identification of more than 1,890 proteins from as few as six single cells. They established a «core-proteome» subset of stable proteins that could serve as normalization factors and identified distinct protein regulation mechanisms.

Further sensitivity enhancements, including reduced flow rates, enabled a ten-fold increase in sensitivity, leading to the identification and quantification of over 3,900 HeLa proteins from just 1 ng of material. They adopted the diaPASEF acquisition mode for increased reproducibility.

Applying this technology, the study explored the cell cycle’s impact on single-cell proteomes. They treated HeLa cells with thymidine and nocodazole, quantifying up to 2,501 proteins per single cell across different cell cycle stages. The data revealed high quantitative precision and allowed the differentiation of cell cycle stages based on proteomic profiles. Using marker proteins, the researchers successfully assigned cellular states and distinguished cell cycle phases.

The single-cell proteome analysis also revealed differential expression of known cell cycle regulators and highlighted novel proteins associated with the G2/M transition.

Figure 1. A novel mass spectrometer allows the analysis of true single-cell proteomes. (Second figure in the article). A: Raw signal increase from standard versus modified TIMS-qTOF instrument (left) and at the evidence level (quantified peptide features in MaxQuant) (right). B: Proteins quantified from one to six single HeLa cells, either with MBR in MaxQuant (orange) or without MBR (blue). The outlier in the three-cell measurement in gray (no MBR) or white (with MBR) is likely due to failure of FACS sorting as it identified a similar number of proteins as blank runs (Horizontal lines within each respective cell count indicate median values). C: Quantitative reproducibility in a rank order plot of a six-cell replicate experiment. D: Same as C for two independent single cells. E: Rank order of protein signals in the six-cell experiment (blue) with proteins quantified in a single cell colored in orange. F: Raw MS1-level spectrum of one precursor isotope pattern of the indicated sequence and shared between the single-cell (top) and six-cell experiments (bottom).

SC proteomes compared to transcriptomes:

The study evaluated over 430 single-cell proteomes and compared them with Drop-seq and SMART-Seq2 scRNA-seq data to gain technology-independent insights. Proteome measurements exhibited higher average correlations among cells compared to scRNA-seq methods. Protein completeness per cell followed a normal distribution, with proteomic 

capturing about 49% of observed proteins, whereas SMART-Seq2 captured only 27% and Drop-seq captured 8%. Detecting limitations in protein measurements, bimodality in lower protein abundance range indicated potential benefits of imputation or tailored parameter estimation methods. While bulk transcript and protein levels showed moderate correlation, single-cell levels diverged, underscoring distinct regulatory mechanisms.

Examining shared gene CVs, single-cell transcriptomes displayed consistent quantitative variation, unlike proteomes. This emphasized different single-cell regulation for protein and RNA abundance. The research concluded that protein and RNA measurements offer complementary information, uncovering unique regulatory mechanisms. A stable core proteome subset of top 200 proteins with low CVs was identified, representing normalization factors and essential cellular processes. This subset’s distribution across the proteome’s dynamic range indicated stability even during remodeling. Overall, the study deepens insights into single-cell proteome dynamics and its interplay with gene expression.

References 

[1]       K. Vandereyken, A. Sifrim, B. Thienpont, and T. Voet, “Methods and applications for single-cell and spatial multi-omics,” Nat Rev Genet, p. 1, Aug. 2023, doi: 10.1038/S41576-023-00580-2.

[2]       E. Flynn, A. Almonte-Loya, and G. K. Fragiadakis, “Single-Cell Multiomics,” https://doi.org/10.1146/annurev-biodatasci-020422-050645, vol. 6, no. 1, Aug. 2023, doi: 10.1146/ANNUREV-BIODATASCI-020422-050645.

[3]       “What is single-cell sequencing? – Single Cell Discoveries.” https://www.scdiscoveries.com/blog/knowledge/what-is-single-cell-sequencing/ (accessed Aug. 21, 2023).

[4]       A. Brunner et al., “Ultra-high sensitivity mass spectrometry quantifies single-cell proteome changes upon perturbation,” Mol Syst Biol, vol. 18, no. 3, Mar. 2022, doi: 10.15252/MSB.202110798.

The post Exploring Cellular Complexity: Unveiling Single-Cell Proteomics appeared first on PROTrEIN.

]]>
Modeling Lower-Order Statistics to Enable Decoy-Free FDR Estimation in Proteomics https://protrein.eu/blog/modeling-lower-order-statistics-to-enable-decoy-free-fdr-estimation-in-proteomics/ Wed, 23 Aug 2023 11:59:13 +0000 https://protrein.eu/?p=1902 In our previous Journal Club blog post, we discussed «nanopore profiling,» a method for identifying proteins within the field of proteomics. Today we proceed with «how to validate the identification statistically other than the traditional target-decoy method with an improved version of the decoy-free approach» based on Dominik et al.’s article «Modeling lower-order statistics to […]

The post Modeling Lower-Order Statistics to Enable Decoy-Free FDR Estimation in Proteomics appeared first on PROTrEIN.

]]>
In our previous Journal Club blog post, we discussed «nanopore profiling,» a method for identifying proteins within the field of proteomics. Today we proceed with «how to validate the identification statistically other than the traditional target-decoy method with an improved version of the decoy-free approach» based on Dominik et al.’s article «Modeling lower-order statistics to enable decoy-free FDR estimation in proteomics«, we present in this blog post a new method for validating identification in proteomics.

There are two methodologies for calculating the false discovery rate (FDR) in proteomics: decoy-based and decoy-free. In decoy-based approaches, spectra are compared to the sequences of naturally occurring (target) proteins and decoys that are in-silico generated, based on peptide sequences of the target database. Using decoy peptide spectrum matches (PSMs), the conventional target-decoy FDR calculation takes into consideration the characteristics of incorrect target PSMs. Depending on the decoy generation method, this increases the cost of computation and decreases the likelihood that accurate PSMs will be considered. In addition, decoy-based methods may under- or overestimate the FDR if the scoring function used in the step of searching the target-decoy database has some bias or if there are insufficient decoys in the region where the models of the correct and incorrect PSMs overlap. Decoy-free statistical validation tools that employ only target PSMs constitute another category of peptide identification validation methods. While the majority of decoy-free statistical validation tools model the score distribution of the highest-scoring PSMs, some also exploit the PSMs with lower scores. The current decoy-free methods typically lack a solid theoretical foundation and rely significantly on the empirical characteristics of the data, or they rely on theoretical assumptions that are not always well-justified.

In this article, the new idea of decoy-free FDR estimation is propose with a semiempirical framework based on lower-scoring target PSMs where as a relationship between the parameters of the distributions of low-order statistics of the log transformed e-value (TEV) score and a necessary empirical optimization to fit a single parameter to real data. The theoretical sharing of parameters μ and β across different orders of Target E-Value (TEV) distributions in proteomics. However, empirical estimation of these parameters reveals slight deviations from the theoretical values. To address this, the article proposes two semiempirical optimization methods for estimating adjusted μ and β values for the Top Null Model (TNM). The methods include a linear regression-based approach and a mean β-based approach. These approaches are applied to data sets, and the best TNM variant is selected based on the Bayesian information criterion (BIC). 

The performance of different Top Null Models (TNMs) estimated using lower-order models generated by the Tide and Comet search engines was evaluated. The evaluation focused on FDR estimation and compared the TNM approaches against Couté’s method, the Gumbel TEV model, and the common decoy distribution (CDD) method. The comparison considered metrics such as false discovery proportion (FDP) and the number of correctly identified spectra at different FDR thresholds. FDR control was performed using the Benjamini-Hochberg (BH) procedure. A ground truth data set was created by searching files against a target-decoy database and generating incorrect PSMs. The TNMs were generated based on the top-scoring PSMs, and the optimal μ and β estimates were determined using the proposed semiempirical estimation methods. FDR control was then applied using the BH procedure, and the results were evaluated using the ground truth labels. The process was repeated on bootstrapped samples, and mean FDP and correct identification values with confidence intervals were calculated.

To test the quality of lower-order models and top null models, data sets of natural peptides from five different species (H. sapiens, M. musculus, A. thaliana, S. cerevisiae, and E. coli) were taken from project repositories in the PRIDE archive. Synthetic human peptide data sets from the ProteomeTools project were used in the validation study. Only MS2 spectra with charge states of 2+, 3+, and 4+ were taken into consideration for the evaluation. The lower-order models were found to fit the empirical data well, with some discrepancies for lower order indices. The maximum likelihood estimation (MLE) method performed better than the method of moments (MM) for parameter estimation in the lower-order models. The proposed models accurately fit the empirical distributions and were not significantly affected by the size difference between the analyzed data sets. The study also compared the performance of the lower-order models with decoy-based models and Couté’s method, and found that the lower-order models estimated FDRs better, particularly for Tide results. However, for Comet results, Couté’s method and decoy-based models overestimated FDRs, while the common decoy distribution method provided second-best estimates. The differences in performance between Tide and Comet can be attributed to the less rigorous e-values produced by Comet. The proposed approach, combining theoretical foundations with empirical optimization, showed resistance to issues associated with statistical scoring in shotgun proteomics.

The proposed approach eliminates the need for decoy sequences and offers improved accuracy compared to alternative methods. While further evaluation and tuning may be necessary for different identification tools, this work highlights the untapped potential of lower-scoring PSMs in enhancing statistical validation methods in proteomics research.

Finally, thanks for reading our post and keep tuned for further content!

The post Modeling Lower-Order Statistics to Enable Decoy-Free FDR Estimation in Proteomics appeared first on PROTrEIN.

]]>
Peptide De Novo Sequencing What are the ingredients of that delicious pizza? https://protrein.eu/blog/peptide-de-novo-sequencing-what-are-the-ingredients-of-that-delicious-pizza/ Tue, 08 Aug 2023 10:07:26 +0000 https://protrein.eu/?p=1870 Proteins, the mighty microscopic marvels present in what we eat, hence, what we are. In milkshakes 🥤, ice creams 🍦… your hair, your skin… you 👤. These molecular machines are present in all living things, from viruses to the dog 🐶 that barked at you the other day. They are so essential that they keep […]

The post Peptide De Novo Sequencing What are the ingredients of that delicious pizza? appeared first on PROTrEIN.

]]>
Proteins, the mighty microscopic marvels present in what we eat, hence, what we are. In milkshakes 🥤, ice creams 🍦… your hair, your skin… you 👤. These molecular machines are present in all living things, from viruses to the dog 🐶 that barked at you the other day. They are so essential that they keep the grand show of life going. Proteins are like the ultimate multi-taskers of the cellular world. They’re the tiny construction workers 🏗, dutiful soldiers, speedy messengers, vigilant guards and the skilled repair crew 🔧 that keep our bodies running smoothly day in, day out.

They’re crafted by our cells using the blueprints written in our DNA. Heard of the genetic code in the DNA? This code is like a huge library 📚, stuffed with cookbooks. These cookbooks contain genes, which are like recipes for whipping up every protein your body needs. And there are at least 10000 different proteins keeping you alive.

These proteins are built inside your cells using tiny molecules called amino acids 🧪, which are like the ingredients in your recipe. These cells are guided by the instructions in the DNA on how to put together particular proteins. Picture a cake 🍰, it needs flour, eggs, sugar, and butter, all in specific quantities and added in a certain order. Similarly, proteins need specific amino acids in the right order to be cooked up right. DNA has the recipe, cookbooks 📚. Amino acids are the ingredients. Proteins are the dishes 🍽. Hungry yet? 😋 

As an example, the following genetic code in your DNA, could be the cookbook instruction 📖 for creating the dish (protein) we want. 

TGT – TAC – ATT – CAA – AAT – TGT – CCT – CTC – GGT 

These instructions would lead the cells to place amino acids (ingredients) one after another leading to a molecular chain of amino acids placed in the following order. 

Cysteine – Tyrosine – Isoleucine – Glutamine – Asparagine – Cysteine – Proline – Leucine – Glycine 


CYIQNCPLG 

In this case, the instruction earlier in the DNA was for a small protein (technically peptide, but we will get to that.) called Oxytocin 😍, often referred to as the «love hormone» or the «cuddle chemical.» It plays a crucial role in social bonding, fostering feelings of trust, empathy, and connection. The amino acid sequence for this oxytocin is CYIQNCPLG.
There are over 20 amino acids, and only using 20 of these, our bodies create 10,000 different types of proteins, with each protein playing a different role. For example, hemoglobin protein 💉 is made by the cells in your bone marrow, because the cells are instructed by your DNA to do so, which is then released into your bloodstream. Tyrosinase is another protein, which is responsible for producing melanin, which gives your hair the color 🌈. Loss of this protein, or cells choosing not to produce it can lead to loss of hair color. 

While oxytocin is just 9 amino acids, hemoglobin is about 300 and Tyrosinase is over 500 amino acids long.

So, What is de novo sequencing?

Here’s the big challenge. Let’s call it «de novo sequencing» which is a fancy way of saying «figuring out the ingredients of a protein by tasting it» 🕵️‍♀️🔍. It’s like heading to a restaurant, eating a lip-smacking dish 🍲, and then trying to identify the ingredients just based on your taste! Sounds fun but tough, doesn’t it? The goal is to identify the composition and sequence of the protein. That is, to identify the amino acid pattern. 

At this point, allow me to introduce the superstars of our story — peptides. Peptides are like mini proteins. Think about a giant pizza 🍕(that’s your protein), peptides are like those cute mini pizza bites, smaller in size but oh-so-crucial. Peptides are just shorter chains of these all-important amino acids. Remember oxytocin, that’s a peptide, a mini protein. 💖 

From now on, we’re shining the spotlight 🎯 on these little strands of amino acids, the peptides. Let’s think of them as the underdogs, small but mighty, and loaded with info about our bodies. 🏋️‍♀️ 

So, why the heck do we want to figure out the sequence of the peptide? 🤷‍♀️ Well, if we crack the sequence of amino acids – the ‘secret ingredients’ 🧪 – in our peptide, we unlock a whole treasure chest 🗝🔓 of information about what it is and what it does. It’s like decoding a secret message written in an alien language! 👽📜 

By identifying a peptide in a sample, we could potentially develop a new wonder drug 💊 that uses the peptide, or a treatment that blocks it. It’s like having a secret weapon in the battle against diseases! 🛡 We could even learn something groundbreaking about how the body works! 🧠⚡ It’s like discovering a new law of physics… but inside our bodies! How cool is that? 🤓🚀

So, how do we do it? 

Well, we could just squint really hard at them under a microscope 🦠🔬. But trust me, even if you have the eyes of an eagle, you wouldn’t see much. They’re way too small. Sure, there have been some incredible advances recently with techniques that read each amino acid, but it’s still not always possible.

How about just weighing it? 🏋️‍♀️ If my peptide weighs, let’s say, 200 (in super-simplified units), then I could whip out a catalog 📚 and see that 200 is the mass of oxytocin. So, it must be oxytocin, right? Ehhh… not so fast. There could be a whole bunch of other peptides out there that also weigh 200. Worse, what if this peptide is from a snake venom 🐍 that’s never been observed before and hence, doesn’t even make it to the catalog?

Aha! Scientists early on had a light bulb 💡 moment. They thought, «Let’s break the peptide into pieces, then weigh the pieces.» It’s like disassembling a Lego tower to understand how it was built.

But then, a new challenge raises its ugly head 🐲. How do we weigh so many tiny pieces at once? It’s like trying to weigh a bunch of confetti 🎉 – all at the same time!


Enter the unsung heroes of our story – mass spectrometers 🌠. These are like high-tech super scales 🧱⚖ that can weigh a ton of things all at once. 

But wait, there’s more! We usually measure a lot of things at once, and mass spectrometers don’t just stop at weighing. They also measure quantities 🔢. It’s like a super-smart scale saying… «Oh, your sample had more M&Ms 🍫 than Coffee beans ☕ in your Mocha coffee bean cookie 🍪.» How cool is that!? 

Let’s say we’ve got our hands on a peptide called CAT (not the furry creature 🐱 but a peptide whose amino acid sequence is C, A, and then T). In fact, let’s say we have a whole pile of them, like a giant CAT party 🎉. But before we can start the weighing party, we’ve got to break them into smaller parts. Kind of like smashing a piñata 🎊. And we have some pretty fancy ways to do this (imagine sophisticated scientific pinata smashers like CID, ETD etc.). 

When we bash our CAT peptide, it can break into C, CA pieces. 

ADITI can break into A, AD, ADI and ADIT and so on. 

So, let’s say, we’ve got our CAT peptide broken into C and CA parts. Now, when we feed these pieces into our trusty mass spectrometer to weigh, we should see two distinct signals 📈📉. One corresponds to the smaller C piece and another to the larger CA piece. 

But hold on, what about the whole, unbroken CAT peptide? Well, it can still hang around in our sample, giving us a third signal 📊. 

So, in the end, we’ve got three peaks on our mass spectrometer’s readout, one for each of the C, CA, and the intact CAT. 

For example, talking about CAT, we know that amino acid C weighs about 103 giving us a peak at 103. Similarly, we know that A weighs about 70, so the mass of CA must be 103 + 70 = 173, where we see the second peak. The mass of amino acid T is about 100, giving us another peak from CAT, at 103 + 70 + 100 = 273. 

The same process happens with a peptide called ADITI. The peak at 300 is due to ADI, and as T is about 100, we get a peak at 400 corresponding to ADIT. 

The values can be generated for any sequence. Here’s a tool to try it out yourself. (use b ions checkbox only in the tool, this will be explained in detail later) 

Fragment Ion Calculator – systemsbiology.net 

But let’s go back to CAT.. Let’s say, you didn’t know this was CAT peptide and we told you that this peptide had three amino acids. Can you figure out the amino acid sequence? Are you ready to perform De Novo Sequencing?

Let’s do De Novo Sequencing 

So, you have taken an unknown peptide, smashed it and then sent it into the mass spectrometer. The mass spectrometer has spit out a spectra with three humps. Time to figure out what was in the sample. 
To begin, let’s just call the unknown peptide X1 X2 X3.

The furthest peak must be from the largest mass, X1 X2 X3 itself, the unbroken peptide.  

So, X1 X2 X3 weighs 273. This cannot directly tell us anything, as we know the masses of amino acids only. And yes, we are approximating here for simplicity.

One can see that the peak for the smallest mass is 103 which should be the smallest fragment. So in our case, X1 is the smallest fragment and must have a mass of 103. This is a single amino acid. But which is it? This is when you go through the amino acid list. C is an amino acid with a mass of 103, so X1 must be C. 

The next peak is at 173 and it must be from X1 X2. So, X1 X2 weighs 173. 

This tells us that X2 must be an amino acid with a weight 70. Go through the amino acid list, there is an amino acid corresponding to 70, it’s A, Alanine. So, X2 must be A. And finally, the peak at 273 must be from the whole peptide X1 X2 X3, so since X1 X2 is 173, X3 must be an amino acid with mass 100. Which our list tells us, X3 must be T.

So.. we now know X1 X2 X3 is CAT. Congrats. That’s your first de novo sequencing. 

Now that you know, wanna try out another? Here’s some help, this one is 5 amino acids long peptides.

And the solution is…

TEAM.

CEAM? That’s possible too.. If we consider 102 to be associated with C. This is a very simplified version. Because, isn’t T, 101.04?  
Yep, you caught us! 🎣 We did some rounding up, sort of like rounding up π to 3 (although not quite as dramatic!). However, it’s important to note there’s another simplification here that matters a lot. The ends of these peptides – imagine them like the caps on a tube of toothpaste 🦷 – can have some effects on the values we see in the spectra. 

We hope you enjoyed this read and got a grasp of what De Novo Sequencing is. Before we dive deeper into the topic, let’s take a breath. Stay tuned for the second part of the post, in which we will also play some fun puzzles 🧩.

References :
1) Main Paper Reference : Medzihradszky, K.F. and Chalkley, R.J. (2013) ‘Lessons inde novopeptide sequencing by Tandem Mass Spectrometry’, Mass Spectrometry Reviews, 34(1), pp. 43–63. doi:10.1002/mas.21406.
2) Secondary Paper : CHONG, K.F. and LEONG, H.W. (2012) ‘Tutorial on de novo peptide sequencing using MS/ms mass spectrometry’, Journal of Bioinformatics and Computational Biology, 10(06), p. 1231002. doi:10.1142/s0219720012310026.
3) Blog Post Referred : https://www.ionsource.com/tutorial/DeNovo/DeNovoTOC.htmI must add. 4) Tutorial «Manual spectra annotation and automatic database/library search» Viktoria Dorfer, FHOOE, PROTrEIN summer School 2021:
https://www.youtube.com/watch?v=ztyglWkY1iI

The post Peptide De Novo Sequencing What are the ingredients of that delicious pizza? appeared first on PROTrEIN.

]]>
Mass spectrometry-based proteomics imputation using self-supervised deep learning https://protrein.eu/blog/mass-spectrometry-based-proteomics-imputation-using-self-supervised-deep-learning/ Thu, 03 Aug 2023 13:51:21 +0000 https://protrein.eu/?p=1858 Hello and welcome back to another issue of the PROTrEIN Journal Club! This occasion we will cover important topics in proteomics, missing values and imputation. We will try to shed light to some of the challenges regarding these matters with the aid of the article titled: “Mass spectrometry-based proteomics imputation using self supervised deep learning” […]

The post Mass spectrometry-based proteomics imputation using self-supervised deep learning appeared first on PROTrEIN.

]]>
Hello and welcome back to another issue of the PROTrEIN Journal Club! This occasion we will cover important topics in proteomics, missing values and imputation. We will try to shed light to some of the challenges regarding these matters with the aid of the article titled: “Mass spectrometry-based proteomics imputation using self supervised deep learning” from Henry Webel et al.1 Also after a few weeks of hiatus, we are back with some machine learning too.

The search for biomarkers and the identification of new drug targets are important use cases of mass spectrometry (MS) based label-free proteomics. However, the downstream analysis of acquired data is largely impacted by missing values. There can be many roots of missing values, but the main contributing factors can be divided into two categories: biological factors such as proteins not existing in the sample or their abundance is below instrument detection limit and analytical factors, for instance poor ionisation efficiency, bad peptides-spectrum matches, stochasticity of precursor selection for fragmentation.2

To address the problem of missing values, different imputation methods were developed. These methods can impute quantification values on various levels such as precursor, aggregated peptides and protein group levels. One common approach is median imputation per feature across samples and another is interpolation of missing features by close replicates. A more sophisticated approach exists that imputes data at protein group level using random draws from down-shifted normal (RSN) distribution. The assumption here is that the values are missing due to absence or lower abundance in the sample than the detection limit. However, that can create biases and skew the downstream analysis. The authors of the paper are presenting three machine learning models to predict missing quantification values. 

These  alternative deep learning (DL) models use different strategies —collaborative filtering (CF), denoising autoencoder (DAE), and variational autoencoder (VAE)— to impute missing values in proteomics data sets. The training objectives, complexity, and therefore capabilities of the models are different which led authors to evaluate their performance in comparison to each other. The CF and autoencoder objective only focuses on reconstruction, whereas the VAE adds a constraint on the latent representation. Furthermore, the first two modeling approaches use a mean-squared error (MSE) reconstruction loss, whereas the VAE uses a probabilistic loss to assess the reconstruction error.

These models were applied to large (N≈450) and small (N≈50) MS-based proteomics data sets of HeLa cell line tryptic lysates acquired over two years during continuous quality control in two different labs at Novo Nordisk Foundation Center for Protein Research (NNF CPR) and Max Planck Institute of Biochemistry. The effectiveness of the models was assessed in comparison to two heuristic-based methods: median imputation and interpolation of missing features. The results show that the self-supervised models, i.e. CF, DAE, and VAE, outperform the heuristic-based approaches, with half of the median imputation mean absolute error (MAE). By identifying (+23.6%) more significantly differentially abundant protein groups, the VAE model in particular is demonstrated to be useful in illness prediction.

The two autoencoder architectures represented a sample in a low-dimensional space using all of the data. The CF model, in contrast, needed to learn a latent embedding space for both the samples and the features. Overall, While the DL methods and median imputation can impute all missing values, interpolation does not replace missing values in case a value is missing in all replicates. The study also discovers that the models’ performance changes based on how frequently a protein group is observed, with better performance for groups observed in more than 80% of the samples. For protein-level data, the models’ overall performance is shown to be the poorest, for aggregated peptides, it is better, and for precursors, it is best.

In the development datasets, while the DL techniques outperform interpolation and have around half the median imputation MAE, The three DL approaches perform about the same. Consequently, when compared to the self-supervised models, the median imputation and interpolation models performed about 1.8–2.4 times worse.

Performance of imputation methods at the level of protein groups, aggregated peptides, and precursors for MaxQuant outputs

The authors were testing the impact of their developed imputation techniques on a real-world dataset of 455 blood plasma proteomics samples from a cohort of alcohol-related liver disease (ALD) and healthy controls. The study3 where the real-world data originated from, was looking for biomarkers of ALD in the proteomics samples that could enable MS-based liver disease testing. One of the key pathological features of alcohol-related liver disease is fibrosis, therefore proteins related to fibrosis were monitored. In addition then they were training machine learning models to predict fibrosis and inflammation from the MS plasma protein groups. In the referenced article3, they used the RSN imputation approach, hence the authors compared their methods to the results of that publication. From the PIMMS methods the authors selected the variational encoder model for imputation and found 23.6% more differentially expressed proteins. They then investigated whether the differentially regulated proteins can be associated with disease using the DISEASE database, to find that 20 of these proteins had an association entry to fibrosis. With the newly found proteins the authors retrained the predictive model from the original study for liver condition development. They found that the retrained model performed as good or slightly better than the original model, concluding that these proteins can have predictive power.

Just like in the previous editions of PROTrEIN Journal Club we have sent our questions to the authors to conduct a short interview with them and gain more insight in their work. Please read our short interview below:

Blog post team: How self-supervised deep learning models that you chose contribute to the imputation of missing values in label-free quantification (LFQ) proteomics data? What advantages do they offer compared to other models?

Henry Webel: The models are in the category of machine learning models. In comparison to let’s say a random forest, you will additionally get embeddings of the features and samples, either separately (collaborative filtering) or joined (Autoencoder based architectures). Clustering in the embedding – also called latent – space could be used to compare it to e.g. hierarchical clustering of the original data. 

Blog post team: The paper suggests assessing if machine learning models can be trained on lower-level data. Could you elaborate on this suggestion and discuss the potential benefits and challenges associated with training machine learning models using lower-level data in the context of MS-based proteomics?

Henry Webel: In mass spectrometry- based bottom-up proteomics the unit of measurement are ions of peptides. The aggregation to protein groups therefore normally implies an implicit imputation using a neutral element as described by Lazar et al.4 Therefore, the elements of interest should be rather measured peptides, but currently the dominant approach is to use aggregated protein groups. 

Blog post team: It is mentioned that performance of the self-supervised models was better than heuristic approaches, which included median, interpolation or shifted normal distribution imputation, what was the main reason for that and could it be affected somehow by features of the data or size of the data? Do you think it is applicable to use supervised models instead of self-supervised models?

Henry Webel: The performance comparison always depends on the design of the comparison. We therefore extended the down-stream comparison in a revised version of the article. Additionally, we added supervised models such as random forests published as R packages. The main idea is to allow users to compare how well the imputation approaches perform on missing completely at random data (MCAR).

Blog post team: What are the future outlooks for PIMMS? Are you planning further developments?

Henry Webel: We added now many R methods to the comparison. Otherwise it would be great to extend the comparison by other methods of creating the validation and test data splits to get an overview of the methods used in the field.

References:

1. Webel, H. et al. Mass spectrometry-based proteomics imputation using self supervised deep learning. bioRxiv 2023.01.12.523792.

2. Jin, L. et al. A comparative study of evaluating missing value imputation methods in label-free proteomics. Sci. Rep. 11, 1760 (2021).

3. Niu, L. et al. Noninvasive proteomic biomarkers for alcohol-related liver disease. Nat. Med. 28, 1277–1287 (2022).

4. Lazar, C., Gatto, L., Ferro, M., Bruley, C. & Burger, T. Accounting for the Multiple Natures of Missing Values in Label-Free Quantitative Proteomics Data Sets to Compare Imputation Strategies. J. Proteome Res. 15, 1116–1125 (2016).

The post Mass spectrometry-based proteomics imputation using self-supervised deep learning appeared first on PROTrEIN.

]]>
Nanopore profiling: a scalable approach to protein identification https://protrein.eu/blog/nanopore-profiling-a-scalable-approach-to-protein-identification/ Mon, 22 May 2023 10:24:17 +0000 https://protrein.eu/?p=1838 Proteomics research relies heavily on mass spectrometry, which has emerged as the most prominent and widely utilized approach for the identification of proteins. However, progress is being made in other analytical methods for peptides and proteins. In this month PROTrEIN ITN’s Journal Club, we are taking a glance at the characterization of proteins with nanopores […]

The post Nanopore profiling: a scalable approach to protein identification appeared first on PROTrEIN.

]]>
Proteomics research relies heavily on mass spectrometry, which has emerged as the most prominent and widely utilized approach for the identification of proteins. However, progress is being made in other analytical methods for peptides and proteins. In this month PROTrEIN ITN’s Journal Club, we are taking a glance at the characterization of proteins with nanopores through a publication from the University of Groningen entitled “Protein identification by nanopore peptide profiling”1.

Nanopores are naturally occurring protein channels in cell membranes through which molecules can pass. Nanopore sequencing and profiling approaches make use of these structures to identify molecules. The nanopore is embedded in a membrane that separates two chambers filled with an electrolyte solution. An electrical potential is then applied across the membrane, creating a current that flows through the nanopore. The analytes that travel through the pore disrupt the ion current, which can then be measured (Fig. 1). The principle behind nanopore sequencing is based on the fact that different molecules (nucleotides in the case of nanopore sequencing) have different sizes and shapes, and therefore induce distinct changes in the ion current. Currently, nanopore technologies are the most commonly used for the sequencing of DNA, but progress is being made in their utilization for the identification of peptides and proteins. 

Fig. 1: Graphical overview of the nanopore protein fingerprinting approach. Peptides are pre-hydrolyzed by a specific protease (e.g. trypsin) and the resulting peptides are measured as they translocate the nanopore. Each peptide entering the nanopore reduces the open pore current (Io) to the blocked pore current (IB). The resulting excluded current (ΔIB = Io − IB) relates to the volume of the peptide. The subsequent histogram of the percent of excluded currents (Iex %= ΔIB/ IO %) is used to identify the protein. Source: Nat Commun (2021) 12, 5795

In the presented article the authors use nanopores to characterize individual proteins following their digestion by trypsin. After digestion, peptides go into the nanopore analyser and generate a change in ion current when they pass through the pore. The protein information is collected in an excluded current spectrum, which is a histogram that summarizes the translocation events of all the peptides. Since each change in the ion current relates mainly to the analyte volumes, the resulting spectrum is a representation of the peptide volumes after protease digestion and can be used for the identification of the original protein.

The identifications obtained with the nanopore analyses are compared to the results collected with ESI-MS (electrospray ionization mass spectrometry). For the comparison, the peptides identified through the mass spectrometry analysis are mapped into an inferred excluded current spectrum. This spectrum is generated through a computational calibration algorithm and compared to the nanopore-obtained one through spectral matching techniques. 

Based on the performed comparisons, the authors observe that the reproducibility of the obtained spectra is quite high. Although for specific proteins the correlation between the nanopore-observed spectrum and the inferred one is lower, this could be due to the computational inference of the spectra based on the peptide mass rather than the volume, which is key in nanopore analyses.In conclusion, the authors observe that although the nanopore analyzers might require improvements in terms of resolution, they could constitute low-cost solutions for performing high-throughput analyses. They especially point out that this technique could offer considerable advantages when having small amounts of material, as is the case of low-abundant proteins or heterogeneous ones. For this reason, mass spectrometry and nanopore techniques could even be used in parallel, as they show different and complementary strength when it comes to protein identifications.

Q&A with Florian Lucas, the first author of the publication

Since nanopore technologies can identify single molecules and do not require large amounts of material, do you foresee a wider use of nanopore peptide profiling in single-cell proteomics?

Nanopore profiling gains its strength from the low sample volumes and high sensitivity able to pick up, previously undetected, unidentified contaminants in our filtered water supply2. They may therefore find their way into some single-cell proteomic pipeline. However, their development is some years from mainstream application due to several engineering challenges remaining for the sensory technology. Nonetheless, these are inevitably solved in due time, as exemplified by commercial nanopore DNA sequencing devices.

More specifically, while single prokaryotic cells may provide currently unachievably small numbers of analytes, it is undeniable that nanopores can be used to detect proteolytic peptides from single eukaryotic cells. These can remain largely undiluted as the total measurement volumes of state-of-the-art nanopore chambers is less than 150 nanolitres, which is still amenable to down-scaling. However, the theoretical resolution of peptide profiling using nanopore will not be sufficient to extract single-protein information in such complex samples. Rather there are two future developments I foresee for the technology. Firstly, the nanopore can be coupled with upstream separation techniques such as liquid chromatography to reduce sample complexity. Secondly, motor proteins can (in theory) be attached on top of the nanopore to directly sequence full proteins (one-by-one), similar to nanopore DNA sequencing. This concept has been shown by the groups of G. Maglia and C. Dekker in two independent ways3, 4.

The current resolution of nanopore technology does not seem to allow for the detection of ‘small’ modifications (e.g methylation). Do you think this approach can one day reach that level of accuracy?

This is an interesting question, as the answer cannot be expressed with a simple yes or no. It is important to acknowledge that nanopores do not detect the mass, rather, the displacement of ions. This makes them ideal for the detection of modifications that alter the ionic current and will therefore be able to detect some, but not all, modifications. Methylation in particular is a modification we can detect using nanopore DNA sequencing, and there is no doubt that this modification on peptides can be detected. Moreover, our recent work shows that glycans can be observed on peptides5. We have also previously shown that conformational differences between leucine and isoleucine can be discriminated6, and our colleagues displayed the discrimination of phosphorylated peptides7.

In the article, only one protein digest is analyzed at a time. How feasible would it be to scale the technique to more complex samples, and what would be the main challenges?

The field is still in the proof-of-concept phase, which is best compared to mass spectrometry in the 1980s. Even mass spectrometry as a stand-alone technique is unable to extract all information from highly complex samples. Instead, it builds on the synergy between analytical separation, e.g. liquid chromatography (LC) or capillary electrophoreses. These methods amplify the separation power to allow complex analysis. Unpublished results show that we can have a steady flow of several milliliters per minute across the nanopore. This is more than compatible with downstream flows of analytical LCs. However, we cannot exclude the possibility of nanopores sequencing proteins one by one when coupled with a motor protein, but this is still in its theoretical and very early proof-of-concept phase3, 4.

References

1. Protein identification by nanopore peptide profiling. Florian Leonardus Rudolfus Lucas, Roderick Corstiaan Abraham Versloot, Liubov Yakovlieva, Marthe T. C. Walvoort and Giovanni Maglia. Nat Commun (2021) 12, 5795. DOI: 10.1038/s41467-021-26046-9

2. The Manipulation of the Internal Hydrophobicity of FraC Nanopores Augments Peptide Capture and Recognition. Florian Leonardus Rudolfus Lucas, Kumar Sarthak, Erica Mariska Lenting, David Coltan, Nieck Jordy van der Heide, Roderick Corstiaan Abraham Versloot, Aleksei Aksimentiev, and Giovanni Maglia. ACS Nano 2021 15 (6), 9600-9613. DOI: 10.1021/acsnano.0c09958

3. Bottom-up fabrication of a proteasome–nanopore that unravels and processes single proteins. Shengli Zhang, Gang Huang, Roderick Corstiaan Abraham Versloot, Bart Marlon Herwig Bruininks, Paulo Cesar Telles de Souza, Siewert-Jan Marrink, and Giovanni Maglia Nat. Chem. 2021 13, 1192–1199. DOI: 10.1038/s41557-021-00824-w

4. Multiple rereads of single proteins at single–amino acid resolution using nanopores. Henry Brinkerhoff, Albert S. W. Kang, Jingqian Liu, Aleksei Aksimentiev, and Cees Dekker. Science 2021 374, 1509-1513(2021). DOI:10.1126/science.abl4381

5. Quantification of Protein Glycosylation Using Nanopores. Roderick Corstiaan Abraham Versloot, Florian Leonardus Rudolfus Lucas, Liubov Yakovlieva, Matthijs Jonathan Tadema, Yurui Zhang, Thomas M. Wood, Nathaniel I. Martin, Siewert J. Marrink, Marthe T. C. Walvoort, and Giovanni Maglia. Nano Letters 2022 22 (13), 5357-5364. DOI: 10.1021/acs.nanolett.2c01338

6. In silico assessment of a novel single-molecule protein fingerprinting method employing fragmentation and nanopore detection. Carlos de Lannoy, Florian Leonardus Rudolfus Lucas, Giovanni Maglia, Dick de Ridder. iScience 2021 24 (10), 103202. DOI: 10.1016/j.isci.2021.103202.

7. Label-Free Detection of Post-translational Modifications with a Nanopore. Laura Restrepo-Pérez, Chun Heung Wong, Giovanni Maglia, Cees Dekker, and Chirlmin Joo. Nano Letters 2019 19 (11), 7957-7964. DOI: 10.1021/acs.nanolett.9b03134

The post Nanopore profiling: a scalable approach to protein identification appeared first on PROTrEIN.

]]>
Papers and patents are becoming less disruptive over time https://protrein.eu/blog/papers-and-patents-are-becoming-less-disruptive-over-time/ Mon, 03 Apr 2023 14:15:46 +0000 https://protrein.eu/?p=1785 The fields of science and technology have been the engines of progress for many years, driving innovation and shaping the world we live in. Contributing to this scientific knowledge and progress is one of the main aspirations most of us researchers have. However, there is a growing number of studies suggesting that scientific progress is […]

The post Papers and patents are becoming less disruptive over time appeared first on PROTrEIN.

]]>
The fields of science and technology have been the engines of progress for many years, driving innovation and shaping the world we live in. Contributing to this scientific knowledge and progress is one of the main aspirations most of us researchers have. However, there is a growing number of studies suggesting that scientific progress is slowing in several fields (1–3). In today’s journal club, we present the work of Ph.D. candidate Michael Park, prof. dr. Erin Leahy and prof. dr. Russel J. Funk, titled “Papers and patents are becoming less disruptive over time” published in Nature last month, where they explore this phenomenon by conducting a large-scale analysis of the innovative activity in science and technology (4). This study is based on 25 million papers from the years 1945-2010 from the Web of Science and 3.9 million patents of the United States Patent and Trademark Office from the years 1976-2010, to try to understand both the extent of the slowdown in innovation and the reasons behind it.

To begin, the authors based their analysis on the distinction of two types of breakthroughs: consolidating contributions and disruptive contributions. The former are works that improve and contribute to further establishing existing knowledge, while the latter challenge current understanding making it obsolete and therefore driving science and technology towards new frontiers. To quantify this characteristic, the authors relied on a metric called the CD index (5), which is based on the number of citations a paper or patent receives and how these citations relate to previous work in the field. As Park et al. best put it: “[…] if a paper or patent is disruptive, the subsequent work that cites it is less likely to also cite its predecessors […] If a paper or patent is consolidating, subsequent work that cites it is also more likely to cite its predecessors”. Figure 1 illustrates the concept of CD index and shows some examples.

Figure 1: This figure shows a schematic visualization of the CD index. a, CD index value of three Nobel Prize-winning papers and three notable patents in our sample, measured as of five years post-publication (indicated by CD5). b, Distribution of CD5 for papers from WoS (n = 24,659,076) between 1945 and 2010 and patents from Patents View (n = 3,912,353) between 1976 and 2010, where a single dot represents a paper or patent. The vertical (up–down) dimension of each ‘strip’ corresponds to values of the CD index (with axis values shown in orange on the left). The horizontal (left–right) dimension of each strip helps to minimize overlapping points. Darker areas on each strip plot indicate denser regions of the distribution (that is, more commonly observed CD5 values). c, Three hypothetical citation networks, where the CD index is at the maximally disruptive value (CDt = 1), midpoint value (CDt = 0), and maximally consolidating value (CDt = −1). The panel also provides the equation for the CD index and an illustrative calculation.

By measuring the CD index of each paper and patent at 5 years after publication, the authors were able to determine that there is a massive decline in disruptiveness in science and technology across all major fields in the last decades (decline 91.9-100% for papers, 78.7-91.5% for patents) (figure 2). The authors found the same trend also using other indicators, namely the linguistic composition of published titles and abstracts. For example, the type-token ratio (unique words to total words) of paper and patent titles has declined significantly, especially before 1970 for papers and 1990 for patents, and a similar decline was observed in the combinatorial novelty of the words used. Finally, there is a decrease in the novelty of the combinations of previous work cited by papers and patents.

Fig 2. Decline in CD5 over time, separately for papers (a, n = 24,659,076) and patents (b, n = 3,912,353)

The authors found that the decline in disruptive activity was not due to a decrease in the quality of science and technology or an artifact of the CD index itself. They observed similar patterns of decline in disruptiveness when they computed the CD index using other data sources, such as JSTOR and PubMed. They also found that the decline was not due to changing publication or citation practices by conducting additional analyses, such as regression models and Monte Carlo simulations. Park et al. also considered the relationship between the growth of knowledge and the decline in disruptiveness. While they found conflicting results, with a positive effect of the growth of knowledge on disruptiveness for papers and a negative effect for patents, they observed a decline in the use of previous knowledge among scientists and inventors. This suggests that scientists and inventors are increasingly focusing on narrower slices of previous work, which could be limiting the potential for disruptive discoveries and inventions.
In conclusion, this study provides evidence of a decline in disruptive activity in the fields of science and technology and suggests that this decline may be related to a decline in the use of previous knowledge among scientists and inventors. This highlights the importance of fostering an environment in which scientists and inventors can engage with a diverse range of knowledge, which is crucial for driving disruptive discoveries and inventions.

References

  1. B. F. Jones, The Burden of Knowledge and the “Death of the Renaissance Man”: Is Innovation Getting Harder? Rev. Econ. Stud. 76, 283–317 (2009).
  2. N. Bloom, C. I. Jones, J. Van Reenen, M. Webb, Are Ideas Getting Harder to Find? Am. Econ. Rev. 110, 1104–1144 (2020).
  3. J. S. G. Chu, J. A. Evans, Slowed canonical progress in large fields of science. Proc. Natl. Acad. Sci. 118, e2021636118 (2021).
  4. M. Park, E. Leahey, R. J. Funk, Papers and patents are becoming less disruptive over time. Nature. 613, 138–144 (2023).
  5. R. J. Funk, J. Owen-Smith, A Dynamic Network Measure of Technological Change. Manag. Sci. 63, 791–817 (2017).

Q&A with Michael Park

Michael Park, one of the authors of the publication, kindly agreed to answer some questions we had:

Ayesha and Marc: Your work suggests that the observed decline in disruptive scientific findings is possibly related to the “publish or perish” culture. A deviation from this would necessitate deeper reforms in policies and is unlikely to occur in the short term. What would be your advice, to us as a network of Ph.D. students, to mitigate the effect of the “publish or perish” culture and increase our chance to produce truly innovative findings?

Michael: I would just point out that this is beyond the scope of our study and the scientific community would benefit from further research on how researchers can better navigate the potential pitfalls of the «publish or perish» culture. Nevertheless, I am happy to share some of my casual thoughts on this. Although it can be difficult, I think trying to not be too limited in the breadth of your research pursuit is important. For example, staying informed about the latest research in not only your own specific field but adjacent fields as well could be helpful. Actively seeking out collaborations with scientists from other disciplines may also be enriching. Although there may be many reasons why disruption is decreasing across time, our paper does suggest that a narrow research focus limited to specific fields is closely linked to nondisruptive research. 

Ayesha and Marc: We are currently seeing major breakthroughs in the field of artificial intelligence and conversational models. Do you think such AI tools can play a catalyzing role to help scientists digest and combine the increasing stock of knowledge? A scaffold, of sorts, that would help us climb on the shoulders of ever taller giants?

Michael: Again, the question is outside of the scope of the paper as well as my main research area since I don’t specialize in AI research. Nevertheless, I am happy to share my non-expert opinion if you would like. I think AI may help make certain parts of the research process more efficient, such as literature review, data analysis, and writing. However, I think the idea generation aspect of research, which is most influential in determining the extent to which a piece of work is disruptive, will largely remain a human task. Therefore, unless there is some AI technology that is capable of producing disruptive new ideas (there may well be as you point out), I personally don’t think AI adoption in research and education will lead to an increase in disruptive discoveries and inventions.

Ayesha and Marc: A more philosophical question: Some support the idea that because humans are biological organisms, they have a specific scope and limits, and that includes their cognitive capacities (for example “What Kinds of Creatures are We?” by Noam Chomsky). What is your opinion on this? Is scientific and technological progress doomed to halt someday? Could the decline in disruptiveness shown in your study be an early indication of this phenomenon?

Michael: I do agree that a human being does have cognitive limitations. This is part of the reason why when researchers are forced to publish many papers quickly, they limit the scope of their specialty, which seems to be linked to the decline in disruptiveness. However, I don’t think the linkage between cognitive limitations and research scope suggests that we are near the «doomed» day yet. Although individual human beings may be limited in their cognitive capacity, science progresses due to the efforts of a community of people. Researchers feed off of the advances but also the limitations of others’ works. So I don’t necessarily think that the cognitive capacities of human beings are detrimental to the creation of disruptive work. In addition, we show in Figure 4 that the number of highly disruptive works is pretty consistent across time. This likely suggests that we are not «running out» of disruptive things to discover.

Below is Figure 4 of the article mentioned above.

This figure shows the number of disruptive papers (a, n = 5,030,179) and patents (b, n = 1,476,004) across four different ranges of CD5 (papers and patents with CD5 values in the range [−1.0, 0) are not represented in the figure). Lines correspond to different levels of disruptiveness as measured by CD5. Despite substantial increases in the number of papers and patents published each year, there is little change in the number of highly disruptive papers and patents, as evidenced by the relatively flat red, green and orange lines. This pattern helps to account for simultaneous observations of both aggregate evidence of slowing innovative activity and seemingly major breakthroughs in many fields of science and technology. The inset plots show the composition of the most disruptive papers and patents (defined as those with CD5 values >0.25) by field over time. The observed stability in the absolute number of highly disruptive papers and patents holds despite considerable churn in the underlying fields of science and technology responsible for producing those works. ‘Life sciences’ denotes the life sciences and biomedicine research area; ‘electrical’ denotes the electrical and electronic technology category; ‘drugs’ denotes the drugs and medical technology category; and ‘computers’ denotes the computers and communications technology category.

Cover Image source: https://ellipse.prbb.org/the-art-of-publishing-a-scientific-article/

The post Papers and patents are becoming less disruptive over time appeared first on PROTrEIN.

]]>
Changing the proteomics shell towards the DIA-world https://protrein.eu/blog/changing-the-proteomics-shell-towards-the-dia-world/ Mon, 27 Mar 2023 15:07:08 +0000 https://protrein.eu/?p=1769 Data-independent acquisition (DIA) proteomics increasingly becomes the method of choice for researchers since it provides better reproducibility, identification rates, and accuracy compared to data-dependent acquisition (DDA). More and more tools are developed for DIA analysis and even established proteomics data processing software now can analyze DIA data. However, analysis of multiplexed spectra, characteristic of DIA, […]

The post Changing the proteomics shell towards the DIA-world appeared first on PROTrEIN.

]]>
Data-independent acquisition (DIA) proteomics increasingly becomes the method of choice for researchers since it provides better reproducibility, identification rates, and accuracy compared to data-dependent acquisition (DDA). More and more tools are developed for DIA analysis and even established proteomics data processing software now can analyze DIA data. However, analysis of multiplexed spectra, characteristic of DIA, remains challenging. This becomes especially crucial once research has more unknown parameters to consider besides just protein sequences – for example in the case of phosphorylation site identification.

In a paper titled “Rapid and site-specific deep phosphoproteome profiling by data-independent acquisition without the need for spectral libraries” Bekker-Jensen et al. strive to develop a reliable approach for DIA phosphorylation data analysis. They started with their optimized instrument settings for DDA and devised in a similar manner the optimized workflow for DIA Then they compared the quantification accuracy and precision of each workflow with a mixed-species approach showing that DIA was able to identify twice as many phosphopeptides as DDA while still accurately estimating the ratios used in mixtures. Next, they developed a PTM localization algorithm specific to DIA and tested its performance on a set of synthetic peptides with known phosphorylation site localization. Again, with DIA they identified more sites identified with lesser error rates compared to DDA. They also adapted a machine learning-based approach for calculating phosphorylation site stoichiometry achieving better precision and accuracy with DIA in comparison with the standard DDA approach in a mixed-species experiment. Last but not least, they did several kinase inhibitor assays to test the workflow in biological conditions and found that the results are consistent with current knowledge.

Figure 1. Comparison of DDA and different types of DIA using a kinase inhibitor assay. a Experimental workflow. bOverview of identified phosphopeptides, localized phosphosites, and ANOVA (s0 = 0.1, FDR 0.5) regulated sites for the different methods. c Heatmap of unsupervised clustering analysis of ANOVA-regulated phosphosites for DDA workflow (d) and for DIA workflow with project-specific library (e). Linear sequence motif analysis for two major clusters marked in colored boxes on heatmaps.

The main reason for choosing this article is to share the promising method for handling DIA data with the other fellows. They proposed directDIA (dDIA) in the paper as an alternative to classical DIA analysis whose processing procedure is more complicated since it does not need an experiment-specific spectral library. Alternatively, dDIA may eliminate the low abundance issue of phospho-site rather than conventional DIA. The number of identified phospho-sites and the combinations of them across all experiments was approximately similar to the DDA method. From this perspective, we commented that this approach still has to be enhanced in terms of sensitivity. On the other hand, In DIA, obtaining more distinctive spectra would be accessible by increasing the number of windows to extend fractionation (sampling). This can be more achievable with a more recent version of the mass spectrometry instrument.

Taking advantage of the latest developments in Large Language Models (LLM), we include here a fictional interview with the authors of the article. Also, the cover image of this post was created using some keywords of the abstract and a text-to-image AI tool. Would you have guessed it shall reflect phosphoproteomics?

PROTrEIN: We are very excited to have with us today Dr. Dorte Bekker-Jensen, Dr. Oliver Bernhardt, Dr. Alexander Hogrebe, Dr. Ana Martinez-Val, Dr. Lynn Verbeke, Dr. Tejas Gandhi, Dr. Christian Kelstrup, Dr. Lukas Reiter and Dr. Jesper Olsen from the University of Copenhagen and Biognosys AG. Welcome!

LLM: Thank you for having us.

PROTrEIN: We really enjoyed reading your paper and we have some questions for you. First of all, can you tell us what motivated you to develop this DIA-based phosphoproteomics method?

LLM: Sure. As you know, protein phosphorylation is one of the most important post-translational modifications that regulates cellular signaling networks. However, it is also very challenging to analyze phosphoproteomes on a large scale using conventional data-dependent acquisition (DDA) mass spectrometry, which relies on stochastic peptide selection and requires extensive spectral libraries for identification and quantification. We wanted to overcome these limitations and enable high-throughput phosphoproteomics with minimal sample preparation and analysis time.

LLM: That’s right. We decided to use data-independent acquisition (DIA) mass spectrometry, which is a more systematic and reproducible way of acquiring MS/MS spectra for all peptides in a given mass range. DIA has been successfully applied to proteomics and metabolomics, but not so much to phosphoproteomics, mainly because of the complexity and diversity of phosphopeptides and the lack of accurate site localization tools for DIA data.

PROTrEIN: Interesting. So how did you solve these challenges?

LLM: Well, we developed a new workflow that combines efficient phosphopeptide enrichment, fast LC-MS analysis, and advanced computational analysis using Spectronaut software. We optimized the DIA settings to achieve high resolution and sensitivity for phosphopeptides, and we incorporated an accurate site localization score into Spectronaut that uses fragment ion intensities and retention time information to assign phosphorylation sites confidently.

LLM: We also compared different strategies for generating spectral libraries for DIA analysis, such as using public databases, project-specific DDA runs, or direct DIA without any library at all. We found that direct DIA performed close to project-specific libraries, quantifying more than 20,000 phosphopeptides in 15 minutes single-shot LC-MS analysis per condition.
Pinar: Wow, that’s impressive. And how did you validate your method?

LLM: We applied our method to various biological systems, such as HeLa cells stimulated with epidermal growth factor (EGF), mouse liver tissue, and human plasma samples. We showed that our method can identify and quantify thousands of phosphorylation sites with high reproducibility and accuracy across different sample types and instruments.

LLM: We also demonstrated the ability of our method to determine phosphorylation site stoichiometry using a 3D multiple regression model-based algorithm that corrects for peptide abundance variations in DIA data. This allows us to measure the fraction of phosphorylated molecules at each site and compare them across different conditions.

PROTrEIN: That’s very cool. And what did you learn from applying your method to EGF signaling?

LLM: We used our method to systematically analyze the effects of 30 kinase inhibitors on EGF signaling in HeLa cells. We quantified more than 25,000 phosphorylation sites across 900 conditions in less than two weeks of LC-MS analysis time. We identified hundreds of kinase inhibitor targets and their downstream effects on EGF-regulated phosphorylation sites.

LLM: We also discovered some unexpected interactions between kinase inhibitors and EGF signaling pathways, such as the cross-talk between PI3K/AKT/mTOR and MAPK/ERK pathways, or the feedback activation of EGFR by some inhibitors. These findings reveal new insights into the complexity and dynamics of EGF signaling network.

LLM: We think that our method can be applied to other signaling systems and drug discovery projects, as well as other PTMs such as ubiquitination or acetylation. We also plan to further improve the speed and sensitivity of our method by using novel MS instruments and data analysis algorithms.

PROTrEIN: That sounds very exciting. Thank you so much for sharing your work with us today. It was a pleasure to talk to you.
We hope our readers enjoyed this episode and learned something new.

The post Changing the proteomics shell towards the DIA-world appeared first on PROTrEIN.

]]>
Are you a morning person? – Chronotypes, Circadian Rhythms, and Questionnaires https://protrein.eu/blog/are-you-a-morning-person-chronotypes-circadian-rhythms-and-questionnaires/ Wed, 22 Feb 2023 14:41:04 +0000 https://protrein.eu/?p=1753 Life as we know it has been shaped by the constraints of the environment. The aerodynamic shape of leaves to prevent the tree from toppling at high winds, or gravity that affects heights of organisms (extraterrestrial humanoids on the fictional pandora in Avatar), or the color of a polar bear. The environment is a critical […]

The post Are you a morning person? – Chronotypes, Circadian Rhythms, and Questionnaires appeared first on PROTrEIN.

]]>
Life as we know it has been shaped by the constraints of the environment. The aerodynamic shape of leaves to prevent the tree from toppling at high winds, or gravity that affects heights of organisms (extraterrestrial humanoids on the fictional pandora in Avatar), or the color of a polar bear. The environment is a critical factor that shapes life. For us humans, along with our living companions on earth, this includes revolving and rotating around a single star, with a single moon. This has enabled us to develop several rhythms, such as circadian cycle (24 hours), ultradian cycle (less than 24 hours), and infradian rhythms such as the human menstrual cycle have periods longer than a day. There are clocks within us, in every organ, every cell. At the cellular level, these are just chemical clocks that play with protein concentrations.

Circadian rhythm is not only important to understand better human pathologies or to find new potential therapeutic candidates, but may also help in personalized medicine. Knowing that the immune system and hepatic metabolism change throughout the day, it is possible to study at which moment a specific drug can be given to a specific patient in order to maximize its effect or reduce its toxicity. Unfortunately the inner clock’s timing can not be generalized, each individual has a different rhythm that depends both on the lifestyle and genetic predisposition [1]. Moreover the rhythm can be totally disrupted, causing each organ’s clock not to be in phase with the body’s one and/or with the day-night cycle. Several factors, especially now with humans’ modern lifestyle, may induce such disruption: examples are the jet-lag, jobs that require frequent night shifts or any kind of stress that precludes a person from having a constant resting-active state cycle [2].

One possible way to investigate a person’s circadian rhythm is through the use of questionnaires. They are used to categorize patients in chronotypes through questions about normal day habits like at what time the subject wakes up, has meals, goes to sleep and how they affect his/her life [3]. Through such data it is possible to determine if the person is an early bird or a night owl (which are called chronotypes) and with this information it is possible to say that, for example, the last type will have its active phase shifted towards the evening. As it is understandable there are several limitations: this approach must rely on what the subject says and do not allow to say with certainty if there is a circadian rhythm disruption.

In this month’s journal club, we decided to take a dive into the circadian rhythm, chronotypes, and the molecules that affect them. We also took the opportunity to develop a web app that helps you figure out your chronotype and what it means for your health, personality, etc. We turned to literature for connections between chronotypes and health, and we zeroed in on a paper by Fabbian et al., 2016 [5]. 

We decided to create an online questionnaire that with simple questions and a nice interface is able to predict the user’s chronotype. The idea of such a questionnaire is to find a way to make data gathering more interesting for those who have to respond to questions, translating a simple questionnaire into a game that in the end may give some information. Such information about chronotypes is given using simple terms in order to be understandable even without any scientific knowledge.

The questions used are from the Horne and Östberg Questionnaire [4] and information about the chronotypes were obtained from the paper F. Fabbian et al. 2016 [5]. There were many other papers that were referenced for this journal club, but these two were instrumental to informing the application design.

Chronotypes

There are several chronotypes that have been theorized by scientists, but for ease we decided to select only three. The selected ones are:

Early Birds/Lack. People having this chronotype tends to go to sleep early in the evening (between 9/10 PM) and wake up very early in the morning (7 AM or before). The peak of activity and energy is reached at the morning, while during the rest of the day they will become more and more tired.

Eagles. This chronotype is a mixture of the other two: peak of activity is reached after noon and they usually go to sleep between 10 and 11 PM and wake up between 8 and 9 AM.

Night Owls. Such chronotype describes people that usually stay awake at night until 12PM or later and wake up in the morning after 9 AM. The peak of activity is reached at the late afternoon, making these people feel less energetic when they wake up while becoming more active and efficient as the day progresses.

Circadian rhythm regulation

To understand better what a chronotype means it is necessary to explain before how the circadian rhythm is regulated, both at the systemic and molecular level.

At the systemic level the main center of regulation is situated in the brain, in particular is a region called the suprachiasmatic nucleus (SNC) in the amygdala. These neurons receive inputs from the whole brain and in particular from eyes and through them they are able to determine if it is day or night [6]. The SNC uses this information to regulate several body functions like animal’s behavior and psychology, body temperature, metabolism and body temperature. All such functions oscillate during the day and do not need inputs from the SNC to do so, although its main function is to reset the whole body circadian clock to ensure that all functions are in phase together and with the environment [7]. To ensure that each organ and tissue is synchronized, there are several genes embedded in the genome that regulate the rhythm in each cell at the molecular level. CLOCK and BMAL1 are the two core genes coding for the two homonym transcription factors, which form a complex that enhances transcription of several other genes [8].

Simple representation of how BMAL1-CLOCK regulate themselves through other transcriptional factors (Courtesy of Pickel L. et Sung H. K. 2020)

CLOCK-BMAL1 complex promotes CRY-PER factors production, which then acts on CLOCK and BMAL1 promoters suppressing their expression, suppressing then indirectly also their own expression and restarting the cycle again. This negative feedback loop is the main mechanism that allows CLOCK-BMAL1 expression to be cyclical with a period of around 24 hours [9]. Other genes promoted by CLOCK-BMAL1 complex are those coding for the nuclear receptors REV-ERB and ROR, which then regulate the expression of several genes involved in cell metabolism and other vital functions. Moreover, REV-ERBα acts also as a suppressor for BMAL1, while RORα enhances it, showing redundancy in the system that regulates core clock genes [10]. An interesting fact is those core clock genes are the same for each cell type, although they regulate a wide number of genes that are different for each tissue. This is in line with the fact that during specific phases of the day organs and tissues may be more or less active, so CLOCK-BMAL1’s effect can not be the same for each cell population. It is although not clear how this exactly happens: some studies found that CLOCK-BMAL1 activates different transcription factors for each cell type, while others suggest that clock genes can activate different promoters thanks to complexes with tissue-specific transcription factors [11, 12].

Considering what said so far, it is possible to reconsider each chronotype as the moment in which a person has the peak of CLOCK-BMAL1: for early birds it will be between late morning and noon while for late owls will be more shifted toward the afternoon. It is not clear although how much a chronotype is determined by genetic predisposition or environmental factor, probably a mixture of both. 

Several studies found strong correlation between night owls and detrimental behaviors, at the point that the chronotype can be considered a risk factor for several pathologies. One important thing to point out is that such correlation may be caused by other factors, like for example circadian rhythm disruption, since night owls’ life-style may conflict with modern world working hours and everyday life [13]. Moreover psychiatric disorders such as depression commonly cause difficulties in falling asleep and waking up early in the morning. This obviously does not mean that such people have an evening-type chronotype, even though through a questionnaire it may seem so.

References
1. Vitaterna MH, Takahashi JS, Turek FW. Overview of circadian rhythms. Alcohol Res Health. 2001;25(2):85-93. PMID: 11584554; PMCID: PMC6707128.
2. Eastman, C. I., Tomaka, V. A., & Crowley, S. J. (2016). Circadian rhythms of European and African-Americans after a large delay of sleep as in jet lag and night work. Scientific Reports, 6(1), 36716.
3. Zavada, A., Gordijn, M. C. M., Beersma, D. G. M., Daan, S., & Roenneberg, T. (2005). Comparison of the Munich Chronotype Questionnaire with the Horne‐Östberg’s Morningness‐Eveningness score. Chronobiology International, 22(2), 267–278.
4. Horne, J. A., & Östberg, O. (1976). A self-assessment questionnaire to determine morningness-eveningness in human circadian rhythms. In International Journal of Chronobiology (Vol. 4, pp. 97–110). Gordon and Breach Science Pub Ltd.
5. Fabbian, F., Zucchi, B., De Giorgi, A., Tiseo, R., Boari, B., Salmi, R., Cappadona, R., Gianesini, G., Bassi, E., Signani, F., Raparelli, V., Basili, S., & Manfredini, R. (2016). Chronotype, gender and general health. Chronobiology International, 33(7), 863–882.
6. Hastings, M. (1998). The brain, circadian rhythms, and clock genes. BMJ, 317(7174), 1704–1707.
7. Moore, R. Y. (2007). Suprachiasmatic nucleus in sleep–wake regulation. Sleep Medicine, 8, 27–33.
8. Trott, A. J., & Menet, J. S. (2018). Regulation of circadian clock transcriptional output by CLOCK:BMAL1. PLOS Genetics, 14(1), 1–34.
9. Yu, W., Nomura, M., & Ikeda, M. (2002). Interactivating Feedback Loops within the Mammalian Clock: BMAL1 Is Negatively Autoregulated and Upregulated by CRY1, CRY2, and PER2. Biochemical and Biophysical Research Communications, 290(3), 933–941.
10. Guillaumond F, Dardente H, Giguère V, Cermakian N. Differential Control of Bmal1 Circadian Transcription by REV-ERB and ROR Nuclear Receptors. Journal of Biological Rhythms. 2005;20(5):391-403.
11. Kondratov, R. V, Shamanna, R. K., Kondratova, A. A., Gorbacheva, V. Y., & Antoch, M. P. (2006). Dual role of the CLOCK/BMAL1 circadian complex in transcriptional regulation. The FASEB Journal, 20(3), 530–532.
12. Qu, M., Duffy, T., Hirota, T., & Kay, S. A. (2018). Nuclear receptor HNF4A transrepresses CLOCK:BMAL1 and modulates tissue-specific circadian networks. Proceedings of the National Academy of Sciences, 115(52), E12305–E12312.
13. Togo, F., Yoshizaki, T., & Komatsu, T. (2022). Interactive effects of job stressor and chronotype on depressive symptoms in day shift and rotating shift workers. Journal of Affective Disorders Reports, 9, 100352.

The post Are you a morning person? – Chronotypes, Circadian Rhythms, and Questionnaires appeared first on PROTrEIN.

]]>
Casanovo, a transformer model to identify De novo mass spectrometry peptide sequencing https://protrein.eu/blog/casanovo-a-transformer-model-to-identify-de-novo-mass-spectrometry-peptide-sequencing/ Fri, 10 Feb 2023 11:49:22 +0000 https://protrein.eu/?p=1740 In the last Journal club, we present a paper by Yilmaz et al. called «De novo mass spectrometry peptide sequencing with a transformer model» [1] introducing a deep learning model for de novo peptide sequencing. What? You do not know exactly what is de novo peptide sequencing? Let me explain it. Imagine that you do […]

The post Casanovo, a transformer model to identify De novo mass spectrometry peptide sequencing appeared first on PROTrEIN.

]]>
In the last Journal club, we present a paper by Yilmaz et al. called «De novo mass spectrometry peptide sequencing with a transformer model» [1] introducing a deep learning model for de novo peptide sequencing.


What? You do not know exactly what is de novo peptide sequencing? Let me explain it. Imagine that you do not have enough prior knowledge about your sample. How can you use database search methodology? Under this condition, we try to identify peptide sequences directly from experimental spectra. The principle of it is to find the specific fragmentation pattern based on the regular breaks in the mass spectrometric detection of the peptide molecules after protease cleavage and calculate the corresponding amino acid information according to the mass difference between the mass spectrum peaks as well as the post-translational modifications of the amino acid.

Figure 1: Casanovo performs de novo peptide sequencing. Source: Yilmaz et al., 2022 [1]

Early de novo methods used the heuristic search or dynamic programming to score peptide sequences. Recently, some efforts have been made to develop Deep learning models to predict peptides sequence from MS2 including DeepNovo, SMS, and PointNovo. However, these models include complex post-processing steps. In addition, their structures are based on recurrent neural networks, which are slow to train and suffer from long dependency issues.
To resolve mentioned drawbacks, they proposed a transformer-based model called Casanovo for de novo peptide sequencing. Casanovo uses the self-attention mechanism to translate from a variable-length sequence of observed spectrum peaks to a variable-length sequence of amino acids, analogous to the neural machine translation model in the natural language processing setting. o consists of a transformer encoder and decoder, where the encoder takes d-dimensional spectrum peak embeddings as input and outputs d-dimensional latent representation vectors.
Casanovo was trained by 30 million labeled spectra which contain Seven different types of variable modifications (methionine oxidation, asparagine deamidation, glutamine deamidation, N-terminal acetylation, N-terminal carbamylation, N-terminal NH3 loss, and the combination of N-terminal carbamylation and NH3 loss).

Figure 2: Casanovo architecture. Source: Yilmaz et al., 2022 [1]

To evaluate the performance of Casanovo and compare it with other de novo peptide sequencing models, they used the nine-species benchmark data set. This data set combines a total of about 1.5 million mass spectra from nine different experiments, each using the same instrument to analyze peptides from a different species.

Casanovo, leverages the transformer architecture to produce a unified solution to translate mass spectra directly into peptide sequences, without resorting to the discretization of the spectrum m/z axis and without complex post-processing.

We had an interview with Melih Yilmaz and asked him the two following questions about the paper:

Do you think what is the most challenging issue to have a better deep model for de novo sequencing?

Thanks for reaching out and for your interest in Casanovo! A challenge that we tried to overcome with our new preprint was that the original version of Casanovo was trained on MS data from peptides digested with trypsin enzyme which didn’t perform as well for samples that were digested using a different enzyme. To mitigate this, we fine-tuned the existing Casanovo model on a non-enzymatic data set which significantly improves performance on non-tryptic data.

Do you have any plan to improve your model by considering more modifications in your training data?

We don’t have short-term plans to increase the number of post-translational modifications in the current version model. However, it would be straightforward to fine-tune the current model with an extended training set containing the new modifications.

References:

[1] Yilmaz et al., bioRxiv (2022), doi.org/10.1101/2022.02.07.479481

The post Casanovo, a transformer model to identify De novo mass spectrometry peptide sequencing appeared first on PROTrEIN.

]]>
Recalling memories of our first in-person project meeting and SciComm training school in the beautiful city of colours ‘Barcelona’ https://protrein.eu/blog/recalling-memories-of-our-first-in-person-project-meeting-at-the-beautiful-city-of-colors-barcelona/ Mon, 12 Dec 2022 16:38:20 +0000 https://protrein.eu/?p=1689 Consortium meeting: After we ESRs had participated in the 13th International MaxQuant Summer School on Computational Mass Spectrometry-based Proteomics, the opening dinner of the consortium meeting was the first time that all PROTrEIN members, incl. supervisors, actually met in person. The venue was amazingly beautiful and close to the sea. Where we all enjoyed a […]

The post Recalling memories of our first in-person project meeting and SciComm training school in the beautiful city of colours ‘Barcelona’ appeared first on PROTrEIN.

]]>
Consortium meeting:

After we ESRs had participated in the 13th International MaxQuant Summer School on Computational Mass Spectrometry-based Proteomics, the opening dinner of the consortium meeting was the first time that all PROTrEIN members, incl. supervisors, actually met in person. The venue was amazingly beautiful and close to the sea. Where we all enjoyed a delicious dinner, with the added bonus of a beautiful night time sea view.

A stunning sea view with a full moon that ESRs enjoyed was next to our dinner location.

ESRs and supervisors attended the consortium meeting on September 12 and 13. The meeting featured presentations of the various work packages, keynote addresses, reports, and a meeting of the supervisory board. The «Management and Coordination» presentation by the PROTrEIN coordination team opened the work package presentations. An overview of previous sessions, including «scientific project planning,» midterm check, and summer school, was given at the beginning of this presentation. It went on to talk about internal reporting before citing the impending deadlines. Lastly, a little overview of Ghent Winter school. The nicest part of the meeting was the speed dating session when all ESRs had been given a chance to briefly introduce themselves to other project supervisors.
The ESRs were well-prepared the following day to provide an update on the status of their research work. Many of us ESRs were rather anxious, not only because of the supervisors, but also because of the «time machine» that was set up to inform us of the allocated time. Nevertheless, everything went smoothly, and we received many useful ideas from the other supervisors to enhance our direction.


«Novel machine learning predictors» was the first work package presented by five of the ESRs. The next six ESRs presented their works about «New algorithms for mass spectrometry raw data processing». This series of presentations was completed in the afternoon with «Integration and visualization of omics data» presented by the final four ESRs.

Sharing one of the memories of our consortium meeting when Shamil was giving updates on his project and while most of us seemed at ease as they finished, several of us were feeling anxious as we were waiting for our turn.

After a long day, Jonas had planned a surprise cooking workshop for us. It was a great team-building experience where we made tasty and colorful tapas. We were instructed by two very nice chefs, they divided us into two groups of warm and cold tapas, and of course, they didn’t forget the vegetarian options. The chefs first provided us with instructions and recipes, which we then heartily enjoyed. Our favorite tapas were the salmon and avocado tartar with pistachios mayo, and cod fritters with honey allioli. Even thinking about those fantastic tapas right now makes us want to go back and try everything over and over again.

The image serves as proof that the PROTrEIN team is capable of doing anything as a team, including cooking, in addition to being excellent researchers.
Sharing the memorable picture of the dinner in which it is clear  that Arthur was thoroughly enjoying his food and didn’t give a damn about the camera.

The second day of the consortium meeting started with keynote lecturer Juan Antonio’s presentation on open-research and data reusability. The day was followed by presenting the remaining three work packages and discussing the upcoming winter school. The consortium meeting ended with supervisory board and coordination meetings, and all ESRs went to ELISAVA to start their science communication course.

A group photo of the PROTrEIN team following a productive consortium meeting.

Science Communication workshop

Design thinking for scientists was the first module of the workshop on science communication, and when we were asked at the beginning of the course “What we are expecting from this course?”, most of us thought it would be some spoken practice for our project presentations and visuals. However, it turned out to be much more than that.

Blanca Guasch steps in to explain how to apply design techniques to make our research more fruitful and efficient.  Design thinking involves several steps, including Empathize, Define, Ideate, Prototype, Test, and Implement. We created individual collages to illustrate the design thinking module Empathize, describing ourselves imaginatively with the aid of printed images, and describing how we are contributing to the PROTrEIN project.

Metaphors used by the ESRs

Another interesting exercise that we ESRs had done is to put our names on the PROTrEIN WPn map according to the importance and relation of our contributions to each module of the design thinking process. Then, for the Ideate section, we discover the links between various ESRs and how they relate to one another in terms of project goals, secondments, and research methodology.

The ESRs created a WPn map to show the connections between their projects

With the help of different metaphors, ESRs described the work bundles and processes like explaining machine learning, cross-linking concepts, etc.  After expressing our analogies to the group, we took into consideration their responses. After this engaging session 1, we all had some constructive ideas to use in our daily lives.

In the picture Zoltan was trying to find metaphors for crosslinking

On the second day of the course on science communication, Blanca began the session with «Presentation designs for scientists,» another fascinating and crucial module for ESRs. This lesson focused on speaking well, taking the initiative, and being creative in our presentations. To keep things simple, avoid presenting too much information at once, emphasize the importance of consistency and typography, employ various styles according to the situation, and utilize a variety of formats and resources. 

Ane Guerra started a fascinating discussion on science communication storytelling and explaining stories. We completed some exercises as part of the story-telling process by describing the major difficulty we encountered when describing our project and the qualities that stand us apart from the others. Also, most of us ESRs defined our project as if we were a superhero, along with its foes, strengths, and nemesis. Being able to connect your project with some superpower heroes was, in my opinion, not an easy task, but most ESRs were able to accomplish it creatively. Then, we got some tips for developing story-telling skills in our research, including the use of a clear message, consideration of the audience, and the communication context. 

In the explaining stories session, we discovered that we should be aware of the audience when discussing our research and should stick to the core idea the entire time. 

“The best storytellers deliberately listen, watch, and read.” 

The session concludes with some helpful advice on how to improve our public speaking abilities, including knowing your audience, taking breaks, dressing appropriately, using our hands/voices effectively, and soliciting feedback.

ESRs along with their mentors were busy doing their group project

Another relevant and interesting session for us ESRs is data visualization which is to embrace scientific complexity and transform it into a sympathetic visual story that improves all of its best features. 

We then discussed the process of mapping facts to visual structures, known as visual encoding. After that we went through how people in ancient times shared information about data I-e, the 1945 Molecular model of penicillin by Dorothy, and also discussed some good books which are based on data and visualizations.

Visual design lecture by Francesc Ribot ,the coordinator of the Graphic Design Area

In the last section of the course, we ESRs demonstrated our creativity by writing a video script while taking the public and expert audiences into consideration. A script for each audience had to be written according to a set of rules and to last for 60 seconds. Furthermore, we ESRs had been given the opportunity to present our posters and received feedback and discussion from our fellows and mentors. 

Last but not least, Jonas Krebs, all the mentors and the funding body deserves praise for organising such a relevant and informative science communication course for us ESRs. From this course, we have learned that better communication abilities enable researchers to share their discoveries with a wider audience and strengthen links within their scientific groups.

The post Recalling memories of our first in-person project meeting and SciComm training school in the beautiful city of colours ‘Barcelona’ appeared first on PROTrEIN.

]]>
Wikipedia Hackathon Experience https://protrein.eu/blog/wikipedia-hackathon-experience/ Mon, 28 Nov 2022 11:36:35 +0000 http://protrein.eu/?p=1655 Most of us do research because we enjoy finding answers, solving problems and even helping people. But there is also a deep seated human tendency to leave a mark. Contribute to a bigger picture, graffiti on your neighbor’s wall, advance a field. This is probably why we write research papers. (Apart from begging for funding […]

The post Wikipedia Hackathon Experience appeared first on PROTrEIN.

]]>
Most of us do research because we enjoy finding answers, solving problems and even helping people. But there is also a deep seated human tendency to leave a mark. Contribute to a bigger picture, graffiti on your neighbor’s wall, advance a field. This is probably why we write research papers. (Apart from begging for funding and completing a PhD) 😛

But the general public cares less about research papers and is more interested in blog posts such as «Is your cow cheating on you?».

But then, there is a that one moment when you have to win a bet against your friend who thinks that «Cows can hear infrasonic sounds» and you need to open Wikipedia (*coughs* Microsoft Bing) to settle it. Or that time you had to write an assignment on «Cow hybrids”. Yep. One of the authors of this blog is obsessed with cows.

Anyways, Wikipedia to this day remains an important repository of high quality articles on almost any topic you can think of. To this end, we participated in a Wikipedia Hackathon that would facilitate addition of proteomics articles to the growing Wiki database.

So what is that we did on Wikipedia?

All PROTrEIN network early stage researchers participated in the wikipedia hackathon. We didn’t just learn about articles, but also approaches to publishing open access data to Wiki data.

Along with the mentor, Toni Hermoso, we created a web page in wikimedia entitled PROTrEIN Editathon (https://meta.wikimedia.org/wiki/PROTrEIN_Editathon_2022) we learned mainly how to edit into a wikipedia page starting from creating paragraph and subparagraphs into adding images into wiki images and then use it for our articles, we created multiple proposals for related terms and articles to the field of proteomics and mass spectrometry.

One article that we wanted to feature as an example is a Maxquant one. So as you see in the following figure, We created a wikimedia image for the software Maxquant, where we added the article for defining Maxquant alongside a figure for the software logo and viewer interface, we also added links and dates.

We learned how to create links, it works similar to hyperlink on the Microsoft Word processor. Alongside to citing, with automatic citation an editor can simply copy-paste the URL, and it will generate a citation. Whereas, manual citation requires the editor to enter the details of a book, journal, article, or website. After making edits, it is time to leave an edit summary about the changes that have been made to the wiki page.

This Wiki Hackathon took place within the framework of Computational Proteomics MaxQuant Summer School 2022 on September 5-6, 2022.

The post Wikipedia Hackathon Experience appeared first on PROTrEIN.

]]>
Organising a large event as a first year PhD student – behind the scenes of the MaxQuant Summer School 2022 https://protrein.eu/blog/organising-a-large-event-as-a-first-year-phd-student-behind-the-scenes-of-the-maxquant-summer-school-2022/ Mon, 21 Nov 2022 15:46:13 +0000 http://protrein.eu/?p=1601 Starting with the premise that I have never organised such a big event in the past, getting actively involved in the organisation of the MaxQuant summer school (MQSS) has been a great opportunity to learn skills that are usually not directly related to pure research in science. The whole process of organising the summer school […]

The post Organising a large event as a first year PhD student – behind the scenes of the MaxQuant Summer School 2022 appeared first on PROTrEIN.

]]>
Starting with the premise that I have never organised such a big event in the past, getting actively involved in the organisation of the MaxQuant summer school (MQSS) has been a great opportunity to learn skills that are usually not directly related to pure research in science.

The whole process of organising the summer school can be summarised in four different steps:
1) finding the venue to host the event
2) advertising the event
3) registration and booking
4) setting up the event

Finding the venue to host the event

In January this year, I joined the summer school organisation committee, which was composed of my supervisor Jürgen Cox, one post-doc and another more experienced PhD student that has already organised summer schools in previous years. When we discussed possible locations of the event, it was immediately decided to go to Barcelona, not only because of the sun, but also due to the logistic easiness (the MQSS took place there already once in 2018). Besides, Jürgen’s lab has contacts there (Eduard and Jonas from CRG) that were of great help. The initial phase consisted mostly of meetings with people from the agency that helped us with the organisation (Crea Congresos) and to decide when and where to have it. Once a venue that fulfilled the main requisites was found the real organisation from our side started and I got more involved. Such main criteria for the decision were: enough space for 200 participants and their posters, audio and video appliances and technical support and lastly being easy to reach. In this part most of the work, like visiting different venues and being sure of which services were ensured, was done by our counterpart in Barcelona. We were lucky that with their help we could convince the «Centre de Cultura Contemporània de Barcelona» (CCCB) to host the event one more time, same as in 20218.

Advertising the event

The first step was to design a nice logo and then set up the webpage. These were my initial tasks. I took this opportunity also to practice my skills in html and Adobe Illustrator. In particular, designing the new logo was something I enjoyed a lot, since I have always drawn and liked art. In other words, it was a great way to apply art to science.

The final logo of the summer school

Once the drafts were approved and the website was set up, it was time to start advertising the summer school through social media (e.g. Cox lab’s and Max Planck’s twitter account) and through sending emails to all previous years’ participants. In the meanwhile, the five main speakers were found and a first conference program drafted, both done in collaborative efforts. Lastly, Crea Congresos prepared the registration form and opened the bank account where to receive the conference fees from participants.

Registration and booking

As soon as we had set the registration deadline on mid May, the priority went to prepare the budget. All expenses were calculated by Crea Congresos, we checked them and added a buffer for eventual last moment expenses (an issue that indeed happened). On top of that, taxes were added and then 170 was set as the minimum required number of participants in order to cover all expenses. The registration fee for each person was calculated accordingly and the registration form opened. Assisting with the budget report was something totally new to me, but also helpful to get an idea of the costs that such event produces. Honestly, I have never dealt before with such an amount of money in my life. From this point on, we spend most of the time with taking care of public relations, since there was an average number of 8-10 emails per day from people asking for more information. The date to close the registration was at the end postponed to August. After that, we could proceed with booking the restaurant and social activities. Once we reached the second half of August, all participants that registered were contacted again in order to know if they wanted to present a poster and which social activity they wanted to join.

The week before the event, the spreadsheets with all this information were sent to Crea Congresos to let them know about the exact numbers. We organised all necessary documents and data (like MaxQuant, Perseus, tutorials and practical exercises) and transferred them to individual USB pen drives that we planned to give to each participant. Lastly, a new Zoom subscription was purchased in order to allow an online participation of the summer school and a new registration form for online participants was opened.

Checking that everything works before the start

Setting up the event

Just few days before the start of the summer school, it was time to move to Barcelona to set up the venue and test appliances to be sure everything was ready for Monday. From that moment on, luck was not very much on our side. One small issue was that there has been some delays in the preparation of the USB sticks that we planned to give to attendants, but that was solved easily through all collaborative efforts during the weekend before. The bigger issue was a strike that caused flights to Barcelona to be cancelled, all except mine. This has been quite a big deal since on the Friday before the start we were supposed to visit the venue and test all the equipment. At the end we managed to solve everything thanks to the help of the people from Barcelona and the technicians, and all was set and ready for the event to start. The flight issue was exactly one thing we were worried about, especially because it was something we could not do much about it. At the end, it was a good reminder of Murphy’s law: “anything that could go wrong will go wrong”. Lastly during the weekend also the rest of the lab managed to arrive and we all prepared the material and rehearse our talks.

Finally it was time for the summer school to begin and about this I will not write much more since everyone from the PROTrEIN-ITN was there. The final program can be found on the MQSS webpage.

Hamid enjoying not being an organizer for one time

Honestly, there has been also a particular moment during which I cursed being one of the organizers and it was during the joint dinner on Wednesday. That day we kept an easy schedule for participants: social activities like tours or sport in the afternoon and then a dinner all together in a nice restaurant close to the sea. Unfortunately, my talk was exactly the next morning and there were some more data needed by participants that we could not give through the usb sticks. This meant that during the free afternoon I had to prepare my talk and then had to leave the dinner earlier in order to prepare the email that was sent to all participants with the aforementioned data. In other words I missed a nice evening and partying while some of my colleagues did not. Have a look at the picture above to get what I mean, for the record I was writing emails at that moment. Well, at least I did not have a hangover the day after.

In summary, it has been a nice experience, even though doing everything took a lot of time and there has been some stressful moments in particular during the last days. It was also a good opportunity to meet other researchers working in proteomics and to get a feedback from them regarding our general work. Lastly, when the event was over, it was also a great satisfaction to have been part of the organisation and seeing that participants appreciated it and had a good time during the week.

Celebrating the event to be over

The post Organising a large event as a first year PhD student – behind the scenes of the MaxQuant Summer School 2022 appeared first on PROTrEIN.

]]>
“Cross-Linking Mass Spectrometry: A sneak peak into its world!! ” https://protrein.eu/blog/cross-linking-mass-spectrometry-a-sneak-peak-into-its-world/ Sun, 26 Jun 2022 15:50:04 +0000 http://protrein.eu/?p=1554 It has been a while since our official PROTrEIN launch. We’ve had monthly ESR meetings during this time, where the majority of us have been able to discuss more details about our projects and present them to the group. Several ESRs are working on developing new tools and data analysis pipelines to improve cross linking […]

The post “Cross-Linking Mass Spectrometry: A sneak peak into its world!! ” appeared first on PROTrEIN.

]]>
It has been a while since our official PROTrEIN launch. We’ve had monthly ESR meetings during this time, where the majority of us have been able to discuss more details about our projects and present them to the group. Several ESRs are working on developing new tools and data analysis pipelines to improve cross linking mass spectrometer data analysis. It is worthy to note that these ESRs are supervised by great supervisors who’ve already made significant contributions to the field of cross linking.
For this month’s journal club, we therefore chose the review paper «Cross-linking mass spectrometry: methods and applications in structural, molecular, and systems biology» by F. O’Reilly and J. Rappsilber[1].
Cross-linking mass spectrometry (CLMS) has emerged as a useful technique in structural biology research, complementing traditional approaches such as x-ray crystallography and electron microscopy. CLMS allows researchers to investigate proteins in solution, capturing them in a dynamic state that is closer to physiological settings. CLMS also allows for the analysis of heterogeneous samples containing compounds in low quantities, showcasing the benefits of CLMS.

Overview of CLMS:

In the cross-linking reaction, covalent bonds are formed between the reactive groups of the cross-linker and surface residues of proteins, peptides and/or nucleic acids (DNA and RNA).
This way, residues that are within a certain reach of each other can be linked, thereby providing information about tertiary structure as well as interactions between protein complexes and/or nucleic acids.
Chemical cross-linkers are molecules with a spacer region flanked by reactive end groups that are extremely specific, while others are not. The chemical properties of the reactive groups and the length of the spacer region establish the limits of a cross-linker and therefore can be designed to accommodate different CLMS workflows. CLMS is not limited to studying the structure and interactions between proteins but can also be performed to gain knowledge about interactions between proteins and nucleic acids, however this blog post focuses specifically on the application and workflows involving proteins and peptides.

The general CLMS workflow is shown in Figure 1.

Figure 1. General cross-linking mass spectrometry (CLMS) workflow. (a) First step is choosing the correct cross-linker for the experiment. Depending on the question you want answered and the workflow, the cross-linker may need to be cleavable in the mass spectrometer, be isotopically labeled or have properties that allow for enrichment. This will be described in further detail in the CLMS workflow section. Once the cross-linker has been added to the sample, inter- and intra-protein crosslinks are formed (b). After the cross-linking reactions, the proteins in the sample are digested by a protease (c) yielding a mix between cross-linked and linear peptides. In some workflows, the cross-linked peptides are enriched (d) before data acquisition by MS/MS (e). The last step is data analysis that aims at identifying cross-linked peptides (f).

CLMS Applications:

The review paper we are discussing focuses on four overall applications of CLMS in protein studies, as illustrated in Figure 2, however CLMS has many more applications.

Modeling of protein complex topology is one of the most common uses of CLMS to investigate how proteins are ordered with respect to one another as they form complexes, a process known as protein complex topology. In circumstances where the complex topology is unknown, information concerning distance limits between surface residues of proteins that are known to be complex can be used with other structural research tools to aid modeling of the complex topology. 

Tertiary protein structure modeling can benefit from High Density (HD) CLMS. HD-CLMS data is acquired by using a cross-linker that is semi specific; one of the reactive groups only binds to specific residues, whereas the other group has no binding restrictions. This type of cross-linker will create a highly dense mapping of distance restraints between residues on the protein surface. Using this information can help exclude certain arrangements of the secondary structure elements of the proteins, thereby supporting modeling of the tertiary structure. Using a semi-specific cross-linker makes the data highly complex due to the large number of crosslinks formed in the reaction. 

Quantitative CLMS based comparative studies can be utilized to study proteins in different conformations. The relative abundance of distinct cross-links generated in each sample can be determined by adding isotopically labeled cross-linkers to samples from different experimental conditions. This information can help detect whether a protein is predominantly in one conformation or another, depending on the conditions of the experiment. These comparative analyses work best for proteins that go through conformational changes that strongly affect the structure since it affects the amount of cross-links that can be produced.

Proteome-wide CLMS studies focus on studying protein-protein interactions (PPIs) in large scale However, due to the enormous number of possible cross-links that might form between peptides, the data from these tests is exceedingly complex, posing some issues. Proteome-wide PPI research can be done in a variety of ways. Targeted pulldown procedures, in which natural protein complexes are identified and examined; cell lysate analysis, in which PPIs in the soluble proteome are explored; and in situ studies of complete cells or organelles are just a few examples.

Figure 2. Four situations where CLMS can be applied and aid modeling of protein-protein interactions and protein structure. (a) studying topology of protein complexes. (b) tertiary structure of single proteins. (c) comparative studies using quantitative CLMS to study protein conformation and (d) proteome-wide studies of protein topology.

CLMS workflows:

Although the overwhelming number of workflows available can be perplexing for newcomers in this field the development of standardized reagents and workflows has significantly boosted the simplicity to use CLMS. For the detection of cross-linked peptides, a number of software solutions are now available. The typical method for gauging confidence, regardless of the search program employed, is to utilize a target-decoy strategy to estimate the false discovery rate (FDR).
Emerging reporting standards and data-visualization tools are facilitating this technique’s accessibility, which are discussed below briefly:

❖ Reporting standards in CLMS: Because this field hasn’t publicly agreed on minimal reporting standards, it’s difficult to evaluate papers and reuse data. “mzIdentML” (http://www.psidev.info/mzidentml/) is an XML-based reporting standard for proteomics data developed by the Human Proteomics Standards Initiative (HUPO-PSI), which includes CLMS.
Raw spectrometric data should be deposited in certain public repository after publication.There is a need for clarification when reporting results when the word ‘cross-link’ is used interchangeably for peptide spectral matching (PSMs), peptide pairings, and residue pairs, because the defined FDR at the PSM or peptide level results in an unknown and typically much larger FDR at the level of residue pair.

❖ Data Visualization and interpretation: Software for visualizing discovered cross-links and the mass spectra that lead to their identification has been developed by numerous laboratories to make CLMS data accessible. Many levels of information is provided by cross-linking studies including:
A. Residue–Residue links
B. 3D structural information
C. Protein–Protein interactions

Figure 3 shows that their combination is one-of-a-kind, necessitating custom visualization.

Figure 3. Visualization solutions for CLMS data. (a) Spectra identified as cross-linked peptides can be manually assessed (b) Cross-linked proteins can be visualized with node and edge graphs to display interconnectivity of proteins (c) Mapping of cross-links on known 3D structures or homology models can score and validate cross-links and show those that violate the distance restraints.

“Notably, the cross-linker spacer’s chemistry can be tweaked, enabling data analysis simpler and boosting confidence in the cross-links found. As a result, before starting a study, it’s important to think about the best cross-linker to be used in conjunction with the analysis pipeline.”

Now we’ll look at some of the most prominent methods for analyzing CLMS.

Universal approach: Most comprehensive method, does not necessitate changing the cross-linker spacer in order to perform downstream analysis, commonly employed in conjunction with commercial cross-linkers, and effective for cross-linkers that can’t be modified in the spacer region , like photo amino acids. Isotope labeling is not crucial for identification and can be employed in quantitative or comparative studies. Using modern mass spectrometers, MS/MS spectra can be recorded at high resolution, which reduces the chance of getting false positive hits in the identification. StavroX[2], Xlink-Identifier[3], and Xcomb[4] generate a database of potentially cross-linked peptide pairs, but as the number of proteins grows, their computational time increases. Modification search combined with experimental heuristics that computationally enrich possible cross-linked peptides, save search time before scoring the spectra in Xi[5], Plink[6], XLSearch[7], Protein Prospector82[8], ECL2[9] , and Kojak[10].

Labeled cross-linker approach: Samples are treated with a mix of a heavy-isotope-labeled cross-linker and its unlabeled equivalent. Cross-linked peptides can then be identified in the MS1 spectra by searching for doublet peaks that are displaced by the mass of the heavy isotopes. MS2 spectra of the ‘light’ (unlabelled cross-linker) and ‘heavy’ (isotopically labeled cross-linker) precursors can reveal which fragment ions can contain the cross-linker, for confident cross-link identification, Hekate[11], StavroX, and the widely used xQuest[12] are just a few examples of search tools that use this approach. Talking about its positive side, this method streamlines data-analysis operations and can even be useful where high-accuracy mass spectrometers are not accessible, but it also increases the complexity of the MS1 spectrum space, potentially lowering recognition rates. Furthermore, requiring both heavy and light precursors for fragmentation can cause problems in complex samples.

MS2-cleavable cross-linker approach: utilizes cross-linkers that are cleavable during MS2 fragmentation, resulting in two peptides per MS2 spectrum that can be observed as unique cross-link-specific fragment ions. As cross-linked peptides are vast and branched, their fragmentation spectra are complex and uneven. The vast number of potential peptide combinations, combined with the frequently poor fragmentation of one of the cross-linked peptides, can make identifying the two peptides challenging, but this can be made easier by separating the two peptides in the mass spectrometer. This technique employs longer duty cycles than MS2-only approaches and requires additionally to execute MS3. Acquisition approaches for these cross-linkers have been designed by several laboratories along with their respective search software, such as ICC-CLASS[13], MeroX[14], X-links/Blinks[15,16] and XlinkX2.0[17,18]. After the review was published in 2018, an additional cross-linking search engine, MS Annika[19] was published in 2021


Figure 4 CLMS data acquisition and analysis workflows.(a) The ‘universal approach’ uses cross-linkers with simple spacers (b) Labeled cross-linker approach using isotopically labeled cross-linkers. (c) Cross-linker approach that uses cleavable cross-linkers in MS2 fragmentation.

Conclusion:

We hope you now understand why CLMS is such a powerful tool for examining protein interactions and topology. CLMS is a blooming field, with new processes and cross-linkers being developed as well as data analysis. Advances in data acquisition should be accompanied by improvements in data analysis. PROTrEIN ESRs, as well as other community members, are working on new tools and analysis pipelines to empower researchers to make new discoveries.

What do you anticipate CLMS will provide next?

References:

  1. O’Reilly, F.J., Rappsilber, J. Cross-linking mass spectrometry: methods and applications in structural, molecular and systems biology. Nat Struct Mol Biol 25, 1000–1008 (2018). https://doi.org/10.1038/s41594-018-0147-0
  2. Götze M, Pettelkau J, Schaks S, Bosse K, Ihling CH, Krauth F, Fritzsche R, Kühn U, Sinz A. StavroX–a software for analyzing crosslinked products in protein interaction studies. J Am Soc Mass Spectrom. 2012 Jan;23(1):76-87. doi: 10.1007/s13361-011-0261-2. Epub 2011 Oct 25. PMID: 22038510.
  3. Du X, Chowdhury SM, Manes NP, Wu S, Mayer MU, Adkins JN, Anderson GA, Smith RD. Xlink-identifier: an automated data analysis platform for confident identifications of chemically cross-linked peptides using tandem mass spectrometry. J Proteome Res. 2011 Mar 4;10(3):923-31. doi: 10.1021/pr100848a. Epub 2011 Feb 16. PMID: 21175198; PMCID: PMC3048902.
  4. Panchaud A, Singh P, Shaffer SA, Goodlett DR. xComb: a cross-linked peptide database approach to protein-protein interaction analysis. J Proteome Res. 2010 May 7;9(5):2508-15. doi: 10.1021/pr9011816. PMID: 20302351; PMCID: PMC2884221.
  5. Giese SH, Fischer L, Rappsilber J. A Study into the Collision-induced Dissociation (CID) Behavior of Cross-Linked Peptides. Mol Cell Proteomics. 2016 Mar;15(3):1094-104. doi: 10.1074/mcp.M115.049296. Epub 2015 Dec 30. PMID: 26719564; PMCID: PMC4813691.
  6. Yang B, Wu YJ, Zhu M, Fan SB, Lin J, Zhang K, Li S, Chi H, Li YX, Chen HF, Luo SK, Ding YH, Wang LH, Hao Z, Xiu LY, Chen S, Ye K, He SM, Dong MQ. Identification of cross-linked peptides from complex samples. Nat Methods. 2012 Sep;9(9):904-6. doi: 10.1038/nmeth.2099. Epub 2012 Jul 8. PMID: 22772728.
  7. Ji C, Li S, Reilly JP, Radivojac P, Tang H. XLSearch: a Probabilistic Database Search Algorithm for Identifying Cross-Linked Peptides. J Proteome Res. 2016 Jun 3;15(6):1830-41. doi: 10.1021/acs.jproteome.6b00004. Epub 2016 May 6. PMID: 27068484; PMCID: PMC5770149.
  8. Trnka MJ, Baker PR, Robinson PJ, Burlingame AL, Chalkley RJ. Matching cross-linked peptide spectra: only as good as the worse identification. Mol Cell Proteomics. 2014 Feb;13(2):420-34. doi: 10.1074/mcp.M113.034009. Epub 2013 Dec 12. PMID: 24335475; PMCID: PMC3916644.
  9. Yu F, Li N, Yu W. Exhaustively Identifying Cross-Linked Peptides with a Linear Computational Complexity. J Proteome Res. 2017 Oct 6;16(10):3942-3952. doi: 10.1021/acs.jproteome.7b00338. Epub 2017 Sep 1. PMID: 28825304. 
  10. Hoopmann MR, Zelter A, Johnson RS, Riffle M, MacCoss MJ, Davis TN, Moritz RL. Kojak: efficient analysis of chemically cross-linked protein complexes. J Proteome Res. 2015 May 1;14(5):2190-8. doi: 10.1021/pr501321h. Epub 2015 Apr 15. PMID: 25812159; PMCID: PMC4428575.
  11. Holding AN, Lamers MH, Stephens E, Skehel JM. Hekate: software suite for the mass spectrometric analysis and three-dimensional visualization of cross-linked protein samples. J Proteome Res. 2013 Dec 6;12(12):5923-33. doi: 10.1021/pr4003867. Epub 2013 Oct 4. PMID: 24010795; PMCID: PMC3859183.
  12. Rinner O, Seebacher J, Walzthoeni T, Mueller LN, Beck M, Schmidt A, Mueller M, Aebersold R. Identification of cross-linked peptides from large sequence databases. Nat Methods. 2008 Apr;5(4):315-8. doi: 10.1038/nmeth.1192. Epub 2008 Mar 9. Erratum in: Nat Methods. 2008 Aug;5(8):748. PMID: 18327264; PMCID: PMC2719781.
  13.  Petrotchenko, E.V., Borchers, C.H. ICC-CLASS: isotopically-coded cleavable crosslinking analysis software suite. BMC Bioinformatics 11, 64 (2010). https://doi.org/10.1186/1471-2105-11-64
  14. Götze M, Pettelkau J, Fritzsche R, Ihling CH, Schäfer M, Sinz A. Automated assignment of MS/MS cleavable cross-links in protein 3D-structure analysis. J Am Soc Mass Spectrom. 2015 Jan;26(1):83-97. doi: 10.1007/s13361-014-1001-1. Epub 2014 Sep 27. PMID: 25261217.
  15. Hoopmann MR, Weisbrod CR, Bruce JE. Improved strategies for rapid identification of chemically cross-linked peptides using protein interaction reporter technology. J Proteome Res. 2010 Dec 3;9(12):6323-33. doi: 10.1021/pr100572u. Epub 2010 Nov 10. PMID: 20886857; PMCID: PMC3018735. 
  16. Anderson GA, Tolic N, Tang X, Zheng C, Bruce JE. Informatics strategies for large-scale novel cross-linking analysis. J Proteome Res. 2007 Sep;6(9):3412-21. doi: 10.1021/pr070035z. Epub 2007 Aug 3. PMID: 17676784; PMCID: PMC2475505.
  17. Liu F, Lössl P, Scheltema R, Viner R, Heck AJR. Optimized fragmentation schemes and data analysis strategies for proteome-wide cross-link identification. Nat Commun. 2017 May 19;8:15473. doi: 10.1038/ncomms15473. PMID: 28524877; PMCID: PMC5454533. 
  18. Liu F, Rijkers DT, Post H, Heck AJ. Proteome-wide profiling of protein assemblies by cross-linking mass spectrometry. Nat Methods. 2015 Dec;12(12):1179-84. doi: 10.1038/nmeth.3603. Epub 2015 Sep 28. PMID: 26414014. 
  19. Pirklbauer GJ, Stieger CE, Matzinger M, Winkler S, Mechtler K, Dorfer V. MS Annika: A New Cross-Linking Search Engine. J Proteome Res. 2021 May 7;20(5):2560-2569. doi: 10.1021/acs.jproteome.0c01000. Epub 2021 Apr 14. PMID: 33852321; PMCID: PMC8155564.

Source of gallery picture of this blog-post: Juan Gaertner/Science Photo Library/Getty Images

The post “Cross-Linking Mass Spectrometry: A sneak peak into its world!! ” appeared first on PROTrEIN.

]]>
Proteins form functional networks, and so do people – the first PROTrEIN summer school https://protrein.eu/blog/11-dec-2020/ Thu, 16 Dec 2021 09:21:00 +0000 http://protrein.eu/?p=810 The first PROTrEIN annual meeting and summer school took place between the 10th and 23rd of September 2021. Here our ESRs Marc and Mostafa share their thoughts and memories about it.

The post Proteins form functional networks, and so do people – the first PROTrEIN summer school appeared first on PROTrEIN.

]]>
Applying and going through the selection for a PhD position in PROTrEIN was challenging enough, but it did not compare to the subsequent excruciating wait for a response. However, the latter arrived approximately one week later and the wait had been well-worth it. It is impossible to forget the moment the notification arrived on the smartphone and the seconds-that-seem-like-years that it took to open the email. We had been selected to be ESRs of PROTrEIN and a major new chapter of our lives was set in motion.Several months later, it was finally time for the first official event of PROTrEIN, where all members of the network were going to gather in one (virtual) place. The first PROTrEIN Summer School was officially launched on September 10th, 2021, and while it was unfortunate that the exciting city of Barcelona could not be the venue for this event, everyone was nonetheless excited to finally have the opportunity to meet the rest of the members. The Summer School was split into three parts: the consortium meeting, the advanced proteomics course and the module on open and responsible research.

After having everyone connected in the Zoom session, Jonas Krebs and Eduard Sabidó launched the day with some introductory slides on the project’s organization, funding and timelines. It was an ideal recap of things that we had already looked up in the previous months (we were impatient, ok?), but also some helpful new information was thrown in, such as a proposed schedule for regular ESR meetings etc. The rest of the day was filled with activities to break the ice between all members and get them to interact in a fun setting. It can be difficult to achieve these goals in a videoconference setting, but we should congratulate all the people that were involved in organizing the event, such as Anna, Imma and Damjana from the CRG Training Unit. They had done an amazing job, not only with all the events and activities that were planned (and that we will go over below), but also for the simpler details such as the music playing in Zoom during the coffee breaks.

A world-wide connection

One of the first ice-breakers we did, was to open up a virtual world map and place a pin on our current location. From Tamper, Finland to Mangalore, India and from Jeonju, South Korea to Rockville, USA, there were pins scattered all around. It is an impressive feat of scientific project such as PROTrEIN, to create such diverse networks of individuals across all these time zones. And soon enough, we would all be on the European continent doing exciting research!

Express your feelings

A couple of weeks before the event, we had been asked to find and post a GIF image each to describe their feelings about ourrespective PhD projects. The old proverb says “one picture is worth a thousand words”, so you would think that an animated one would be adequate to describe these feelings. Well, wrong! There were so many feelings and thoughts in our minds regarding our forthcoming PhD project, and some of them do not even have a word dictionary. Nonetheless, we all tried our best to focus on one of those feelings and we ended up with a highly expressive wall of GIF images.

Speed-dating

Writing this post several weeks after the events, we still vividly remember the reaction of Eduard at the end of this particular section: “Wow! This was… intense!”. It was the most suitable reaction to what was otherwise an incredibly fun activity. Thanks to Zoom’s functionality, pairs were formed and placed in separate breakout rooms with precisely two minutes to quickly get to know each other, before being abruptly interrupted and thrown into another breakout room with someone new. Do you remember the beginning of this post where we described how can a few seconds feel so long when you open an important email? Well, these 2 minutes in a breakout room with someone you barely know felt like the complete opposite! It provided mixed feelings of excitement to discover who your new partner was, while at the same time carrying over the frustration from being interrupted mid-sentence with the previous. Now we look forward to the next event (1st Winter School in Zurich) and hope that we are able to meet in person and complete all these unfinished conversations.

Advanced Proteomics course

After getting familiar with all ESRs and supervisors, it was time to focus on the concept of proteomics. We took the online advanced proteomics course from the 13th to the 17th of September, and each day consisted of two main parts. In the first part (morning), a series of lectures were provided by our supervisors and other invited professors. These lectures were provided in-depth information on various topics such as data-dependent and –independent acquisition, manual spectra annotation and even reconstruction of ancient protein sequences. We gained a lot of insight into some core, but also advanced, concepts of proteomics. One of the most interesting sessions was about the applications of machine learning to proteomics data. In the next part (afternoon), we experienced some hands-on workshops where ESRs were separated into groups to work together on a particular problem. This experience was really challenging and also exciting, mainly because we had to network with other ESRs, coming from different backgrounds, to resolve an issue at a specific time.  

Open & Responsible research

The last three days of the Summer School were dedicated to the topic of open and responsible research. The topic itself is broad enough, and through some excellent seminars/workshops we had the chance approach the matter from a variety of angles. Going into details would make what is an already pretty long post interminable, but we had the chance to learn about concepts such as data openness (and particularly the FAIR principles), result reproducibility and the various options and procedures that go into authoring and publishing scientific articles. Best of all was the opportunity we had to openly question and discuss these important matters among ourselves.We hope these are topics that we will have the chance to revisit in future posts of this blog!

We hope these are topics that will be revisited in future posts of this blog!

The post Proteins form functional networks, and so do people – the first PROTrEIN summer school appeared first on PROTrEIN.

]]>
How can proteomics data become more reproducible? https://protrein.eu/blog/1-jan-2021/ Thu, 14 Oct 2021 15:59:35 +0000 http://protrein.eu/?p=802 In this Journal Club, our ESRs Louise and Shamil present the article "Strategies to enable large-scale proteomics for reproducible research" by R. Poulos et al. from July 2020

The post How can proteomics data become more reproducible? appeared first on PROTrEIN.

]]>
In the first PROTrEIN Summer School we were introduced to proteomics. We learned about many different aspects of proteomics, including some of the challenges associated with proteomics experiments. However, as this was only an introductory course, everything concerning proteomics could not be highlighted. We therefore saw this first journal club as an opportunity to highlight data reproducibility, which is important not only for proteomics experiments but also in all other areas of research. Research reproducibility is indeed one of the major science problems, and it becomes even more significant when scientists are dealing with large quantities of noisy data. This paper focuses specifically on the reproducibility of DIA MS data in large-scale proteomics studies.

The researchers performed 1560 DIA-MS runs on six different instruments (of the same type) over the course of four months, with one major instrument cleaning after three months. The purpose of this was to investigate the reproducibility of the data over the extensive time period. They found that the data became less reproducible over time, but that the reproducibility was ‘restored’ after the cleaning. Using their novel normalization method, the researchers were able to mitigate the unwanted variation of the data that arose due to the decrease in instrument performance over time. Furthermore, they demonstrated that using technical replicates of samples allows for data imputation for a more complete dataset, hence making the data more reproducible. 

We think this paper is relevant for us as computational proteomics researchers, because even though we are not performing experiments in the laboratory, we must still be aware of data reproducibility and what measures can be taken to improve it, both during and after data acquisition. As large-scale proteomics research becomes increasingly common, the strategies mitigating instrumental and other non-biological variation grow in importance.

Figure 1. Subsection of Fig. 3 in the article. The top figure shows how the non-normalized peptide intensities decrease over time due to a decline in instrument performance, but that the peptide intensities are restored after the major instrument clean. Decreasing peptide intensities are an indicator of the data becoming less reproducible. The bottom figure shows the peptide intensity after normalization by their novel normalization method (RUV-III-C). Using the normalization method, the unwanted variation observed in the top figure is mitigated, improving the data reproducibility of the data.

Reference: Nat Commun. 2020 Jul 30;11(1):3793. doi: 10.1038/s41467-020-17641-3.

The post How can proteomics data become more reproducible? appeared first on PROTrEIN.

]]>