PeptideStack
5.2kquestions
20kanswers
220users

How do I read a sequence coverage map, and why is 100% coverage sometimes worthless?

Asked 30 Oct 2025Modified 7 months agoViewed 4.6k times
17

I received a mapping report on tirzepatide. It states "sequence coverage: 100%" in bold at the top, and then the fragment table has exactly two entries in it. Thirty-nine residues, two peptides.

Something about that feels wrong but I cannot say what. Technically every residue is inside one of the two peptides, so 100% is arithmetically true. But intuitively a map with two pieces seems like it has told me almost nothing that the intact mass did not, since one of the pieces is more than half the molecule.

Is my intuition right, and if so what is the metric I should be looking at instead of coverage percentage? And is there something about tirzepatide specifically that makes trypsin a bad choice here, because on semaglutide I have seen maps with more fragments than this.

peptide-mapping
peptide-mapping

Enzymatic digestion followed by LC-MS/MS to confirm sequence rather than just mass. The method that catches a scrambled sequence or a D-amino-acid…

83 questions
lc-ms
lc-ms

Liquid chromatography coupled to mass spectrometry: separation and identification in one run, and the method of choice when you need to know not…

14 questions
coa
coa

Certificates of analysis: what fields a useful one carries, how to tell a real analytical report from a marketing document, batch and lot…

749 questions
tirzepatide
tirzepatide

A dual GIP and GLP-1 receptor agonist. Questions here cover the SURPASS and SURMOUNT programmes, the practical differences from a pure GLP-1…

162 questions
batch-testing
batch-testing

Testing at the batch or lot level: sampling plans, how many vials from a lot need testing to say anything about the lot, and the difference…

865 questions
shareeditfollowflag
AL
askeda_lindgren46k13830 Oct 2025
6Your intuition is exactly right and the reason is tirzepatide has no arginine at all. – low_dead_space 34 days ago
7Coverage percentage is the most quoted and least informative number on a mapping report. – Dr_Nadia_Farsi 3 months ago
add a comment

3 Answers

Accepted answer first, then by votes
48

Accepted answer

Your intuition is correct, the coverage figure is honest and nearly meaningless, and the cause is a structural feature of tirzepatide: it contains no arginine, and its only unblocked lysine is at position 16.

Why trypsin gives two fragments

Tirzepatide is 39 residues, Aib at 2 and 13, acylated Lys20, C-terminal amide. Trypsin cleaves after Lys and Arg. There is no Arg. Lys20 bears the AEEA-AEEA-gamma-Glu-C20-diacid side chain on its epsilon-amine, so trypsin will not touch it. That leaves Lys16, followed by Ile17, which is not proline, so it cleaves.

FragmentResiduesLengthMonoisotopic mass (Da)Observed as
T11-16161807.852+ at 904.93
T217-39233020.682+ at 1511.35, 3+ at 1007.90

Those two masses sum correctly: 1807.85 + 3020.68 - 18.01 = 4810.52, which is the monoisotopic mass of intact tirzepatide. So the report is internally consistent and the arithmetic checks out. Coverage is genuinely 39 of 39.

Why 100% coverage is the wrong metric

Coverage tells you which residues appeared inside some identified peptide. It does not tell you how precisely a defect could have been located, and localisation is the entire reason to run a map. The useful way to think about it is average fragment length, or more directly, the number of residues over which a modification would be indistinguishable.

On this report, a mass shift anywhere in T2 is localised to a 23-residue window that contains the acylation site, the Trp, all three prolines and the C-terminal amide. A +16 on that fragment could be Trp25 oxidation or several other things. A -18 could be Ser dehydration at 32, 33 or 39. You have not localised anything; you have subdivided the molecule into two halves.

Compare a Glu-C digest. Glu-C in phosphate buffer cleaves after both Glu and Asp; tirzepatide has Glu3, Asp9 and Asp15:

FragmentResiduesLengthLocalises
E11-33N-terminal Tyr1, Aib2
E24-96Thr5, Phe6, Thr7, Ser8, Asp9
E310-156Tyr10, Ser11, Ile12, Aib13, Leu14, Asp15
E416-3924Lys16, the acylated Lys20, Trp25, the PPP region, the C-terminal amide

Better at the N-terminus, still bad at the C-terminus. The real answer is that no single enzyme handles this sequence, and a serious map uses two or three in parallel and stitches overlapping fragments together:

  • Chymotrypsin cleaves after aromatic residues and Leu: Tyr1, Phe6, Tyr10, Leu14, Phe22, Trp25, Leu26. That is seven sites, and crucially several of them are in the C-terminal half that trypsin and Glu-C leave intact. Chymotrypsin is the enzyme that breaks up T2.
  • Asp-N cleaves N-terminal to Asp, giving fragments offset by one from the Glu-C set, which is exactly what you want for overlap.
  • Non-specific digestion with pepsin or a short thermolysin incubation gives a shotgun of overlapping fragments; messier to interpret, excellent for localisation.

The PPP motif near the C-terminus is genuinely hard for everything. Proline resists cleavage on its N-terminal side by most proteases and suppresses CID fragmentation, so residues 36 to 39 tend to be the least well characterised part of the molecule in any map. A report that quietly puts them inside a large fragment and calls it 100% has not lied; it has just not done the difficult part.

What to look for on a mapping report instead of coverage

  1. Number of fragments and the longest fragment. If the longest fragment is over about 15 residues, the map has poor localising power in that region.
  2. Whether MS/MS was acquired on every fragment or only on some. Accurate mass on a fragment confirms composition; MS/MS confirms sequence. A "map" with only fragment masses and no fragment spectra is a digest mass list, and it cannot distinguish a transposition or place a modification within a fragment.
  3. Whether more than one enzyme was used. Single-enzyme maps on peptides with sparse cleavage sites are structurally limited, and this molecule is the textbook case.
  4. Explicit confirmation of the modification site. The one thing you most want stated is that the acyl side chain was localised to Lys20 by MS/MS, not merely that a fragment containing Lys20 had the right total mass. Those are different claims.
  5. Missed-cleavage products, assigned as such. Their presence at high abundance means the digest was incomplete and the coverage claim is soft.
  6. Digest conditions: enzyme ratio, pH, temperature, duration. Without these you cannot evaluate any modification finding, for the artefact reasons covered elsewhere in this tag.

Your report is not fraudulent. It is a single-enzyme map on a sequence that punishes single-enzyme maps, presented with the most flattering metric in front. Ask whether they can add a chymotryptic digest, and ask for the MS/MS confirmation of the Lys20 acylation site specifically. Those two additions turn it into a real characterisation.

edited 26 Dec 2025 by pierce_count — added a caveat about sampling

shareimprove this answerflag
PC
answered · acceptedpierce_count15k2830 Nov 2025
8Longest-fragment length as the headline metric instead of coverage percent. That is the correction the whole field needs. – ben_akintola 9 months ago
The distinction between "a fragment containing Lys20 had the right mass" and "the side chain was localised to Lys20" is the crux. Vendors conflate them constantly. – tess_amankwah 32 days ago
6Chymotrypsin plus trypsin on tirzepatide is what I have seen done properly. Two digests, one report. – h_pergande 3 months ago
add a comment
Sponsored

PeptideMeter - Independent Peptide Analytics

Aggregated, published test results and vendor ratings built from submitted batches. Methodology stated, dataset browsable, no listing fees.

Browse results
17

To make the semaglutide comparison in the question explicit, since the difference is instructive rather than accidental.

Semaglutide has Arg34 and Arg36 in addition to the blocked Lys26. Trypsin therefore gives three fragments rather than two — but look at where they fall: 7-34, 35-36 and a free Gly at 37. One 28-residue fragment carrying 93% of the molecule, plus a dipeptide and a single amino acid. That is worse localisation than tirzepatide's two-fragment map, not better, even though it has more entries in the table.

Which is the general lesson: fragment count is no better a metric than coverage percentage. What matters is whether the fragments are short enough, in the regions you care about, to place a modification. Both of these molecules concentrate their interesting chemistry — the acylated lysine, the Asp that forms aspartimide, the Trp that oxidises — inside the large fragment that trypsin cannot break up.

For semaglutide the enzyme that fixes it is Glu-C in phosphate, which puts the acylated Lys26 into a six-residue fragment on its own. For tirzepatide, Glu-C does not help at the C-terminus and chymotrypsin does. There is no universal answer; the enzyme choice follows from where the Lys, Arg, Glu, Asp and aromatic residues sit in the specific sequence, which is a thing you can work out from the sequence in five minutes before you commission the test.

Doing that in advance is also the best way to evaluate a quote. If a lab proposes a tryptic map on tirzepatide and cannot explain how they will localise anything in residues 17 to 39, they have not thought about your molecule.

shareimprove this answerflag
IB
answeredines_brandt93k24812 Dec 2025
8

One more thing to check on any coverage map that people skip: the false discovery rate and the identification criteria, if the report came out of a database search engine rather than by hand.

Automated mapping software assigns peptides by matching observed masses and fragment spectra against an in-silico digest. On a 39-residue single-protein search space that is a nearly trivial problem and the assignments will be right. But the software will also happily assign a peptide from a mass match alone at wide tolerance, and it reports coverage from whatever it assigned. Two things to confirm:

  • Mass tolerance used. Precursor tolerance of 10 ppm and fragment tolerance of 20 ppm is reasonable on a high-resolution instrument. A precursor tolerance of 0.5 Da means the software could not have distinguished a deamidated peptide from a correct one, and any deamidation finding from that search is noise.
  • Whether variable modifications were enabled, and which. A search with oxidation, deamidation, dehydration and the acyl side chain enabled as variable modifications will find them if present. A search with none enabled will assign every modified peptide as unmatched and then not report it, and coverage will look lower rather than the modification being found. Conversely a search with twenty variable modifications enabled will find spurious ones by chance.

The tell for a hand-checked map versus an unreviewed software output: a hand-checked report lists the unassigned peaks and says what was done about them. Software output lists only what matched. On a peptide this small there is no excuse for not accounting for every peak above a stated threshold, and asking for that list is a fair request.

shareimprove this answerflag
DF
answeredDr_Nadia_Farsi90k2588 Nov 2025

Your answer

Ask PeptideStack is a static archive. Posting is closed, but the norms are worth stating: answer the question that was asked, show your working, cite the trial or the certificate, and say plainly where the evidence runs out.

Not medical advice. Research-use-only compounds are not approved for human use.