Chemistry / Organic Chemistry Biosynthetic Pathways, Heterocycles & Bioactive Scaffolds 100% Free Open Access
Chapter 6 โ€ข Theory & Derivations

Unit 6: Amino Acids, Peptides & Protein Architectures

Comprehensive physical organic and biochemical foundations of amino acids, peptides, and proteins: classic and asymmetric chemical syntheses (Strecker, Gabriel, diethyl acetamidomalonate), multi-prototropic equilibria and isoelectric point ($pI$) calculus, amide bond stereodynamics and Ramachandran dihedral landscapes ($\phi, \psi$), Merrifield solid-phase peptide synthesis (SPPS), sequence determination via Edman degradation and tandem mass spectrometry, and the thermodynamics of protein folding.

ยง6.1 Natural $\alpha$-Amino Acids: Classification, Stereochemistry & Side Chains

Amino acids are bifunctional organic compounds containing both an amino group ($-\text{NH}_2$) and a carboxyl group ($-\text{COOH}$). In $\alpha$-amino acids, both groups are bonded to the same $\alpha$-carbon atom:

$$\text{H}_2\text{N}-\text{CH(R)}-\text{COOH}$$

Stereochemical Configuration: The L-Series

With the exception of glycine (where $\text{R}=\text{H}$, which is achiral), all 19 standard proteinogenic amino acids possess a chiral $\alpha$-carbon:

  • In Fischer projections with the $-\text{COOH}$ group at the top and side chain $\text{R}$ at the bottom, the $\alpha$-amino group points to the left in all naturally occurring proteinogenic amino acids. Thus, they belong to the L-configuration.
  • Under the Cahn-Ingold-Prelog $(R/S)$ system, 18 of the 19 chiral L-amino acids are $(S)$-enantiomers.
  • L-Cysteine is the sole exception: Because the sulfur atom in the $-\text{CH}_2\text{SH}$ side chain has a higher atomic number than the oxygen atoms of the carboxyl group, the priority of the side chain exceeds that of $-\text{COOH}$, rendering natural L-cysteine (R)-cysteine.

Classification of the 20 Proteinogenic Side Chains

1. Non-Polar, Aliphatic: Glycine (Gly, G), Alanine (Ala, A), Valine (Val, V), Leucine (Leu, L), Isoleucine (Ile, I), Proline (Pro, P, a cyclic secondary imino acid), Methionine (Met, M).

2. Aromatic: Phenylalanine (Phe, F), Tyrosine (Tyr, Y), Tryptophan (Trp, W).

3. Polar, Uncharged: Serine (Ser, S), Threonine (Thr, T), Cysteine (Cys, C), Asparagine (Asn, N), Glutamine (Gln, Q).

4. Positively Charged (Basic): Lysine (Lys, K, $\epsilon\text{-NH}_3^+$), Arginine (Arg, R, guanidinium), Histidine (His, H, imidazole).

5. Negatively Charged (Acidic): Aspartate (Asp, D, $\beta\text{-COO}^-$), Glutamate (Glu, E, $\gamma\text{-COO}^-$).

Physical Reference Table: Acid-Base Dissociation Constants ($pK_a$) & Isoelectric Points ($pI$)

The twenty standard proteinogenic amino acids exhibit precise ionization constants in aqueous solution at 298.15 K:

| Amino Acid | Symbol | $pK_{a1}$ ($\alpha\text{-COOH}$) | $pK_{a2}$ ($\alpha\text{-NH}_3^+$) | $pK_{aR}$ (Side Chain) | Isoelectric Point $pI$ | Hydropathy Index | | :--- | :--- | :--- | :--- | :--- | :--- | :--- | | Glycine | Gly (G) | $2.34$ | $9.60$ | โ€” | $5.97$ | $-0.4$ | | Alanine | Ala (A) | $2.35$ | $9.69$ | โ€” | $6.02$ | $+1.8$ | | Valine | Val (V) | $2.32$ | $9.62$ | โ€” | $5.97$ | $+4.2$ | | Leucine | Leu (L) | $2.36$ | $9.60$ | โ€” | $5.98$ | $+3.8$ | | Isoleucine | Ile (I) | $2.36$ | $9.68$ | โ€” | $6.02$ | $+4.5$ | | Proline | Pro (P) | $1.99$ | $10.96$ | โ€” | $6.48$ | $-1.6$ | | Methionine | Met (M) | $2.28$ | $9.21$ | โ€” | $5.75$ | $+1.9$ | | Phenylalanine | Phe (F) | $1.83$ | $9.13$ | โ€” | $5.48$ | $+2.8$ | | Tryptophan | Trp (W) | $2.38$ | $9.39$ | โ€” | $5.89$ | $-0.9$ | | Tyrosine | Tyr (Y) | $2.20$ | $9.11$ | $10.07$ (Phenolic) | $5.66$ | $-1.3$ | | Serine | Ser (S) | $2.21$ | $9.15$ | $\sim 13.6$ | $5.68$ | $-0.8$ | | Threonine | Thr (T) | $2.11$ | $9.62$ | $\sim 13.6$ | $5.87$ | $-0.7$ | | Cysteine | Cys (C) | $1.96$ | $10.28$ | $8.18$ ($-\text{SH}$) | $5.07$ | $+2.5$ | | Asparagine | Asn (N) | $2.02$ | $8.80$ | โ€” | $5.41$ | $-3.5$ | | Glutamine | Gln (Q) | $2.17$ | $9.13$ | โ€” | $5.65$ | $-3.5$ | | Aspartate | Asp (D) | $1.88$ | $9.60$ | $3.65$ ($\beta\text{-COOH}$) | $2.77$ | $-3.5$ | | Glutamate | Glu (E) | $2.19$ | $9.67$ | $4.25$ ($\gamma\text{-COOH}$) | $3.22$ | $-3.5$ | | Lysine | Lys (K) | $2.18$ | $8.95$ | $10.53$ ($\epsilon\text{-NH}_3^+$) | $9.74$ | $-3.9$ | | Arginine | Arg (R) | $2.17$ | $9.04$ | $12.48$ (Guanidinium) | $10.76$ | $-4.5$ | | Histidine | His (H) | $1.82$ | $9.17$ | $6.00$ (Imidazole) | $7.59$ | $-3.2$ |

ยง6.2 Chemical Synthesis of $\alpha$-Amino Acids: Strecker, Gabriel & Acetamidomalonate

Industrial and laboratory syntheses provide racemic or enantiopure amino acids for pharmaceuticals and peptide chemistry.

The Strecker Synthesis (1850)

One of the oldest multi-component reactions in organic chemistry, converting an aldehyde into an $\alpha$-amino acid:

1. Imine Formation: An aldehyde condenses with ammonia to form an imine (or iminium ion):

$$\text{R-CHO} + \text{NH}_3 \rightleftharpoons \text{R-CH}=\text{NH} + \text{H}_2\text{O}$$

2. Cyanide Addition: Nucleophilic attack by cyanide ion ($\text{CN}^-$) affords an $\alpha$-aminonitrile:

$$\text{R-CH}=\text{NH} + \text{HCN} \to \text{R-CH}(\text{NH}_2)\text{-CN}$$

3. Acidic Hydrolysis: Exhaustive hydrolysis of the nitrile with aqueous $\text{HCl}$ yields the racemic $\alpha$-amino acid:

$$\text{R-CH}(\text{NH}_2)\text{-CN} + 2\text{ H}_2\text{O} + \text{HCl} \to \text{R-CH}(\text{NH}_3^+)\text{-COOH} \cdot \text{Cl}^- + \text{NH}_4\text{Cl}$$

Gabriel Phthalimide Synthesis

Potassium phthalimide is alkylated with an $\alpha$-halo ester (e.g., ethyl $\alpha$-bromoacetate), followed by hydrazinolysis (Ing-Manske procedure) or acidic hydrolysis to yield pure primary amino acids without over-alkylation to secondary/tertiary amines.

Diethyl Acetamidomalonate Synthesis

The most general and reliable laboratory protocol for complex amino acids:

  1. Deprotonation of diethyl acetamidomalonate with sodium ethoxide generates a resonance-stabilized enolate:
$$\text{CH}_3\text{CONH-CH}(\text{COOEt})_2 + \text{NaOEt} \to [\text{CH}_3\text{CONH-C}(\text{COOEt})_2]^- \text{Na}^+ + \text{EtOH}$$
  1. Nucleophilic $S_N2$ alkylation with an alkyl halide ($\text{R-X}$) introduces the desired side chain:
$$\to \text{CH}_3\text{CONH-C}(\text{R})(\text{COOEt})_2$$
  1. Vigorous refluxing with concentrated aqueous $\text{HCl}$ or $\text{HBr}$ simultaneously hydrolyzes the amide, saponifies both ethyl esters to carboxylic acids, and triggers spontaneous decarboxylation of the geminal dicarboxylic acid, yielding the racemic $\alpha$-amino acid:
$$\to \text{R-CH}(\text{NH}_3^+)\text{-COOH} + \text{CO}_2 + 2\text{ EtOH} + \text{CH}_3\text{COOH}$$

ยง6.3 Acid-Base Properties, Zwitterionic Equilibrium & Isoelectric Point ($pI$) Calculus

In both solid state and aqueous solution, amino acids exist predominantly as dipolar internal salts, known as zwitterions (German for 'hybrid ions'):

$$\text{H}_3\text{N}^+-\text{CH(R)}-\text{COO}^-$$

This explains their physical properties: high melting points ($>250^\circ\text{C}$ with decomposition), large dipole moments, and high solubility in polar water but insolubility in non-polar organic solvents.

Multi-Prototropic Speciation

An amino acid with an ionizable side chain undergoes multiple sequential deprotonations:

$$\text{H}_3\text{A}^+ \underset{K_{a1}}{\rightleftharpoons} \text{H}_2\text{A}^{\pm} \underset{K_{a2}}{\rightleftharpoons} \text{HA}^- \underset{K_{a3}}{\rightleftharpoons} \text{A}^{2-}$$

The fractional population $\alpha_i$ of each species is governed by the Henderson-Hasselbalch equation:

$$\text{pH} = pK_a + \log\left(\frac{[\text{Base}]}{[\text{Acid}]}\right)$$

Rigorous Calculus of the Isoelectric Point ($pI$)

The isoelectric point ($pI$) is the precise pH at which the net electrical charge of the amino acid ensemble is identically zero:

$$\langle z \rangle = \sum z_i \cdot \alpha_i = 0$$

At this pH, the molecule exhibits zero electrophoretic mobility.

1. Simple Amino Acids (Diprotic, Non-Ionizable Side Chain):

$$pI = \frac{pK_{a1}(\alpha\text{-COOH}) + pK_{a2}(\alpha\text{-NH}_3^+)}{2}$$

2. Acidic Amino Acids (Aspartate, Glutamate):

The zwitterionic form with net charge zero lies between the two carboxylic acid deprotonations:

$$pI = \frac{pK_{a1}(\alpha\text{-COOH}) + pK_{aR}(\text{side-chain } -\text{COOH})}{2}$$

3. Basic Amino Acids (Lysine, Arginine, Histidine):

The neutral zwitterionic form lies between the two basic nitrogen deprotonations:

$$pI = \frac{pK_{aR}(\text{side-chain}) + pK_{a2}(\alpha\text{-NH}_3^+)}{2}$$

ยง6.4 Peptides & The Peptide Bond: Partial Double-Bond Character & Ramachandran Plots

A peptide bond is an amide linkage formed by the condensation of the $\alpha$-carboxyl group of one amino acid with the $\alpha$-amino group of another:

$$\text{R}_1\text{-COOH} + \text{H}_2\text{N-R}_2 \to \text{R}_1\text{-CO-NH-R}_2 + \text{H}_2\text{O}$$

Partial Double-Bond Character and Planarity

Linus Pauling and Robert Corey (1951) deduced the structural constraints of the peptide bond through X-ray crystallography:

  • Resonance delocalization between the carbonyl $\pi$-electrons and the nitrogen lone pair creates a significant partial double bond:
$$\text{O}=\text{C}-\text{N}-\text{H} \longleftrightarrow ^-\text{O}-\text{C}=\text{N}^+-\text{H}$$
  • The $\text{C-N}$ bond length is $1.32\text{ \AA}$, intermediate between a standard $\text{C-N}$ single bond ($1.47\text{ \AA}$) and a $\text{C}=\text{N}$ double bond ($1.28\text{ \AA}$), possessing approximately $40\%$ double-bond character.
  • The rotational energy barrier about the $\text{C-N}$ bond is substantial ($\Delta G^\ddagger \approx 84\text{ kJ/mol}$), locking the six atoms of the peptide group ($\text{C}_\alpha^i, \text{C}, \text{O}, \text{N}, \text{H}, \text{C}_\alpha^{i+1}$) into a rigid, planar peptide plane.
  • The trans conformation (dihedral angle $\omega = 180^\circ$) is thermodynamically favored over the cis conformation ($\omega = 0^\circ$) by $\sim 10\text{ kJ/mol}$ due to steric clash between adjacent side chains (with proline being a notable exception where $\sim 10-20\%$ adopts cis).

The Ramachandran Plot ($\phi, \psi$)

Conformational freedom in the protein backbone is restricted to rotation around two single bonds per residue:

  • $\phi$ (Phi): Torsion angle around the $\text{N}-\text{C}_\alpha$ bond.
  • $\psi$ (Psi): Torsion angle around the $\text{C}_\alpha-\text{C}$ bond.

G. N. Ramachandran calculated sterically allowed regions of $(\phi, \psi)$ space by modeling atoms as hard spheres. Steric clash between carbonyl oxygens, amide hydrogens, and $\text{C}_\beta$ atoms excludes over $75\%$ of conformation space, restricting stable secondary structures to narrow permissible islands:

  • Right-handed $\alpha$-helix: $\phi \approx -57^\circ, \psi \approx -47^\circ$
  • $\beta$-pleated sheets: $\phi \approx -120^\circ \text{ to } -140^\circ, \psi \approx +135^\circ \text{ to } +150^\circ$
  • Left-handed $\alpha$-helix: $\phi \approx +57^\circ, \psi \approx +47^\circ$ (primarily glycine)

Advanced Research Monograph: AI Structure Prediction & De Novo Protein Design

The resolution of the 50-year-old protein folding problem by deep learning has transformed chemical biology:

1. Deep Learning Architectural Principles (AlphaFold & ESMFold):

By training deep neural networks on $>200,000$ crystallographic structures in the Protein Data Bank (PDB) and billions of metagenomic sequences:

  • Invariant point attention layers and evoformer modules extract evolutionary covariation signals between amino acid pairs across phylogenetic alignments.
  • AlphaFold accurately predicts three-dimensional coordinates of backbone and side-chain atoms with sub-Angstrom root-mean-square deviation (RMSD $< 1.0\text{ \AA}$), matching high-resolution X-ray crystallography and Cryo-EM.

2. De Novo Protein Design (Baker Laboratory):

David Baker and coworkers inverted the prediction process using deep learning (RFdiffusion and ProteinMPD):

  • Rather than predicting the structure of an existing natural sequence, generative models design entirely novel, stable tertiary folds with zero natural homologs from scratch.
  • De novo designed mini-proteins bind viral targets (such as SARS-CoV-2 spike protein or influenza hemagglutinin) with picomolar affinities, functioning as synthetic neutralizing therapeutics.

ยง6.5 Solid-Phase Peptide Synthesis (SPPS): The Merrifield Methodology & Coupling Reagents

R. Bruce Merrifield revolutionized protein chemistry in 1963 by inventing Solid-Phase Peptide Synthesis (SPPS, 1984 Nobel Prize), allowing automated, stepwise synthesis of long peptides on an insoluble polymeric resin support.

The Merrifield Reaction Cycle

1. Resin Anchoring: The C-terminal amino acid is covalently attached via its carboxylate to an insoluble chloromethylated polystyrene resin (Merrifield resin) or functionalized Wang/2-chlorotrityl resin.

2. Deprotection: The temporary N-terminal protecting group is selectively cleaved:

  • Boc Strategy: Cleaved by $50\%$ trifluoroacetic acid (TFA).
  • Fmoc Strategy (Modern Standard): Cleaved by $20\%$ piperidine in DMF via base-catalyzed E1cB elimination.

3. Coupling: The incoming amino acid, possessing an Fmoc-protected amino group and activated carboxyl group, is introduced in excess along with a coupling reagent:

  • Carbodiimides: DCC or DIC in the presence of HOBt or Oxyma.
  • Phosphonium / Uronium Salts: HBTU, HATU, or PyBOP in the presence of DIPEA base.

These reagents convert the carboxylate into an activated ester that couples rapidly ($>99.5\%$ yield per cycle) without racemization.

4. Washing: Excess reagents and soluble byproducts are washed away through a sintered glass filter, eliminating the need for chromatographic purification after each step.

5. Global Cleavage and Deprotection: Once the full sequence is assembled, the peptide is cleaved from the resin and permanent side-chain protecting groups ($t\text{Bu}$, Trt, Pbf) are removed simultaneously using concentrated TFA ($95\%$) containing scavengers (triisopropylsilane, EDT, $\text{H}_2\text{O}$).

ยง6.6 Protein Sequence Determination: Edman Degradation & Mass Spectrometry

Determining the primary amino acid sequence of a polypeptide is essential for structural biology and proteomics.

The Edman Degradation (Pehr Edman, 1950)

A cyclic, stepwise chemical sequencing method that removes one residue at a time from the unblocked N-terminus:

1. Coupling: Phenyl isothiocyanate (PITC, Edman's reagent) reacts with the free N-terminal amino group under mildly alkaline conditions ($\text{pH } 8.5 - 9.0$) to form a phenylthiocarbamoyl (PTC) peptide:

$$\text{Ph-N}=\text{C}=\text{S} + \text{H}_2\text{N-CH(R}_1)\text{-CONH}\cdots \to \text{Ph-NH-CS-NH-CH(R}_1)\text{-CONH}\cdots$$

2. Cleavage: Anhydrous trifluoroacetic acid (TFA) protonates the sulfur atom and induces nucleophilic attack of the thiocarbonyl sulfur onto the first peptide carbonyl carbon. The first peptide bond is cleaved under anhydrous conditions, releasing an anilinothiazolinone (ATZ) amino acid and leaving the intact remaining peptide chain shortened by one residue.

3. Conversion: The unstable ATZ-amino acid is extracted into organic solvent and heated with aqueous acid to rearrange into a stable phenylthiohydantoin (PTH) amino acid.

4. Identification: The PTH-amino acid is identified by reversed-phase HPLC or LC-MS against calibrated standards. The cycle is repeated sequentially for 30โ€“50 residues.

Modern Mass Spectrometry: MALDI-TOF & ESI-MS/MS

  • Electrospray Ionization (ESI) and Matrix-Assisted Laser Desorption/Ionization (MALDI) generate intact gas-phase peptide ions without thermal fragmentation.
  • Tandem Mass Spectrometry (MS/MS): Selected peptide ions are accelerated into a collision cell containing argon gas (Collision-Induced Dissociation, CID). Cleavage occurs predominantly at peptide backbone amide bonds, generating characteristic $b$-ions (retaining the N-terminus) and $y$-ions (retaining the C-terminus). The mass differences between successive peaks directly read out the amino acid sequence.

ยง6.7 Protein Structural Hierarchy: Secondary, Tertiary & Quaternary Folding Landscapes

Proteins fold into unique three-dimensional conformations governed by thermodynamic stability and non-covalent interactions.

The Four Levels of Protein Structure

1. Primary Structure: The covalent sequence of amino acids linked by peptide bonds.

2. Secondary Structure: Local spatial conformations stabilized by hydrogen bonds between backbone carbonyl oxygens ($\text{C}=\text{O}$) and amide nitrogens ($\text{N}-\text{H}$):

  • $\alpha$-Helix: Polypeptide chain coils in a right-handed helix with $3.6$ residues per turn (pitch $= 5.4\text{ \AA}$). Every backbone $\text{C}=\text{O}$ of residue $i$ forms an optimal linear hydrogen bond with the $\text{N}-\text{H}$ of residue $i+4$.
  • $\beta$-Pleated Sheet: Extended polypeptide chains aligned side-by-side, forming inter-strand hydrogen bonds. In antiparallel $\beta$-sheets, hydrogen bonds are linear and perpendicular ($180^\circ$); in parallel $\beta$-sheets, hydrogen bonds are slightly angled and weaker.

3. Tertiary Structure: The complete three-dimensional folding of a single polypeptide chain, driven by:

  • Hydrophobic Effect: Entropic release of structured water clathrates as non-polar aliphatic and aromatic side chains bury into the anhydrous protein core ($\Delta S_{\text{water}} > 0$).
  • Electrostatic Interactions (Salt Bridges): Attractive ionic bonds between oppositely charged side chains (e.g., $\text{Lys}^+ \cdots \text{Glu}^-$).
  • Disulfide Bridges: Covalent $-\text{S}-\text{S}-$ crosslinks formed by the oxidation of two cysteine sulfhydryl groups.

4. Quaternary Structure: The assembly of multiple folded polypeptide subunits into a functional oligomeric complex (e.g., hemoglobin $\alpha_2\beta_2$ tetramer).

ยง6.8 Protein Folding: Anfinsen's Dogma, Disulfide Pairing & Chaperones

The spontaneous self-assembly of a linear polypeptide into a biologically functional tertiary structure represents a central triumph of molecular thermodynamics.

Anfinsen's Dogma (Thermodynamic Hypothesis)

Christian Anfinsen (1972 Nobel Prize) demonstrated using bovine pancreatic ribonuclease A (124 residues, 4 native disulfide bridges) that:

  1. Denaturing ribonuclease in $8\text{ M urea}$ containing $\beta$-mercaptoethanol completely unfolds the protein and reduces all 4 disulfide bonds (loss of enzymatic activity).
  2. Removing urea and mercaptoethanol by dialysis under aerobic conditions allowed the protein to spontaneously refold into its native, catalytically active conformation with $100\%$ recovery of activity.
  3. If re-oxidation was conducted in the presence of $8\text{ M urea}$, the 8 cysteine residues paired randomly into 105 possible scrambled disulfide combinations, yielding $<1\%$ enzymatic activity. Subsequent addition of catalytic trace mercaptoethanol restored native folding.

Anfinsen's Conclusion: The native three-dimensional conformation of a protein is encoded entirely in its primary amino acid sequence and represents the global thermodynamic free energy minimum ($\Delta G^\circ_{\text{fold}} < 0$).

Levinthal's Paradox and Folding Funnels

Cyrus Levinthal calculated that if an unfolded 100-residue protein sampled all possible conformations randomly ($3^{200} \approx 10^{95}$ states at $10^{-13}\text{ s}$ per state), finding the native state would take $10^{75}$ years!

  • Proteins resolve this paradox by folding along a cooperative energy landscape (folding funnel):
  • Local secondary structural elements form rapidly ($10^{-6}\text{ s}$), guiding the polypeptide down a funnel of decreasing free energy through molten globule intermediates into the native minimum within milliseconds to seconds.

Molecular Chaperones

In the crowded cellular cytoplasm ($300\text{ mg/mL}$ protein), exposed hydrophobic patches of nascent folding intermediates risk irreversible aggregation. Heat shock proteins (Hsp70, GroEL/GroES chaperonins) bind exposed hydrophobic surfaces in an ATP-dependent cycle, isolating the folding polypeptide inside a hydrophilic chamber to allow safe folding without aggregation.

Intermediate Example 6.1: Strecker Synthesis Mechanism and Asymmetric Synthesis of (S)-Phenylalanine

(a) Write out the complete stepwise mechanism for the synthesis of racemic phenylalanine starting from phenylacetaldehyde ($\text{PhCH}_2\text{CHO}$), ammonium chloride ($\text{NH}_4\text{Cl}$), and sodium cyanide ($\text{NaCN}$). (b) To prepare enantiomerically pure $(S)$-phenylalanine directly, modern pharmaceutical chemistry employs chiral auxiliaries or chiral organocatalysts. Outline the catalytic asymmetric Strecker reaction using a chiral cyclic guanidine or BINOL-derived catalyst, explaining how facial selectivity is achieved during cyanide addition.

Step 1: Classical Strecker Mechanism

1. Imine Formation:

Phenylacetaldehyde reacts with ammonia (from $\text{NH}_4\text{Cl} + \text{NaCN}$ equilibrium):

$$\text{PhCH}_2\text{CHO} + \text{NH}_3 \rightleftharpoons \text{PhCH}_2\text{CH}=\text{NH} + \text{H}_2\text{O}$$

Protonation yields the reactive electrophilic iminium ion: $\text{PhCH}_2\text{CH}=\text{NH}_2^+$.

2. Nucleophilic Cyanide Addition:

Cyanide ion ($\text{CN}^-$) attacks the iminium carbon:

$$\text{PhCH}_2\text{CH}=\text{NH}_2^+ + \text{CN}^- \to \text{PhCH}_2\text{CH}(\text{NH}_2)\text{-CN} \quad (\alpha\text{-Aminonitrile})$$

3. Acidic Hydrolysis:

Refluxing with aqueous $\text{HCl}$ protonates the nitrile nitrogen, followed by nucleophilic addition of water to form the amide intermediate, which hydrolyzes to the carboxylic acid:

$$\text{PhCH}_2\text{CH}(\text{NH}_2)\text{-CN} + 2\text{ H}_2\text{O} + \text{H}^+ \to \text{PhCH}_2\text{CH}(\text{NH}_3^+)\text{-COOH} + \text{NH}_4^+$$

This yields racemic $(\pm)$-phenylalanine.

Step 2: Asymmetric Strecker Reaction

  1. In an asymmetric organocatalytic Strecker reaction (e.g., Jacobsen chiral thiourea catalyst or chiral BINOL-phosphoric acid):
  • The chiral catalyst forms a dual hydrogen-bonding network with the imine nitrogen and the cyanide nucleophile (e.g., $\text{HCN}$ or $\text{TMSCN}$).
  • The bulky BINOL/thiourea scaffold completely blocks one face of the imine (the si-face).
  • Nucleophilic attack by cyanide is restricted exclusively to the re-face.
  1. Hydrolysis of the resulting chiral aminonitrile yields (S)-phenylalanine in $>95\%$ enantiomeric excess ($ee$) without requiring racemic resolution.
Advanced Example 6.2: Multi-Prototropic Speciation and Isoelectric Point ($pI$) of Histidine

L-Histidine is a triprotic amino acid possessing three ionizable functional groups:

  • $\alpha\text{-COOH}$: $pK_{a1} = 1.82$
  • Imidazole ring side chain: $pK_{aR} = 6.00$
  • $\alpha\text{-NH}_3^+$: $pK_{a2} = 9.17$

(a) Write down the structures and net charges of all four ionic species ($\text{H}_3\text{His}^{2+}, \text{H}_2\text{His}^+, \text{HHis}^0, \text{His}^-$). (b) Derive the exact mathematical formula for the isoelectric point ($pI$) and calculate its numerical value for histidine. (c) At physiological $\text{pH } 7.40$, calculate the percentage of histidine molecules that have a positively charged imidazole side chain.

Step 1: Speciation and Charges of Histidine

  1. $\text{H}_3\text{His}^{2+}$ (at $\text{pH} < 1.82$): $\alpha\text{-COOH}$, protonated imidazole ($-\text{ImH}^+$), $\alpha\text{-NH}_3^+$. Net charge = $+2$.
  2. $\text{H}_2\text{His}^+$ (at $1.82 < \text{pH} < 6.00$): $\alpha\text{-COO}^-$, protonated imidazole ($-\text{ImH}^+$), $\alpha\text{-NH}_3^+$. Net charge = $+1$.
  3. $\text{HHis}^0$ (at $6.00 < \text{pH} < 9.17$): $\alpha\text{-COO}^-$, neutral imidazole ($-\text{Im}$), $\alpha\text{-NH}_3^+$. Net charge = $0$ (Zwitterion).
  4. $\text{His}^-$ (at $\text{pH} > 9.17$): $\alpha\text{-COO}^-$, neutral imidazole ($-\text{Im}$), unprotonated $\alpha\text{-NH}_2$. Net charge = $-1$.

Step 2: Derivation and Calculation of $pI$

At the isoelectric point, the concentrations of positively charged species must balance negatively charged species:

$$[\text{H}_2\text{His}^+] + 2[\text{H}_3\text{His}^{2+}] = [\text{His}^-]$$

Near $pI$ (between $6.0$ and $9.2$), $[\text{H}_3\text{His}^{2+}]$ is negligible. Thus:

$$[\text{H}_2\text{His}^+] \approx [\text{His}^-]$$

Expressing both in terms of the neutral zwitterion $[\text{HHis}^0]$:

$$[\text{H}_2\text{His}^+] = \frac{[\text{H}^+] [\text{HHis}^0]}{K_{aR}}, \quad [\text{His}^-] = \frac{K_{a2} [\text{HHis}^0]}{[\text{H}^+]}$$

Equating:

$$\frac{[\text{H}^+] [\text{HHis}^0]}{K_{aR}} = \frac{K_{a2} [\text{HHis}^0]}{[\text{H}^+]}$$
$$[\text{H}^+]^2 = K_{aR} \cdot K_{a2}$$

Taking negative logarithms:

$$pI = \frac{pK_{aR} + pK_{a2}}{2} = \frac{6.00 + 9.17}{2} = \frac{15.17}{2} = \mathbf{7.585} \approx \mathbf{7.59}$$

Step 3: Imidazole Protonation State at $\text{pH } 7.40$

Using the Henderson-Hasselbalch equation for the imidazole side chain ($pK_{aR} = 6.00$):

$$\text{pH} = pK_{aR} + \log\left(\frac{[\text{Im}]}{[\text{ImH}^+]}\right)$$
$$7.40 = 6.00 + \log\left(\frac{[\text{Im}]}{[\text{ImH}^+]}\right) \implies \log\left(\frac{[\text{Im}]}{[\text{ImH}^+]}\right) = 1.40$$
$$\frac{[\text{Im}]}{[\text{ImH}^+]} = 10^{1.40} = 25.12$$

Fraction protonated ($f_{\text{pos}}$):

$$f_{\text{pos}} = \frac{[\text{ImH}^+]}{[\text{Im}] + [\text{ImH}^+]} = \frac{1}{25.12 + 1} = \frac{1}{26.12} = 0.0383 \implies \mathbf{3.83\%}$$

At physiological pH 7.40, approximately $3.8\%$ of histidine side chains are protonated/positively charged, making histidine uniquely suited as a versatile general acid-base catalyst in enzyme active sites.

Intermediate Example 6.3: Peptide Bond Rotational Barrier and Double-Bond Resonance Energy

The rotational barrier about the central $\text{C}-\text{N}$ bond in formamide ($\text{HCONH}_2$) and model dipeptides is experimentally determined to be $\Delta G^\ddagger = 84.0\text{ kJ/mol}$ at $300\text{ K}$. (a) Calculate the rate constant of cis-trans isomerization $k_{\text{iso}}$ using the Eyring equation:

$$k = \frac{k_B T}{h} \exp\left(-\frac{\Delta G^\ddagger}{R T}\right)$$

(b) Explain why peptidyl-prolyl cis-trans isomerases (PPIases) are essential cellular enzymes during nascent protein folding in the endoplasmic reticulum.

Step 1: Rate Constant Calculation via Eyring Equation

At $T = 300\text{ K}$:

$$\frac{k_B T}{h} = \frac{(1.381\times 10^{-23}\text{ J/K})(300\text{ K})}{6.626\times 10^{-34}\text{ J}\cdot\text{s}} = 6.25\times 10^{12}\text{ s}^{-1}$$

The exponential factor:

$$\frac{\Delta G^\ddagger}{R T} = \frac{84,000\text{ J/mol}}{(8.314\text{ J/(mol}\cdot\text{K)})(300\text{ K})} = \frac{84,000}{2494.2} = 33.678$$
$$\exp(-33.678) = 2.36\times 10^{-15}$$

The isomerization rate constant is:

$$k_{\text{iso}} = (6.25\times 10^{12}\text{ s}^{-1}) \times (2.36\times 10^{-15}) = \mathbf{1.48\times 10^{-2}\text{ s}^{-1}}$$

The half-life for non-catalyzed cis-trans isomerization is:

$$t_{1/2} = \frac{\ln 2}{k_{\text{iso}}} = \frac{0.6931}{0.0148\text{ s}^{-1}} \approx \mathbf{46.8\text{ seconds}}$$

Step 2: Biological Role of Peptidyl-Prolyl Isomerases (PPIases)

  • In standard amino acids, steric clash forces $>99.9\%$ of peptide bonds into the trans conformation.
  • In proline residues, because the pyrrolidine ring bridges back to the nitrogen, the energy difference between cis and trans peptide bonds ($\Delta G^\circ$) is only $4-8\text{ kJ/mol}$, resulting in $10-20\%$ of native prolyl bonds occupying the cis conformation.
  • Because spontaneous uncatalyzed isomerization requires nearly a minute ($t_{1/2} \approx 47\text{ s}$), incorrect prolyl isomerization represents a severe kinetic bottleneck in protein folding.
  • PPIases (such as cyclophilins and FKBP) accelerate prolyl cis-trans isomerization by lowering the $\text{C-N}$ double-bond barrier, allowing proteins to achieve their native functional tertiary fold on physiological millisecond timescales.
Intermediate Example 6.4: Solid-Phase Peptide Synthesis Yield Compounding and Stepwise Efficiency

A peptide containing $N = 50$ amino acid residues is to be synthesized by automated solid-phase peptide synthesis (SPPS). (a) The overall yield of the target peptide is given by $Y_{\text{overall}} = (y_{\text{step}})^{N-1}$, where $y_{\text{step}}$ is the average fractional coupling yield per cycle. Calculate the overall yield if $y_{\text{step}} = 95.0\%$, and compare it to an optimized protocol where $y_{\text{step}} = 99.5\%$. (b) Calculate the maximum chain length $N$ that can be synthesized while maintaining an overall yield of at least $50.0\%$ when the coupling efficiency is $99.2\%$. (c) Explain the function of acetic anhydride 'capping' in preventing deletion peptides and simplifying chromatographic purification.

Step 1: Overall Yield for 50-Residue Peptide

For an $N = 50$ residue peptide, exactly $49$ coupling cycles are required:

1. At $95.0\%$ Coupling Efficiency ($y = 0.950$):

$$Y_{\text{overall}} = (0.950)^{49} = \mathbf{0.081} \implies \mathbf{8.1\%}$$

Over $91.9\%$ of the resin-bound material consists of truncated or deletion impurities!

2. At $99.5\%$ Optimized Coupling Efficiency ($y = 0.995$):

$$Y_{\text{overall}} = (0.995)^{49} = \mathbf{0.782} \implies \mathbf{78.2\%}$$

An increase of just $4.5\%$ in stepwise efficiency produces a nearly 10-fold increase in the recovery of the target peptide.

Step 2: Maximum Chain Length for $50\%$ Overall Yield

Given $y = 0.992$ and target $Y \ge 0.50$:

$$(0.992)^{N-1} \ge 0.50$$

Taking natural logarithms:

$$(N - 1) \ln(0.992) \ge \ln(0.50)$$
$$(N - 1) (-0.008032) \ge -0.69315$$
$$N - 1 \le \frac{0.69315}{0.008032} = 86.3$$
$$N \le 87.3 \implies \mathbf{N = 87\text{ residues}}$$

Step 3: Role of Capping with Acetic Anhydride

If an unreacted amino group fails to couple during cycle $k$, it will remain available to couple in cycle $k+1$, producing a deletion peptide missing residue $k$.

  • Deletion peptides differ from the full-length target by only a single amino acid, making their chromatographic separation (by HPLC) virtually impossible.
  • Treating the resin after each coupling step with acetic anhydride ($\text{Ac}_2\text{O} / \text{pyridine}$) acetylates all unreacted amine termini, permanently terminating their growth.
  • These capped, truncated fragments have vastly different molecular weights and retention times, enabling straightforward HPLC purification of the target peptide.
Intermediate Example 6.5: Edman Degradation Phenylthiohydantoin (PTH) Mechanism and Sequence Deduction

A purified pentapeptide isolated from an amphibian skin secretion was subjected to sequential automated Edman degradation. Chromatographic analysis of the released PTH-amino acid derivatives yielded:

  • Cycle 1: PTH-Tyr
  • Cycle 2: PTH-Gly
  • Cycle 3: PTH-Gly
  • Cycle 4: PTH-Phe
  • Cycle 5: PTH-Leu

(a) State the primary sequence of the pentapeptide. Identify this physiologically active neuropeptide. (b) Draw the chemical structure and curved-arrow mechanism for the conversion of the anilinothiazolinone (ATZ) intermediate into the stable phenylthiohydantoin (PTH) ring for Cycle 1.

Step 1: Sequence Deduction and Identification

Because the Edman degradation sequentially cleaves residues starting from the free N-terminus:

  • N-terminus (Residue 1) = Tyrosine (Tyr)
  • Residue 2 = Glycine (Gly)
  • Residue 3 = Glycine (Gly)
  • Residue 4 = Phenylalanine (Phe)
  • C-terminus (Residue 5) = Leucine (Leu)

The primary amino acid sequence is:

$$\mathbf{H_2N-Tyr-Gly-Gly-Phe-Leu-COOH \quad (YGGFL)}$$

This peptide is [Leu]-Enkephalin, an endogenous opioid pentapeptide that binds with high affinity to $\mu$- and $\delta$-opioid receptors in the central nervous system.

Step 2: Mechanism of ATZ $\to$ PTH Conversion

1. ATZ Structure:

Cleavage of the PTC-peptide with anhydrous TFA yields the 5-membered anilinothiazolinone (ATZ) intermediate:

$$\text{ATZ-Tyr contains a sulfur-containing 5-membered ring with a } \text{C}=\text{N-Ph bond and C}=\text{O bond.}$$

2. Hydrolytic Ring Opening:

In aqueous acid ($1.0\text{ M HCl}$ at $80^\circ\text{C}$), water attacks the carbonyl carbon of the ATZ ring, opening it to form the acyclic phenylthiocarbamoyl amino acid intermediate:

$$\text{ATZ} + \text{H}_2\text{O} \to \text{Ph-NH-CS-NH-CH}(\text{CH}_2\text{C}_6\text{H}_4\text{OH})\text{-COOH}$$

3. Recyclization through Nitrogen:

The aniline nitrogen atom ($\text{Ph-NH}-$) attacks the carboxylic acid carbonyl carbon with elimination of water ($\text{H}_2\text{O}$). This closes a stable five-membered phenylthiohydantoin (PTH) ring with the sulfur atom remaining exocyclic ($\text{C}=\text{S}$) and the carbonyl oxygen exocyclic ($\text{C}=\text{O}$). The resulting PTH-Tyrosine is chemically stable, UV-active at $269\text{ nm}$, and readily quantified by reversed-phase HPLC.

Advanced Example 6.6: Ramachandran Steric Contour Mapping and Dihedral Angle Constraints

Consider an L-alanine dipeptide model ($\text{Ac-Ala-NHMe}$). (a) Define the dihedral angles $\phi$ (Phi), $\psi$ (Psi), and $\omega$ (Omega) by specifying the four consecutive backbone atoms that define each torsion angle. (b) Explain why the conformation $(\phi = 0^\circ, \psi = 0^\circ)$ is strictly forbidden on the Ramachandran plot, identifying the specific steric clash that occurs. (c) Why does glycine occupy a much larger permissible area on the Ramachandran plot compared to all other amino acids, and why is proline severely restricted to $\phi \approx -65^\circ \pm 15^\circ$?

Step 1: Definition of Backbone Dihedral Angles

1. $\phi$ (Phi): Defined by atoms $\text{C}_{i-1} - \text{N}_i - \text{C}_{\alpha, i} - \text{C}_i$. Rotation about the $\text{N}-\text{C}_\alpha$ bond.

2. $\psi$ (Psi): Defined by atoms $\text{N}_i - \text{C}_{\alpha, i} - \text{C}_i - \text{N}_{i+1}$. Rotation about the $\text{C}_\alpha-\text{C}$ bond.

3. $\omega$ (Omega): Defined by atoms $\text{C}_{\alpha, i} - \text{C}_i - \text{N}_{i+1} - \text{C}_{\alpha, i+1}$. Rotation about the peptide amide bond (constrained by partial double bond to $\omega \approx 180^\circ$ for trans).

Step 2: The Forbidden $(\phi = 0^\circ, \psi = 0^\circ)$ Conformation

When $\phi = 0^\circ$ and $\psi = 0^\circ$:

  • The $\text{C}_{i-1}=\text{O}_{i-1}$ carbonyl group is eclipsed with the $\text{C}_i=\text{O}_i$ carbonyl group.
  • The distance between the carbonyl oxygen of residue $i-1$ and the amide nitrogen/hydrogen of residue $i+1$ drops below $1.8\text{ \AA}$, far below the sum of their van der Waals radii ($r_{\text{vdW}}(\text{O}) + r_{\text{vdW}}(\text{N}) \approx 3.0\text{ \AA}$).
  • This severe van der Waals overlap generates an enormous steric repulsion penalty ($>100\text{ kJ/mol}$), making $(0^\circ, 0^\circ)$ completely forbidden.

Step 3: Glycine vs Proline Conformational Flexibility

1. Glycine:

  • The side chain of glycine is a single hydrogen atom ($\text{R}=\text{H}$).
  • Lacking a bulky $\text{C}_\beta$ carbon, glycine experiences minimal steric hindrance across almost the entire $(\phi, \psi)$ landscape.
  • Glycine's Ramachandran plot is centrosymmetric, occupying over $60\%$ of conformational space, allowing it to adopt conformations forbidden to all other residues (e.g., tight turns and left-handed helices).

2. Proline:

  • In proline, the aliphatic side chain forms a rigid five-membered pyrrolidine ring that is covalently bonded to the backbone nitrogen atom ($\text{C}_\delta-\text{N}$).
  • This covalent ring locks rotation around the $\text{N}-\text{C}_\alpha$ bond, restricting $\phi$ strictly to $-65^\circ \pm 15^\circ$.
  • Consequently, proline acts as a conformational 'helix breaker' and is predominantly found at the initiation of helices or in tight $\beta$-turns.
Advanced Example 6.7: Protein Thermal Denaturation Thermodynamics and Melting Temperature ($T_m$)

A globular protein undergoes reversible two-state thermal unfolding:

$$\text{Native (N)} \rightleftharpoons \text{Denatured (D)}$$

Calorimetric measurements establish a denaturation enthalpy $\Delta H^\circ_{\text{unf}} = +420.0\text{ kJ/mol}$ and a denaturation entropy $\Delta S^\circ_{\text{unf}} = +1.280\text{ kJ/(mol}\cdot\text{K)}$ at the transition midpoint ($T_m$). (a) Calculate the thermal melting temperature ($T_m$) in degrees Celsius. (b) Calculate the Gibbs free energy of conformational stability ($\Delta G^\circ_{\text{unf}}$) at physiological temperature $T = 37.0^\circ\text{C}$ ($310.15\text{ K}$), assuming that the change in heat capacity $\Delta C_p$ is negligible over this temperature range. (c) Calculate the fraction of unfolded protein ($f_D$) present at $37.0^\circ\text{C}$.

Step 1: Melting Temperature ($T_m$) Calculation

At the denaturation midpoint ($T_m$), exactly half of the protein is native and half is unfolded ($[\text{N}] = [\text{D}] \implies K_{\text{unf}} = 1$ and $\Delta G^\circ = 0$):

$$\Delta G^\circ = \Delta H^\circ - T_m \Delta S^\circ = 0$$
$$T_m = \frac{\Delta H^\circ_{\text{unf}}}{\Delta S^\circ_{\text{unf}}} = \frac{420.0\text{ kJ/mol}}{1.280\text{ kJ/(mol}\cdot\text{K)}} = 328.125\text{ K}$$

Converting to Celsius:

$$T_m = 328.125 - 273.15 = \mathbf{54.98^\circ\text{C}} \approx \mathbf{55.0^\circ\text{C}}$$

Step 2: Gibbs Free Energy of Stability at $37.0^\circ\text{C}$

At $T = 310.15\text{ K}$ ($37.0^\circ\text{C}$):

$$\Delta G^\circ_{\text{unf}}(310.15\text{ K}) = \Delta H^\circ_{\text{unf}} - T \Delta S^\circ_{\text{unf}}$$
$$\Delta G^\circ_{\text{unf}} = 420.0\text{ kJ/mol} - (310.15\text{ K} \times 1.280\text{ kJ/(mol}\cdot\text{K)})$$
$$\Delta G^\circ_{\text{unf}} = 420.0 - 396.992 = \mathbf{+23.01\text{ kJ/mol}}$$

At body temperature, the native folded state is favored over the denatured state by $23.0\text{ kJ/mol}$ (the equivalent of approximately 4 to 5 hydrogen bonds).

Step 3: Fraction of Unfolded Protein ($f_D$)

The unfolding equilibrium constant is:

$$K_{\text{unf}} = \exp\left(-\frac{\Delta G^\circ_{\text{unf}}}{R T}\right) = \exp\left(-\frac{23,010}{(8.314)(310.15)}\right) = \exp(-8.923) = 1.333\times 10^{-4}$$

The fraction unfolded ($f_D$) is:

$$f_D = \frac{K_{\text{unf}}}{1 + K_{\text{unf}}} = \frac{1.333\times 10^{-4}}{1 + 1.333\times 10^{-4}} \approx \mathbf{1.33\times 10^{-4}} \implies \mathbf{0.0133\%}$$

Only about 1 in every 7,500 protein molecules is denatured at physiological temperature.

Advanced Example 6.8: Anfinsen's Ribonuclease A: Disulfide Pairing Combinatorics and Folding Yield

Bovine pancreatic ribonuclease A contains 124 amino acid residues including 8 cysteine residues that form 4 specific native disulfide crosslinks: Cys26-Cys84, Cys40-Cys95, Cys58-Cys110, and Cys65-Cys72. (a) Calculate the total number of mathematically possible, distinct pairings of 8 cysteines into 4 disulfide bonds. (b) If unfolded ribonuclease is re-oxidized under denaturing conditions (8 M urea), what is the theoretical probability of randomly forming the exact native set of 4 disulfide bonds? (c) When catalytic trace amounts of $\beta$-mercaptoethanol are added in the absence of urea, the scrambled inactive protein converts quantitatively into active ribonuclease A. Explain the thermodynamic driving force.

Step 1: Combinatorial Pairing Calculation

For $2n = 8$ cysteine residues to form $n = 4$ disulfide bonds:

  • The first cysteine can pair with any of the remaining $7$ cysteines ($7$ choices).
  • The next available cysteine can pair with any of the remaining $5$ cysteines ($5$ choices).
  • The next can pair with any of the remaining $3$ cysteines ($3$ choices).
  • The last two must pair together ($1$ choice).

The total number of distinct pairings is:

$$N = (2n - 1)!! = 7 \times 5 \times 3 \times 1 = \mathbf{105\text{ possible disulfide isomers}}$$

Alternatively, using factorials:

$$N = \frac{8!}{2^4 \times 4!} = \frac{40,320}{16 \times 24} = \frac{40,320}{384} = \mathbf{105}$$

Step 2: Random Probability in Urea

If oxidation occurs in $8\text{ M urea}$, the polypeptide chain has zero secondary/tertiary structure preference and behaves as a random coil. Each disulfide isomer forms with equal statistical probability:

$$P(\text{native}) = \frac{1}{105} \approx 0.00952 \implies \mathbf{0.95\%}$$

Less than $1\%$ of the scrambled protein possesses enzymatic activity.

Step 3: Thiol-Catalyzed Reshuffling and Thermodynamic Minimum

In native buffer without denaturant:

  • Addition of a catalytic trace of reducing thiol ($\text{R-SH}$, e.g., $\beta$-mercaptoethanol or protein disulfide isomerase, PDI) initiates reversible thiol-disulfide exchange:
$$\text{Protein-S-S-Protein} + \text{R-S}^- \rightleftharpoons \text{Protein-S-S-R} + \text{Protein-S}^-$$
  • This allows mismatched non-native disulfides to break and reform reversibly.
  • As the polypeptide explores conformational space, non-covalent interactions (hydrophobic burial, salt bridges, hydrogen bonds) stabilize the native tertiary fold.
  • Once the native pairing forms, it is locked inside the rigid tertiary structure where cysteines are shielded from further reduction.

Because the native fold represents the global thermodynamic free energy minimum ($\Delta G^\circ < 0$), the entire scrambled population is thermodynamically pulled into the $100\%$ active native state.

Advanced Example 6.9: Tandem Mass Spectrometry (MS/MS) CID Sequence Deduction of an Octapeptide

A bioactive antimicrobial peptide isolated from a frog skin secretion was analyzed by electrospray ionization tandem mass spectrometry (ESI-MS/MS). The singly protonated molecular ion $[M+H]^+$ has $m/z = 948.5$. Collision-Induced Dissociation (CID) generated the following prominent $b$-type fragment ion series ($m/z$):

  • $b_1 = 88.0$
  • $b_2 = 187.1$
  • $b_3 = 300.2$
  • $b_4 = 413.3$
  • $b_5 = 576.4$
  • $b_6 = 675.5$
  • $b_7 = 838.6$

Given the residue masses of standard amino acids: Gly = $57.0$, Ala = $71.0$, Val = $99.1$, Leu/Ile = $113.1$, Tyr = $163.1$, Phe = $147.1$, Trp = $186.1$, Ser = $87.0\text{ Da}$. (a) Determine the amino acid sequence of the peptide from the $b$-ion series. (b) Identify the C-terminal amino acid by calculating the mass difference between the intact $[M+H]^+$ ($948.5$) and the $b_7$ ion ($838.6$), and confirm the complete sequence.

Step 1: Sequence Deduction from Successive $b$-Ions

The mass differences between consecutive $b$-ions ($\Delta m = b_k - b_{k-1}$) correspond to the monoisotopic masses of the incorporated amino acid residues:

1. Residue 1 (N-terminus):

$b_1 = 88.0$. Since $b_1 = M(\text{Res}_1) + 1.0\text{ (proton)}$:

$$M(\text{Res}_1) = 88.0 - 1.0 = 87.0\text{ Da} \implies \mathbf{Serine\ (Ser)}$$

2. Residue 2:

$$\Delta m_2 = b_2 - b_1 = 187.1 - 88.0 = 99.1\text{ Da} \implies \mathbf{Valine\ (Val)}$$

3. Residue 3:

$$\Delta m_3 = b_3 - b_2 = 300.2 - 187.1 = 113.1\text{ Da} \implies \mathbf{Leucine\ or\ Isoleucine\ (Leu/Ile)}$$

4. Residue 4:

$$\Delta m_4 = b_4 - b_3 = 413.3 - 300.2 = 113.1\text{ Da} \implies \mathbf{Leucine\ or\ Isoleucine\ (Leu/Ile)}$$

5. Residue 5:

$$\Delta m_5 = b_5 - b_4 = 576.4 - 413.3 = 163.1\text{ Da} \implies \mathbf{Tyrosine\ (Tyr)}$$

6. Residue 6:

$$\Delta m_6 = b_6 - b_5 = 675.5 - 576.4 = 99.1\text{ Da} \implies \mathbf{Valine\ (Val)}$$

7. Residue 7:

$$\Delta m_7 = b_7 - b_6 = 838.6 - 675.5 = 163.1\text{ Da} \implies \mathbf{Tyrosine\ (Tyr)}$$

Step 2: C-Terminal Residue Identification

The intact peptide molecular ion is $[M+H]^+ = b_7 + M(\text{Res}_8) + M(\text{H}_2\text{O})$:

$$M(\text{Res}_8) = [M+H]^+ - b_7 - 18.02 = 948.5 - 838.6 - 18.0 = 91.9 \approx 92\text{ Da}$$

Wait, let us check:

$$[M+H]^+ - b_7 = 948.5 - 838.6 = 109.9\text{ Da}$$

In standard CID fragmentation, $[M+H]^+ - b_n = y_1 = M(\text{Res}_n) + 18.02 + 1.008 = M(\text{Res}_n) + 19.03$:

$$M(\text{Res}_8) = 109.9 - 19.0 = 90.9 \approx 91\text{ Da} \quad \text{or for amidated C-terminus: } M(\text{Res}_8) + 17.0 = 109.9 \implies M = 92.9$$

Wait, if the terminal amino acid is Alanine ($71.0$) or Glycine ($57.0$), what matches? If $y_1$ is $109.9$, notice $113.1 - 18.0 = 95.1$, but if it is an unmodified C-terminal amino acid:

$$\Delta m = 948.5 - 838.6 = 109.9\text{ Da}$$

For a free carboxylate C-terminus, the fragment missing from $b_7$ is $-\text{NH-CH(R)-COOH}$, which equals $M(\text{residue}) + \text{H}_2\text{O} = M + 18.02$. Thus:

$$M(\text{residue}) = 109.9 - 18.02 = 91.88 \approx 92\text{ Da}$$

(Wait, if it is an amidated C-terminal $-\text{NH}_2$, $109.9 - 17.03 = 92.8\text{ Da}$; or if it is a Valine derivative). The primary sequence through the 7 $b$-ions is unambiguously established as:

$$\mathbf{H_2N-Ser-Val-Leu-Leu-Tyr-Val-Tyr-[C-term]}$$

Solved Honors Problems & Derivations

Step-by-step rigorous solutions with full chemical, thermodynamic, and mechanistic validation.