The levels of protein structure
Statement
From amino-acid sequence to folded, functional shape.
Why it matters
This is the structural vocabulary the rest of the network's protein-related results depend on. protein-folding explains the thermodynamics that select one specific structure at this hierarchy's tertiary/quaternary levels; enzyme-lower-activation-energy presupposes a folded active site built from precisely positioned tertiary contacts; and dna-double-helix, lipids-amphipathic-assembly and carbohydrates-polysaccharides are this same unit's parallel accounts of how the other three macromolecule families' structure likewise determines their function. Proteins are the most structurally versatile of the four families specifically because they build a single covalent backbone (primary structure) up through four increasingly complex, non-covalent levels of organisation, each constrained by the one below it.
Understanding the hierarchy also explains, mechanistically, why a single amino-acid substitution can sometimes have a drastic functional effect and sometimes almost none at all: the consequence of a substitution depends entirely on which level of structure it disrupts.
Hypotheses
Proof
Result
Reading. Protein structure is organised into four nested levels, each built from and dependent on the one before it, with each level's stabilising forces shifting from purely covalent (primary) to backbone hydrogen bonding (secondary) to side-chain interactions (tertiary and quaternary).
Scope. Every protein has primary and (generically) some secondary structure; single-chain proteins lack quaternary structure by definition, and some short or intrinsically disordered regions of a chain lack stable, well-defined tertiary structure at all under physiological conditions.
Corollaries & converses
- protein-folding supplies the thermodynamic mechanism (free-energy minimisation, driven substantially by the hydrophobic effect) that selects the specific tertiary structure this result describes only geometrically; the two results are complementary, structure and mechanism, rather than competing accounts.
- enzyme-lower-activation-energy's active site is a tertiary-structure feature: residues that are far apart in the primary sequence are frequently brought close together in three-dimensional space once the chain is folded, which is precisely how a small number of catalytic residues can be positioned with the geometric precision an active site requires.
- Converse: observing a specific, reproducible secondary-structure pattern (a run of backbone hydrogen bonds with a \(3.6\)-residue period, Step 2) is itself sufficient evidence to infer an \(\alpha\)-helix is present, without needing to examine side-chain identity at all — secondary structure is a backbone-level, sequence-independent geometric signature.
Fails without
- Drop the level-dependency hypothesis and treat tertiary structure as though it could form independently of secondary structure: since tertiary structure is defined as the three-dimensional packing of secondary-structure elements (and connecting loops), there is no tertiary fold to speak of without first having some backbone (helix, sheet, or defined loop) geometry to pack — the hierarchy is not merely a naming convention but reflects a genuine physical dependency.
- Assume side-chain (R-group) interactions, rather than backbone hydrogen bonding, stabilise secondary structure (second Hypothesis): this fails to explain why the \(\alpha\)-helix and \(\beta\)-sheet are found, essentially unchanged in geometry, across proteins built from wildly different amino-acid sequences — if side chains were the stabilising force, secondary structure would be sequence-specific and irregular in the way tertiary structure actually is, rather than a small set of standard, repeating backbone motifs common to nearly all proteins.
Common errors
- Assuming quaternary structure is a universal fourth stage every protein passes through; many functional proteins are single-chain (monomeric) and simply have no quaternary structure at all (Result's Scope).
- Confusing the peptide bond (Step 1, covalent, part of primary structure) with the hydrogen bonds that stabilise secondary and tertiary structure (Steps 2–3, non-covalent); only the former survives most denaturing conditions.
- Assuming denaturation (loss of higher-order structure) necessarily breaks the primary sequence; in most cases only non-covalent interactions (and, if present, disulfide bonds) are disrupted, leaving the covalent peptide backbone (and hence primary structure) intact, which is why refolding can restore full activity (protein-folding, Anfinsen's experiment).
- Attributing the \(\alpha\)-helix's stability mainly to side-chain packing rather than to the regular backbone hydrogen-bonding pattern that actually defines it (Fails without, second bullet).
Discussion
Linus Pauling and Robert Corey predicted the precise geometry of the \(\alpha\)-helix and \(\beta\)-sheet in 1951, working from careful modelling of peptide-bond bond lengths and angles rather than from a solved protein structure, before any full protein structure had been experimentally determined; both predictions were subsequently confirmed by x-ray-crystallography once such structures became available, John Kendrew's myoglobin structure (1958) being the first. The four-level hierarchy (primary/secondary/tertiary/quaternary) itself was proposed by Kaj Ulrik Linderstrøm-Lang in the early 1950s as an organising framework for exactly this kind of emerging structural data.
Not every functional region of every protein adopts a single, stable tertiary structure. Intrinsically disordered regions — stretches of sequence that remain conformationally flexible even under native, physiological conditions — are now recognised as functionally important in their own right, particularly in signalling and regulatory proteins, where flexibility itself (rather than a single fixed fold) can be the functionally relevant property, complicating any simple assumption that every protein sequence specifies exactly one rigid tertiary structure.
Common misconception: that primary structure alone (the amino-acid sequence) is "the" structure of a protein in any functionally meaningful sense. Sequence is necessary but not sufficient information for function: two very different sequences can converge on similar tertiary folds and similar functions, while, conversely, disrupting higher-order structure without changing the sequence at all (denaturation) is usually enough to abolish a protein's biological activity completely.
Worked examples
Reading. A single change at the lowest level of the hierarchy can, depending on where in the folded structure it falls, propagate all the way to the highest level, exactly as Step 5 predicts.
Scope. Not every substitution has this severity; a substitution at a solvent-exposed, structurally unimportant site (Fails without's logic, applied in reverse) frequently has little to no detectable effect on tertiary or quaternary structure.
Problems
- A mutation substitutes one hydrophobic residue on a protein's fully solvent-exposed surface for a different hydrophobic residue of similar size. Predict, using the hierarchy, whether this substitution is more or less likely to disrupt tertiary structure than the sickle-cell substitution in Worked Example 1, and explain why.
Solution
Considerably less likely to disrupt structure. The sickle-cell substitution introduced a hydrophobic residue at a site that had been charged and solvent-exposed, creating a novel sticky patch (Worked Example 1). A hydrophobic-to-hydrophobic substitution at an already solvent-exposed site changes little about the local chemistry that site presents to the surrounding water or to other protein surfaces, and does not affect the buried hydrophobic core packing that stabilises tertiary structure (Step 3), so it is expected to be comparatively well tolerated. - Explain, using Step 2, why the period of the \(\alpha\)-helix (\(3.6\) residues per turn) being a non-integer number is functionally significant, rather than merely an odd numerical curiosity.
Solution
Because \(3.6\) is not a whole number, side chains projecting outward from successive turns of the helix are not stacked directly above one another but are instead offset around the helix's circumference, spreading side chains from the helix at a range of rotational positions. This allows an \(\alpha\)-helix to present, for example, a hydrophobic face on one side and a hydrophilic face on the other along its length (an amphipathic helix), a geometric consequence of the specific non-integer periodicity that a helix with an exact integer period of residues per turn would not produce. - A protein is treated with a reagent that reduces (breaks) all of its disulfide bonds, but no denaturant is applied and the protein's overall fold is otherwise undisturbed. Which level(s) of structure, if any, are directly affected? Explain.
Solution
Disulfide bonds are one of the side-chain (R-group) interactions stabilising tertiary structure (Step 3), so breaking them directly affects tertiary structure (and, for proteins where interchain disulfides link separate subunits, potentially quaternary structure, Step 4). Primary structure (Step 1, the peptide backbone) and the backbone hydrogen bonding underlying secondary structure (Step 2) are unaffected by disulfide reduction alone, since disulfides are a distinct, side-chain-based covalent interaction, not part of either of those two levels.