What is a protein?

Your body is, in a very real sense, a protein machine. Proteins are the molecular workhorses of every living cell — they digest your food, carry oxygen in your blood, fight off infections, build your muscles, transmit nerve signals, copy your DNA, and regulate when genes turn on or off. Without proteins, life as we know it wouldn't exist.

At the most basic level, a protein is a long chain of smaller molecules called amino acids, folded into a specific three-dimensional shape. That 3-D shape is not decorative — it is the protein's function. Change the shape, and you change (or destroy) what the protein can do.

💡
Analogy: Swiss army knives Think of proteins as microscopic Swiss army knives. The chain of amino acids is like the raw metal, but it's only once it folds into its unique shape that the individual tools — the blade, the scissors, the screwdriver — can do anything useful. A protein "unfolded" is mostly useless; a protein in its correct 3-D shape is extraordinarily precise.

Amino acids — the 20-letter alphabet

All proteins in all living things are built from the same set of just 20 standard amino acids. Each amino acid is a small organic molecule with a central carbon atom, an amino group (–NH₂), a carboxyl group (–COOH), and a side chain that gives it its unique chemical personality.

You can think of amino acids as an alphabet of 20 letters. The sequence of amino acids in a protein is its primary structure — like a word written using this 20-letter alphabet. A typical protein might be anywhere from 50 to several thousand amino acids long.

🔬
Why exactly 20? The 20 standard amino acids are encoded by the genetic code — triplets of DNA bases (codons) each specify one amino acid. Evolution settled on 20 because they collectively provide a rich enough chemical diversity (hydrophobic, polar, charged, ringed, flexible, rigid…) to build essentially any molecular machine. Some exotic organisms use additional non-standard amino acids, but the core 20 are universal.

When amino acids are strung together in a chain, each one links to the next via a peptide bond — a covalent chemical bond between the carboxyl group of one amino acid and the amino group of the next, releasing a water molecule in the process. A chain of amino acids connected this way is called a polypeptide. A protein may consist of one or more polypeptide chains.

Leu — Gly — Ala — Val — Phe — Lys — Glu — Pro — … ↑ ↑ N-terminus C-terminus (start) (end) Each "—" is a peptide bond. The sequence runs from left (N) to right (C).

The side chains of the amino acids project outward from this backbone. Some are hydrophobic (water-fearing), some hydrophilic (water-loving), some positively charged, some negatively charged, some form disulfide bridges, some are ring-shaped and rigid. These chemical properties govern how the chain folds — they are the instructions hidden inside the sequence itself.

Four levels of protein structure

Structural biologists describe protein architecture at four levels of organisation. Each level builds on the previous one, like zooming out from individual bricks to a whole building.

Level 1
🔗
Primary
The linear sequence of amino acids. The raw "recipe."
Level 2
🌀
Secondary
Local folding into helices and sheets held by hydrogen bonds.
Level 3
🔮
Tertiary
The full 3-D shape of a single polypeptide chain.
Level 4
🧩
Quaternary
Multiple polypeptide chains assembled together.

Primary structure: the sequence

Primary structure is simply the order of amino acids along the polypeptide chain — nothing more. It's determined by the gene encoding the protein: DNA is transcribed into messenger RNA (mRNA), and the ribosome reads the mRNA to assemble amino acids in the coded order. Primary structure contains all the information needed to determine the final 3-D shape — this is Anfinsen's central insight from the 1950s, and why computational structure prediction is theoretically possible at all.

Secondary structure: helices and sheets

Once a segment of polypeptide has been synthesised, nearby portions of the chain begin to interact via hydrogen bonds — weak but numerous attractions between the slightly positive hydrogen atoms on one part of the backbone and slightly negative oxygen or nitrogen atoms elsewhere.

🔬
Hydrogen bonds: weak individually, powerful collectively A single hydrogen bond has about 1/20th the strength of a covalent bond. But proteins can form hundreds or thousands of them simultaneously. Together they create stable, specific structures — like how a zip works: each individual tooth is trivial to separate, but the whole zip holds firmly.

Two dominant patterns emerge from local hydrogen bonding:

  • Alpha helices (α-helices): The chain coils into a right-handed helix, like a spiral staircase. Every backbone N–H points toward a C=O four residues earlier in the chain. These are compact, rigid, and often found in membrane-spanning regions where the protein needs to cross the oily interior of a cell membrane.
  • Beta sheets (β-sheets): Stretches of chain line up side by side (either parallel or antiparallel) and zip together with hydrogen bonds between them, forming a flat, pleated sheet. Beta sheets provide structural rigidity — silk fibers are made almost entirely of beta sheets.
α-helix (viewed from the side): β-sheet (two strands): →→→ Strand 1: ————→———— ↗ ↘ | | | | | ↑ ↓ Strand 2: ←———————— ↖ ↗ ←←← (hydrogen bonds shown as | )

Regions without regular secondary structure are called loops or coils. These flexible linker regions are important too — they connect helices and sheets and often form the active sites where the protein does its work.

Tertiary structure: the full 3-D shape

Tertiary structure is the complete three-dimensional arrangement of all atoms in a single polypeptide chain. This is the "shape" people usually mean when they talk about protein structure.

It is stabilised by several types of interactions:

  • Hydrophobic effect: The dominant force. Hydrophobic side chains cluster in the protein's interior to avoid contact with water, collapsing the chain inward. This is why proteins fold spontaneously — it's thermodynamically favourable to bury the greasy bits.
  • Hydrogen bonds: Between side chains and between side chains and the backbone.
  • Ionic interactions: Attraction between oppositely charged side chains (salt bridges).
  • Disulfide bonds: Covalent S–S bonds between cysteine side chains — particularly important in secreted proteins that face harsh extracellular conditions.
  • Van der Waals forces: Very weak short-range attractions that add up across the tightly packed protein interior.

Quaternary structure: assemblies

Many functional proteins are not single chains but assemblies of multiple polypeptide subunits (called protomers or monomers). Haemoglobin, for example, has four subunits (two α-globin and two β-globin chains) that work cooperatively to pick up and release oxygen. The ribosome — the cell's protein factory — is an enormous complex of dozens of proteins plus RNA molecules.

🗝️
Key insight The information specifying a protein's shape is entirely encoded in its primary sequence. No external "shape template" is needed — given the right conditions, the chain will fold itself. This was demonstrated experimentally by Christian Anfinsen in the 1950s, earning him the 1972 Nobel Prize. It's also precisely why AlphaFold is possible: if sequence determines structure, a computer should be able to learn the mapping.

Shape = Function

A protein's three-dimensional shape is not an accident — it is the protein's function. The shape creates specific surfaces, grooves, and cavities that allow the protein to bind to other molecules with exquisite selectivity.

💡
Analogy: locks and keys Emil Fischer (1894) described enzyme-substrate interactions as a lock and key: only the right key (substrate molecule) fits the right lock (enzyme's active site). Modern biology has refined this to an "induced fit" model where the lock actually flexes slightly to grip the key — but the core point stands. The 3-D shape of the protein's active site determines precisely which molecules it can bind, react with, or transport.

Examples of shape-determined function:

Enzymes

Catalytic proteins whose active site provides a precisely shaped environment to speed up chemical reactions, sometimes by factors of 1012.

Antibodies

Y-shaped proteins whose variable tips (complementarity-determining regions) are shaped to grip specific antigens on pathogens.

Receptors

Membrane proteins whose extracellular domain is shaped to bind specific hormones or neurotransmitters, triggering a signal inside the cell.

Structural proteins

Collagen's triple-helix shape gives connective tissue its tensile strength. Keratin's coiled-coil shape makes hair tough.

When folding goes wrong: protein misfolding diseases

Now comes the part that makes protein structure prediction medically urgent. If a protein doesn't fold into its correct 3-D shape — whether due to a mutation in its gene, cellular stress, aging, or a chance error — it becomes misfolded. A misfolded protein can't do its job, and worse, it often becomes toxic.

⚠️
What "misfolding" actually means Misfolding doesn't mean the protein collapses into a random blob. It usually means the protein adopts an alternative stable shape — a different fold that is thermodynamically feasible but biologically useless or actively harmful. The new shape often exposes hydrophobic surfaces that should be buried, causing proteins to clump together into toxic aggregates.

Alzheimer's disease

A protein called amyloid-beta (Aβ) is normally produced in the brain and cleared away. In Alzheimer's disease, Aβ misfolds from a normal soluble form into a beta-sheet-rich shape that stacks into long, insoluble fibres called amyloid plaques. These plaques accumulate between neurons, disrupting communication. Concurrently, the tau protein inside neurons misfolds into neurofibrillary tangles. Together these cause the progressive neurodegeneration characteristic of Alzheimer's.

Parkinson's disease

A protein called alpha-synuclein normally plays a role in neurotransmitter release. In Parkinson's disease, it misfolds and aggregates into structures called Lewy bodies inside dopamine-producing neurons. Knowing the misfolded structure helps researchers design small molecules that might prevent the aggregation.

Prion diseases (mad cow disease, Creutzfeldt-Jakob)

Prions represent the most dramatic example of misfolding. The prion protein (PrPC) exists in a normal cellular form. When it misfolds into the pathological PrPSc form, it becomes infectious — it can physically recruit and convert nearby normal PrPC molecules into the misfolded form, spreading like a conformational chain reaction. This is genuinely caused by protein shape alone: no DNA, no RNA, just a misfolded protein propagating its aberrant shape.

Cystic fibrosis

Here misfolding comes from a genetic mutation. The most common CF mutation (ΔF508) deletes a single amino acid (phenylalanine at position 508) from the CFTR protein. This prevents CFTR from folding correctly in the cell, so it gets flagged for disposal before it ever reaches the cell surface. Cells lining the lungs and pancreas then lack functional CFTR channels, leading to thick, sticky mucus. Crucially, the drug ivacaftor/lumacaftor works partly by stabilising the misfolded CFTR protein, helping it fold correctly enough to function — a landmark example of structure-based drug design.

Type 2 diabetes

Islet amyloid polypeptide (IAPP / amylin), co-secreted with insulin by pancreatic beta cells, can misfold and form amyloid deposits in the islets of Langerhans. These deposits are found in the majority of type 2 diabetes patients and contribute to beta cell dysfunction.

🗝️
Why structure prediction helps To design drugs that prevent misfolding or break up aggregates, researchers need to know the exact atomic positions of both the normal and misfolded structures. Before AlphaFold, getting that information experimentally was prohibitively slow and often impossible for the misfolded forms. Now predictions can seed hypothesis-driven drug design within days.

Why we need to predict structure computationally

By the end of 2020, biology had catalogued roughly 180 million distinct protein sequences in databases — proteins from bacteria, viruses, plants, fungi, animals, humans. Yet the number of proteins with experimentally determined 3-D structures was only around 170,000.

That staggering gap — over 1,000 times more known sequences than known structures — existed because experimental structure determination is hard, slow, and expensive. X-ray crystallography can take months or years per protein. NMR spectroscopy is limited to small proteins. Even cryo-electron microscopy (cryo-EM), the revolutionary newcomer, requires expensive instruments and specialist expertise.

The genomics revolution had given us sequences faster than we could ever characterise them experimentally. Computational prediction was the only realistic path to closing the gap. That's the context for AlphaFold — not a curiosity, but an urgent medical and scientific necessity.

We'll examine the experimental methods in detail in Part 2, and trace the history of computational prediction attempts in Part 3. But first, let's understand precisely why this problem was considered so hard that solving it earned a Nobel Prize.

Key points from this chapter

  • Proteins are chains of amino acids that fold into precise 3-D shapes determining their function.
  • All proteins are built from the same 20 amino acids — the "20-letter alphabet" of life.
  • Structure has four levels: primary (sequence), secondary (helices/sheets), tertiary (full 3-D shape), quaternary (multi-chain complexes).
  • The dominant folding force is the hydrophobic effect: oily amino acids cluster in the interior to avoid water.
  • Anfinsen showed that amino acid sequence alone encodes the final structure — no external template needed.
  • Misfolded proteins cause diseases ranging from Alzheimer's and Parkinson's to cystic fibrosis and prion diseases.
  • The huge gap between known sequences (180M+) and known structures (~170K) motivated computational prediction.