What is a protein?
Your body is, in a very real sense, a protein machine. Proteins are the molecular workhorses of every living cell — they digest your food, carry oxygen in your blood, fight off infections, build your muscles, transmit nerve signals, copy your DNA, and regulate when genes turn on or off. Without proteins, life as we know it wouldn't exist.
At the most basic level, a protein is a long chain of smaller molecules called amino acids, folded into a specific three-dimensional shape. That 3-D shape is not decorative — it is the protein's function. Change the shape, and you change (or destroy) what the protein can do.
Amino acids — the 20-letter alphabet
All proteins in all living things are built from the same set of just 20 standard amino acids. Each amino acid is a small organic molecule with a central carbon atom, an amino group (–NH₂), a carboxyl group (–COOH), and a side chain that gives it its unique chemical personality.
You can think of amino acids as an alphabet of 20 letters. The sequence of amino acids in a protein is its primary structure — like a word written using this 20-letter alphabet. A typical protein might be anywhere from 50 to several thousand amino acids long.
When amino acids are strung together in a chain, each one links to the next via a peptide bond — a covalent chemical bond between the carboxyl group of one amino acid and the amino group of the next, releasing a water molecule in the process. A chain of amino acids connected this way is called a polypeptide. A protein may consist of one or more polypeptide chains.
The side chains of the amino acids project outward from this backbone. Some are hydrophobic (water-fearing), some hydrophilic (water-loving), some positively charged, some negatively charged, some form disulfide bridges, some are ring-shaped and rigid. These chemical properties govern how the chain folds — they are the instructions hidden inside the sequence itself.
Four levels of protein structure
Structural biologists describe protein architecture at four levels of organisation. Each level builds on the previous one, like zooming out from individual bricks to a whole building.
Primary structure: the sequence
Primary structure is simply the order of amino acids along the polypeptide chain — nothing more. It's determined by the gene encoding the protein: DNA is transcribed into messenger RNA (mRNA), and the ribosome reads the mRNA to assemble amino acids in the coded order. Primary structure contains all the information needed to determine the final 3-D shape — this is Anfinsen's central insight from the 1950s, and why computational structure prediction is theoretically possible at all.
Secondary structure: helices and sheets
Once a segment of polypeptide has been synthesised, nearby portions of the chain begin to interact via hydrogen bonds — weak but numerous attractions between the slightly positive hydrogen atoms on one part of the backbone and slightly negative oxygen or nitrogen atoms elsewhere.
Two dominant patterns emerge from local hydrogen bonding:
- Alpha helices (α-helices): The chain coils into a right-handed helix, like a spiral staircase. Every backbone N–H points toward a C=O four residues earlier in the chain. These are compact, rigid, and often found in membrane-spanning regions where the protein needs to cross the oily interior of a cell membrane.
- Beta sheets (β-sheets): Stretches of chain line up side by side (either parallel or antiparallel) and zip together with hydrogen bonds between them, forming a flat, pleated sheet. Beta sheets provide structural rigidity — silk fibers are made almost entirely of beta sheets.
Regions without regular secondary structure are called loops or coils. These flexible linker regions are important too — they connect helices and sheets and often form the active sites where the protein does its work.
Tertiary structure: the full 3-D shape
Tertiary structure is the complete three-dimensional arrangement of all atoms in a single polypeptide chain. This is the "shape" people usually mean when they talk about protein structure.
It is stabilised by several types of interactions:
- Hydrophobic effect: The dominant force. Hydrophobic side chains cluster in the protein's interior to avoid contact with water, collapsing the chain inward. This is why proteins fold spontaneously — it's thermodynamically favourable to bury the greasy bits.
- Hydrogen bonds: Between side chains and between side chains and the backbone.
- Ionic interactions: Attraction between oppositely charged side chains (salt bridges).
- Disulfide bonds: Covalent S–S bonds between cysteine side chains — particularly important in secreted proteins that face harsh extracellular conditions.
- Van der Waals forces: Very weak short-range attractions that add up across the tightly packed protein interior.
Quaternary structure: assemblies
Many functional proteins are not single chains but assemblies of multiple polypeptide subunits (called protomers or monomers). Haemoglobin, for example, has four subunits (two α-globin and two β-globin chains) that work cooperatively to pick up and release oxygen. The ribosome — the cell's protein factory — is an enormous complex of dozens of proteins plus RNA molecules.
Shape = Function
A protein's three-dimensional shape is not an accident — it is the protein's function. The shape creates specific surfaces, grooves, and cavities that allow the protein to bind to other molecules with exquisite selectivity.
Examples of shape-determined function:
Enzymes
Catalytic proteins whose active site provides a precisely shaped environment to speed up chemical reactions, sometimes by factors of 1012.
Antibodies
Y-shaped proteins whose variable tips (complementarity-determining regions) are shaped to grip specific antigens on pathogens.
Receptors
Membrane proteins whose extracellular domain is shaped to bind specific hormones or neurotransmitters, triggering a signal inside the cell.
Structural proteins
Collagen's triple-helix shape gives connective tissue its tensile strength. Keratin's coiled-coil shape makes hair tough.
When folding goes wrong: protein misfolding diseases
Now comes the part that makes protein structure prediction medically urgent. If a protein doesn't fold into its correct 3-D shape — whether due to a mutation in its gene, cellular stress, aging, or a chance error — it becomes misfolded. A misfolded protein can't do its job, and worse, it often becomes toxic.
Alzheimer's disease
A protein called amyloid-beta (Aβ) is normally produced in the brain and cleared away. In Alzheimer's disease, Aβ misfolds from a normal soluble form into a beta-sheet-rich shape that stacks into long, insoluble fibres called amyloid plaques. These plaques accumulate between neurons, disrupting communication. Concurrently, the tau protein inside neurons misfolds into neurofibrillary tangles. Together these cause the progressive neurodegeneration characteristic of Alzheimer's.
Parkinson's disease
A protein called alpha-synuclein normally plays a role in neurotransmitter release. In Parkinson's disease, it misfolds and aggregates into structures called Lewy bodies inside dopamine-producing neurons. Knowing the misfolded structure helps researchers design small molecules that might prevent the aggregation.
Prion diseases (mad cow disease, Creutzfeldt-Jakob)
Prions represent the most dramatic example of misfolding. The prion protein (PrPC) exists in a normal cellular form. When it misfolds into the pathological PrPSc form, it becomes infectious — it can physically recruit and convert nearby normal PrPC molecules into the misfolded form, spreading like a conformational chain reaction. This is genuinely caused by protein shape alone: no DNA, no RNA, just a misfolded protein propagating its aberrant shape.
Cystic fibrosis
Here misfolding comes from a genetic mutation. The most common CF mutation (ΔF508) deletes a single amino acid (phenylalanine at position 508) from the CFTR protein. This prevents CFTR from folding correctly in the cell, so it gets flagged for disposal before it ever reaches the cell surface. Cells lining the lungs and pancreas then lack functional CFTR channels, leading to thick, sticky mucus. Crucially, the drug ivacaftor/lumacaftor works partly by stabilising the misfolded CFTR protein, helping it fold correctly enough to function — a landmark example of structure-based drug design.
Type 2 diabetes
Islet amyloid polypeptide (IAPP / amylin), co-secreted with insulin by pancreatic beta cells, can misfold and form amyloid deposits in the islets of Langerhans. These deposits are found in the majority of type 2 diabetes patients and contribute to beta cell dysfunction.
Why we need to predict structure computationally
By the end of 2020, biology had catalogued roughly 180 million distinct protein sequences in databases — proteins from bacteria, viruses, plants, fungi, animals, humans. Yet the number of proteins with experimentally determined 3-D structures was only around 170,000.
That staggering gap — over 1,000 times more known sequences than known structures — existed because experimental structure determination is hard, slow, and expensive. X-ray crystallography can take months or years per protein. NMR spectroscopy is limited to small proteins. Even cryo-electron microscopy (cryo-EM), the revolutionary newcomer, requires expensive instruments and specialist expertise.
The genomics revolution had given us sequences faster than we could ever characterise them experimentally. Computational prediction was the only realistic path to closing the gap. That's the context for AlphaFold — not a curiosity, but an urgent medical and scientific necessity.
We'll examine the experimental methods in detail in Part 2, and trace the history of computational prediction attempts in Part 3. But first, let's understand precisely why this problem was considered so hard that solving it earned a Nobel Prize.
Key points from this chapter
- Proteins are chains of amino acids that fold into precise 3-D shapes determining their function.
- All proteins are built from the same 20 amino acids — the "20-letter alphabet" of life.
- Structure has four levels: primary (sequence), secondary (helices/sheets), tertiary (full 3-D shape), quaternary (multi-chain complexes).
- The dominant folding force is the hydrophobic effect: oily amino acids cluster in the interior to avoid water.
- Anfinsen showed that amino acid sequence alone encodes the final structure — no external template needed.
- Misfolded proteins cause diseases ranging from Alzheimer's and Parkinson's to cystic fibrosis and prion diseases.
- The huge gap between known sequences (180M+) and known structures (~170K) motivated computational prediction.