MB&B 7200 Macromolecular Structure and Biophysical Analysis (Fall 2026)
Lecture 1 Primary structure and directed evolution
Summary
This lecture introduces the fundamental building blocks of proteins and post-translational modifications and turns to primary structure as an evolutionary record and how directed evolution and related methods exploit the protein sequence space.
Key Principles
Explain how the chemical properties of amino acids (charge, electronegativity, size, hydrophobicity) determine their behavior in protein structures and their roles in protein function
Explain how post-translational modifications expand the chemical repertoire of the 20 amino acids
Distinguish homologs, orthologs, and paralogs; explain how the significance of an alignment is assessed and why a substitution matrix detects distant relationships that identity-only scoring misses
Explain the concept of sequence space, calculate theoretical diversity, and describe how directed evolution and phage display sample a small fraction of it through cycles of selections
Briefly explain why a site-independent model of a multiple sequence alignment is incomplete, and how protein language models learn context-dependent rules directly from sequence data
Terms & Facts
1. Chemical properties of amino acids
The structures, three-letter and one-letter codes, and chemical classification of the 20 standard amino acids
Amino acids are zwitterions at neutral pH; know the pKa values of the ionizable sidechains
van der Waals radius: the closest approach of another atom without forming a bond. Sidechain van der Waals volumes
2. Ionic and dipole forces
Ionic interactions (salt bridges): electrostatic forces between charged residues; strong and long-range (energy decays as r⁻²); much weaker in water; and provide structural specificity
Hydrogen bonds: arise from the charge dipoles in bonds between atoms of different electronegativity (O > N > C > H); moderate strength; short-range (r⁻³ for fixed dipoles); directional; and weakened by water, which competes
London dispersion forces (van der Waals attraction): induced dipole–induced dipole; very short range (r⁻⁶); weak individually but dominant between non-charged groups such as the hydrophobic core of proteins
3. Post-translational modifications (PTMs)
PTMs act mostly on sidechains and extend protein chemistry: phosphorylation, N- and O-linked glycosylation, acetylation, methylation, ubiquitination, lipidation, and etc.
Cysteine forms reversible disulfide bonds by oxidation/reduction, providing covalent cross-links that stabilize structure.
Proteolytic cleavage as maturation, and modifications that create chemistry the polypeptide cannot supply
4. Homology, alignment, and BLOSUM
Orthologs are homologs in different species; paralogs are homologs within one species
Alignments score matches against a gap penalty. 15–25% identity is ambiguous where identity alone cannot establish homology
Shuffling: permute the residues of one sequence (length and composition preserved, order destroyed), realign and compare the real score with the resulting null distribution
BLOSUM (BLOcks SUbstitution Matrix): a log ratio of observed to random substitution frequency, so a score >0 means the substitution is tolerated more often than random and <0 less often. BLOSUM62 is the default matrix for protein BLAST, which reports an E value, the number of alignments of that quality expected by chance
[Altschul et al. Lipman. 1990 PMID: 2231712]
4. Sequence space, directed evolution, and phage display
Protein sequence space: 20^N possible sequences for a protein of N residues.
Directed evolution: iterative laboratory cycles of mutagenesis, selection, and amplification. On a fitness landscape, selection climbs toward peaks but can become trapped at a local maximum when diversification cannot cross a fitness valley
[Frances Arnold: Nobel Lecture in Chemistry 2018]
Deep mutational scanning (DMS): the functional effect of every possible single substitution, displayed as a fitness map or a Weblogo
[Molina et al. Liu. 2022 PMID: 37073402]
[Fowler and Fields. 2014 PMID: 25075907]
[Packer and Liu. 2015 PMID: 26055155]
Phage display links genotype (DNA encoded in the virion) to phenotype (peptide fused to the coat protein); used for peptides and for scFv or Fab antibodies
[George Smith: Nobel Lecture in Chemistry 2018]
[Sir Gregory P. Winter: Nobel Lecture in Chemistry 2018]
5. Learning evolutionary rules from data
A substitution matrix is position-independent; a multiple sequence alignment adds position-specific frequencies but a site-independent model still treats positions as unlinked
Epistasis: the effect of a substitution depends on sequence context, so site-independent models are incomplete
Protein language models estimate how likely a sequence is given all protein sequence data collected so far. Autoregressive models generate and score sequences; masked models learn generalizable representations; and sampling high-likelihood mutations concentrates a library on evolutionarily plausible variants
Lecture 2 Secondary structure and tertiary structure
Summary
This lecture explored peptide bond geometry, secondary structure elements, and the forces driving tertiary protein folding. Key topics included Ramachandran plots and specific structural motifs like coiled coils.
Key Principles
Explain how peptide bond planarity constrains protein backbone geometry to two variable angles (φ and ψ) per residue. Interpret Ramachandran plots to identify allowed φ–ψ angle combinations and recognize regions corresponding to α-helices, β-strands, and other conformations
Describe the structural characteristics of α-helices and β-sheets, including their hydrogen bonding patterns and geometric parameters
Explain how hydrophobic forces drive protein folding by minimizing unfavorable water-hydrophobic interfaces, while salt bridges and hydrogen bonds provide structural specificity
Describe how the oligomeric state of coiled coils can be controlled by varying the size and identity of core residues. Leucine at a and d positions favors dimers, while smaller residues like isoleucine or valine can favor trimers or tetramers.
Terms & Facts
1. Peptide bond and backbone geometry
Resonance gives the peptide bond partial double-bond character that restricts rotation around the C–N bond
The amide plane is planar with ω = 0° for cis, 180° for trans; trans configuration is strongly preferred (~99.97% of non-proline bonds); Proline has 5.2% cis bonds due to ring constraints
Dihedral angles φ (phi) and ψ (psi) define the local backbone geometry for each residue
Ramachandran plots: φ vs ψ angles showing sterically allowed conformations: α-helices cluster around φ = -60°, ψ = -60°; β-strands cluster around φ = -120°, ψ = +120°; Glycine has much greater conformational freedom due to lack of side chain
2. Secondary structure elements
[Eisenberg. 2003 PMID: 12966187]
α-helix: right-handed helix geometry; 3.6 residues per complete turn; 1.5 Å rise per residue; Hydrogen bonds between backbone C=O of residue i and N-H of residue i+4
β-sheet: antiparallel sheets: one residue to one residue hydrogen bonding; parallel sheets: one residue to two residues hydrogen bonding; side chains alternate above and below the sheet plane; Cα–Cα distance of ~3.5 Å between adjacent strands
Common secondary structures can be assessed by circular dichroism spectroscopy
3. Tertiary structure and protein folding
Hydrophobic core formation is driven by unfavorable entropy of water ordering around hydrophobic surfaces; burial of hydrophobic residues releases ordered water molecules, which creates a densely packed hydrophobic core in globular proteins
Four major protein fold classes: all α, all β, α/β, α+β
4. Coiled coils and leucine zippers
[Lupas & Bassler. 2017 PMID: 27884598]
Super-helical structure of two or more α-helices
Heptad repeat: Seven-residue repeat (abcdefg) characteristic of coiled coils
Positions a and d are typically hydrophobic
These hydrophobic residues form the core interface between helices
Other positions (b c e f g) are usually polar or charged
Knobs-into-holes: side chains from one helix (knobs) pack into spaces on partner helix (holes) that provides optimal van der Waals contacts between helices
Leucine zipper: characterized by leucines at every 7th position, e.g. GCN4
Be able to draw helical wheel presentations depicting multi-stranded coiled-coil structures: identify the heptad repeats (abcdefg), determine the wheel’s orientation, draw the wheel, and illustrate the coiled-coil interactions
Lecture 3 Quaternary structure, glycans, and glycoproteins
Summary
Proteins can assemble into higher-order structures through specific interactions between subunits. This lecture explored coiled-coil structures like the leucine zipper, quaternary protein assemblies with defined symmetries, and dynamic protein filaments. Carbohydrates can be attached to proteins and lipids on the surface of cells to present specific features that other cells can recognize.
Key Principles
Explain the spring-loaded mechanism of viral membrane fusion where a metastable pre-fusion conformation stores energy that drives membrane fusion when triggered, and how the 2P (proline) substitutions stabilize the pre-fusion state for vaccine applications
Distinguish between different symmetry groups (cyclic, dihedral, and cubic) and understand certain symmetries are favored in biological assemblies for coding efficiency, error control, and finite assembly.
Explain the difference between isologous and heterologous interactions in oligomer assembly and predict whether each type leads to closed or open (potentially infinite) structures
Explain how glycosylation is a post-translational modification that occurs co-translationally (N-linked in ER) and post-translationally (O-linked in Golgi)
Terms & Facts
1. Coiled coils and viral membrane fusion
Coiled-coil prediction from protein sequence using the observation that hydrophobic residues (especially leucine) appear with a periodicity of 3–4 residues in coiled-coil forming sequences
Viral fusion protein: a protein that mediates merger of viral and cellular membranes, containing a fusion peptide that insert into the target cell membrane and a coiled-coil region that extends for membrane fusion, e.g. influenza hemagglutinin (HA), HIV gp41, and SARS-CoV-2 Spike
2. Protein symmetry
Cyclic symmetry (Cn), dihedral symmetry (Dn), and cubic symmetry
Triangulation number (T): In icosahedral virus capsids, quantifies how many quasi-equivalent positions the subunit must adopt. Number of subunits = 60×T, where T = 1, 3, 4, 7, 13...
Quasisymmetry: In icosahedral viruses, the principle that allows more than 60 subunits by having subunits adopt slightly different but similar conformations to fill out larger capsids
[Goodsell & Olson. 2000 PMID: 10940245]
[Stenkamp. 2014 Protein Quaternary Structure: Symmetry Patterns. In eLS, John Wiley & Sons, Ltd (Ed.)]
3. Principles of protein oligomerization
Isologous interactions and heterologous interactions
Critical concentration (Kc): For polymerizing systems, the monomer concentration at which the rates of subunit addition and loss are balanced. Below this concentration, filaments depolymerize; above it, they grow
4. Glycans and protein glycosylation
Know the common monosaccharides in glycoproteins
N-linked glycosylation occurs on Asn in the -NXS/T- motif co-translationally in the ER; and O-linked glycosylation occurs on Ser/Thr post-translationally in the Golgi
Glycoconjugates on the surface of cells are recognized by other cells, viruses, and proteins.
Know the concept of biorthogonal chemistry as a method for studying glycans and glycoproteins in living systems
Lecture 6 Membranes and membrane proteins
Summary
This lecture covered the structural and functional principles of biological membranes, including lipid diversity and organization, lipid polymer assemblies (droplets, micelles, bilayers), and how proteins associate with and shape membranes through various mechanisms.
Key Principles
Explain how the hydrophobic effect drives lipid assembly into polymers based on molecular shape
Describe how phosphatidylinositol (PtdIns, PI) serves as subcellular localization signal, and how phosphatidylserine (PtdSer, PS) exposure on the outer leaflet serves as an "eat me" signal
Explain multiple mechanisms of curvature generation: lipid composition, protein scaffolding, helix insertion, and protein oligomerization
Describe how BAR (Bin/Amphiphysin/Rvs) domains can both sense and generate membrane curvature
Explain how ESCRT-III generates positive curvature through amphipathic helix insertion and negative curvature through polymerization
Terms & Facts
1. Phospholipids and sphingolipids
The structure of diacylglycerol phospholipids with two fatty acid chains attached to glycerol via ester bonds and the major head groups: choline, ethanolamine, serine, glycerol, and inositol
PI can be phosphorylated at positions 3, 4, and 5 of the inositol ring to yield PI3P, PI4P, PI(3,4)P₂, PI(4,5)P₂, and PI(3,4,5)P₃; different PI species localize to specific cellular compartments
Sphingolipids are based on sphingosine rather than glycerol, including ceramide, sphingomyelin, and glycosphingolipids; cholesterol as a sterol lipid that modulates membrane fluidity; and cardiolipin is a dimeric phospholipid found in mitochondrial membranes
2. Lipid polymers
Triglycerides (3 fatty acids on glycerol) form lipid droplets for energy storage
Detergents are micelle-forming amphiphiles used to solubilize and extract membrane proteins
Double-chain lipids form bilayers that assemble into vesicles
3. Membrane dynamics
[James Rothman, Randy Schekman, & Thomas Südhof, 2013 Nobel Prize in Physiology or Medicine]
Membrane asymmetry is actively maintained: lipids have high lateral mobility within a monolayer
Lipids have very slow flip-flop between leaflets
Flippases are ATP-dependent enzymes that move phospholipids from outer to inner leaflet
Floppases are ATP-dependent enzymes that move phospholipids from inner to outer leaflet
Scramblases are ATP-independent enzymes that equilibrate phospholipids between leaflets
PS is normally restricted to the cytoplasmic leaflet; PS exposure on the outer leaflet serves as an "eat me" signal and can facilitate cell–cell fusion
4. Membrane proteins
Membrane protein types: Single-pass transmembrane proteins (Type I: N-terminus out; Type II: C-terminus out), multi-pass transmembrane proteins with α-helical segments, β-barrel transmembrane proteins (cylindrical, typically 8–22 β-strands), monotopic membrane proteins associate with one leaflet via amphipathic helices, lipid-anchored proteins attach via covalent lipid modifications: myristoylation (Gly), S-palmitoylation (Cys), S-prenylation (Cys), GPI anchor (C-terminus), etc
Predicting transmembrane α-helices by identifying stretches of ~20+ consecutive hydrophobic residues
5. Membrane curvature
Positive (convex) and negative (concave) membrane curvatures shape organelles
Lipid composition can induce curvature: conical lipids favor positive curvature
Scaffolding proteins can impose curvature through their intrinsic shape
Amphipathic helix insertion, or wedging, creates asymmetry between leaflets, inducing curvature
BAR domains: banana-shaped dimeric domains that bind curved membranes; both sense existing curvature and generate curvature through scaffolding
ESCRT-III: an amphipathic helix generates positive membrane curvature; ESCRT-III forms filaments to generate negative curvature through polymerization
About the Study Guides
There is extensive information available about macromolecular structure and biophysical methods, but mastery comes from understanding core principles, not memorizing exhaustive details. Your goal should be to develop a conceptual framework for how biomacromolecules are organized and analyzed, focusing on the central concepts that unite biochemistry, biophysics, and structural biology.
My study guides for my lectures in MB&B 7200 will help you focus on the important principles and essential vocabulary. Each lecture highlights key concepts that form the foundation for the subsequent lectures in the course. Each lecture also introduces critical terms and facts that enable discussion of these concepts. These key terms are typically emphasized in lecture slides and are defined in the study guides. You should master the concepts and terms to the point where you can:
Recognize which principles apply to new biochemical, biophysical, or structural biology problems
Integrate concepts across lectures to solve problems we have not explicitly presented in class
Use appropriate terminology correctly when explaining your reasoning
Apply these principles to analyze macromolecular structure and function
Simply memorizing the study guide content is insufficient. You should expect to understand the underlying principles and be able to apply them. All homework and exams will emphasize the key concepts and terms in the study guides. I focus less on peripheral details from lectures and not at all on information from readings that isn’t covered in lectures or study guides. When studying, prioritize mastering the study guide material first. Additional details can enrich your understanding but should be secondary to core concepts. I design my homework and exams to test the important factual and conceptual knowledge covered in my lectures. I hope that by explaining to you what I test on and how I do it, you will spend your study time learning this important knowledge rather than focusing on minutia that you think might be on the exams.
Principles of macromolecules are fields of continuous discovery. Building expertise requires starting with foundational principles. These study guides are designed to help you establish that biochemical and biophysical foundation efficiently and effectively.
I learn alongside you during this course.
Steven Tang, October 3, 2025