Skip to content

Rob Phillips and the Physical Biology of the Cell

Woolf Software
A studio-lit feathered dinosaur specimen with feathers in a strict grid pattern, facing brass calipers and weights on wires.

This is the first in a series for readers who are comfortable with equations and want to see biology treated as a system that can be modeled, measured, and predicted. The prompt is a talk by Rob Phillips of Caltech, “Rob Phillips: Physical Biology of the Cell”. He gave it at the Kavli Institute for Theoretical Physics on February 26, 2011 at a conference for high school physics and biology teachers. KITP posted the recording in December 2017. Phillips wrote the textbook of the same name with Jane Kondev and Julie Theriot and Hernan Garcia (Garland Science, 2nd edition, 2012). His program has two explicit moves. First, estimate the numbers before writing any equation. Second, when you do write an equation, make it statistical mechanics and find an experiment with a knob you can turn to test it.

The talk is aimed at teachers and is light on derivations. The sections below follow its order and fill in the published work behind each step.

The argument of the talk

Phillips opens with an analogy. Tycho Brahe was dissatisfied with the astronomical tables of his day and built better instruments. He measured the position of Mars carefully enough that Kepler could later extract the laws of planetary motion from the data. Phillips’s claim is that biology is living through the same kind of period now. The ability to read and write DNA has rewritten the subject within a lifetime, and an entire third domain of life was found by comparing ribosomal RNA sequences. He points the audience to a six-page paper estimating that there are on the order of 103010^{30} prokaryotic cells on Earth 1.

The point of the analogy is that new instruments produce quantitative data, and quantitative data changes what theory has to do. His example is a plot of gene expression in bacteria in which two repressor binding sites were moved apart one base pair at a time. Expression rose and fell with a period of about 10.5 base pairs 2. The period is the helical repeat of DNA. Move a site five base pairs and it has rotated roughly 180 degrees around the helix, so a protein bound at one site can no longer reach the other. Phillips says he understands the period but not the position of the peak. The question he wants students to ask is whether either one surprises them, because in his words it is “inappropriate to respond to this kind of data with words alone”.

To be surprised you need a prior expectation. His illustration is the specific heat of solids, which Dulong and Petit measured as three times the gas constant and which classical statistical mechanics predicts exactly. The later observation that it vanishes at low temperature was therefore an anomaly that only a theory could have flagged. Phillips’s word for this kind of prior is “prejudice”, and he says his goal as a teacher is to give students prejudices of this sort.

Order of magnitude first

The second move is estimation. Before any statistical mechanics, Phillips wants a number for every size and rate and copy number. His main case study in the talk is how a bacterial virus packs its genome. A lambda phage genome is roughly 16 micrometers of DNA, and it sits inside a capsid about 50 nanometers across at about 50 percent volume fraction (golf balls poured into a bathtub reach 73 percent). Two physical facts set the cost of doing this. DNA is stiff on the 50 nanometer scale, meaning its persistence length (the distance over which the chain remembers its direction) is comparable to the capsid itself. It also carries one negative charge every 0.17 nanometers, so folding it against itself costs electrostatic energy.

Those two energies add up to a measurable force, and the experiment that Phillips says changed his career measured it. Carlos Bustamante’s group held a single phi29 phage in an optical trap and watched its motor pack DNA at about 100 base pairs per second against a rising internal force, consuming one ATP per base pair 3. Phillips’s description of the force curve is that you should think of it as compressing a spring, and that the area under it estimates the work done. His own group then reversed the experiment and watched single lambda phages eject their DNA into solution. That let them measure the ejection rate as a function of two knobs: genome length (the 48.5 kilobase wild type against a shorter mutant) and the counterion (sodium screens the charge less well than magnesium and gives faster ejection) 4.

The random walk is the second estimate in the talk. Phillips describes asking students to chain paper clips together, shake the chain on a table, and measure the size of the blob against the number of clips. The size scales as the square root of the number of segments, and the same scaling applied to a micrograph of a genome that has burst out of its cell gives the genome length from a single picture. This kind of reasoning is the whole content of Phillips’s book with Ron Milo, Cell Biology by the Numbers (Garland Science, 2015). The raw inputs live in the BioNumbers database that Milo and colleagues built, where every entry carries a literature source 5.

Counting molecules with coin flips

Between the estimate and the model there is a measurement problem. Fluorescence microscopy gives an intensity in arbitrary units, and Phillips says plainly in the talk that he hates seeing “a.u.” on an axis. The thermodynamic model of gene regulation needs the number of transcription factor molecules in the cell. Without a conversion factor between the two, the model cannot be tested.

The solution he describes comes from Rosenfeld, Elowitz, and colleagues 6. When a bacterium divides, each protein molecule ends up in one daughter or the other. If the partitioning is random it follows a binomial distribution, whose variance is set by the number of trials. So the squared mismatch in fluorescence between two daughters, averaged over many divisions, is proportional to the total fluorescence with a slope equal to the intensity per molecule. Phillips shows a plot with one point per division event and a fitted line giving the calibration. He mentions flipping a hundred coins on his bedroom floor the previous morning to make the point that coin-flip statistics let you count molecules you cannot see.

The thermodynamic model of simple repression

The model that Phillips’s group has tested most thoroughly is the thermodynamic model of transcription. The idea goes back to a 1982 model of the lambda phage repressor by Ackers and colleagues, which assigned each arrangement of proteins on the DNA a Boltzmann weight 7. Bintu and colleagues generalized it to a catalog of regulatory architectures 8. The central assumption is that the expression of a gene is proportional to the equilibrium probability that RNA polymerase is bound to its promoter. Statistical mechanics gives that probability as a sum over states, and every regulatory protein enters through a “regulation factor” multiplying the polymerase weight.

For the simplest architecture, a single repressor binding site overlapping the promoter, the sum can be done by hand. The observable is the fold-change: expression with repressor divided by expression without it. The slide in the talk gives it as

fold-change=(1+RNexp(ΔεkBT))1\text{fold-change} = \left(1 + \frac{R}{N}\,\exp\left(-\frac{\Delta\varepsilon}{k_B T}\right)\right)^{-1}

Here RR is the number of repressors in the cell and NN is the number of nonspecific binding sites on the genome (a few million for E. coli). Δε\Delta\varepsilon is the binding energy of the repressor at its specific site relative to a nonspecific one, and kBTk_B T is thermal energy. The derivation is short: the specific site is occupied with a probability set by its Boltzmann weight relative to the nonspecific background, and the promoter is available only when it is not.

Garcia and Phillips tested this equation on the lac operon in a way that reverses the usual logic 9. Instead of fitting the model to expression data, they used it as a predictive tool. Given the binding energy for an operator sequence, the model predicts how many repressors a strain must contain to produce a given fold-change. They built strains spanning nearly four orders of magnitude in fold-change and about two in repressor number, then checked the predicted copy numbers independently with quantitative immunoblots. The outcomes were in vivo binding energies for several repressor binding sites, a census of Lac repressor in E. coli, and a map of where the simple-repression input-output relation holds.

A later review from the same group spells out the methodological point 10. It describes a hierarchy of experiments in which the small parameter set learned at one level is used to make predictions at the next, with no refitting. That is what separates a physical model from a curve fit.

Allostery: the Monod-Wyman-Changeux model and induction

The equation above treats the repressor as always able to bind. Real repressors are switched off by small molecules, and the lac repressor is switched off by allolactose or its synthetic analog IPTG. In 1965 Monod and Wyman and Changeux proposed that such proteins exist in two conformations in equilibrium with each other, and that a ligand shifts the balance by binding the two states with different affinities 11. That is the Monod-Wyman-Changeux (MWC) model. It is a two-state model in exactly the sense Phillips praises the Ising model for in the talk: you get a great deal from two states.

Razo-Mejia and colleagues put the MWC model inside the thermodynamic model 12. The repressor’s probability of being active becomes a function of inducer concentration, with parameters for the ligand’s affinity for each conformation and for the free energy difference between them. That probability multiplies RR in the fold-change formula, so the induction curve of any strain follows from the binding energies and copy numbers already measured. They tested the predictions across strains with different repressor copy numbers and operator strengths and found that the model captured the induction profiles, including the dynamic range and the EC50 (the inducer concentration at which the response is halfway between its minimum and maximum). A single expression for the free energy of the repressor collapsed all of the data onto one master curve: many curves are one curve in the right coordinate.

Chure and colleagues then extended the framework to mutations 13. A point mutation in the repressor’s DNA-binding domain changed only the DNA-binding energy, and a mutation in the inducer-binding domain changed only the allosteric parameters. The induction profiles of double mutants could be predicted from the single mutants.

Napoleon is in equilibrium

The obvious objection is that a cell is not at equilibrium, since it consumes energy continuously. Phillips addressed the objection in a 2015 review with the provocative title “Napoleon Is in Equilibrium” 14. The argument is about timescales. Equilibrium statistical mechanics applies to a subprocess whenever that subprocess relaxes much faster than the processes driving the system as a whole. Transcription factor binding is fast compared with transcription, and conformational switching of a receptor is fast compared with the cascade downstream of it. So the occupancy of a promoter or the fraction of active receptors can be computed with Boltzmann weights even though the cell itself is far from equilibrium. He works through examples in cell signaling and in the organization of proteins and membranes. He is also explicit that the assumption fails when the driven process and the modeled one share a timescale.

The question to ask of any biological model is therefore whether the variable being computed equilibrates quickly relative to the variables being held fixed. If it does, the partition function is available and the parameter count stays small. If it does not, you need kinetics, and the parameter count grows.

What this would look like applied to your own data

Everything above was done in E. coli with strains built to turn one knob at a time. A person supplies no such knobs, so it is worth being precise about which parts transfer to an individual’s molecular data.

The estimation step transfers completely, and it is where we would start. Before interpreting an RNA sequencing result, get the orders of magnitude. A typical human cell holds a few hundred thousand mRNA molecules, so a transcript at one count per million corresponds to a fraction of a molecule per cell. The copy number of a given transcription factor sets whether the RR in the fold-change equation is 10 or 10,000. BioNumbers holds many of these figures for human cells 5. The purpose of the estimate is Phillips’s purpose: to know whether a measured value should surprise you.

The fold-change equation transfers as a language rather than as a fitted model. The observable in the lac experiments is expression relative to a reference state, and the natural observable in a person’s data is the same thing: the level of a transcript today relative to a baseline sample from the same person. What does not transfer is the ability to vary RR and Δε\Delta\varepsilon independently. In an individual the regulator copy numbers and binding energies are whatever they are, and a change in fold-change between two timepoints could come from either. The model does predict that a change in a transcription factor’s abundance should shift all of its targets together according to their binding energies. That is a pattern that can be checked in a person’s own data, and it is a pattern only.

The glucose response is where the Napoleon argument earns its place. A continuous glucose monitor records a curve over hours after a meal, and that curve is a driven kinetic process. Phillips’s criterion says an equilibrium model applies only to the fast subprocesses inside it, such as insulin binding its receptor. An MWC-style dose-response with an EC50 and a dynamic range is a reasonable descriptor of those. The curve as a whole needs a model with rates. The parameters you would extract from it are kinetic: how fast glucose rises, how quickly it returns to baseline, and how those numbers drift across months. Phillips’s approach would insist that they be estimated before they are fit and compared against a prior expectation, so that a change in them is something you can be surprised by.

What is established and what is a program

The established part is narrow and solid. For simple repression in E. coli, a thermodynamic model with a handful of parameters predicts fold-change across four orders of magnitude. Its binding energies transfer unchanged to predict induction curves, and its allosteric parameters transfer to predict the effect of mutations 9 12 13. The estimation habits in Cell Biology by the Numbers are a method anyone can apply today.

The program is everything beyond that. The Phillips group’s review is candid that simple repression is a template and that most regulatory arrangements in nature are more complex 10. Eukaryotic transcription involves chromatin and enhancers acting at a distance, and thermodynamic models there are a research topic rather than a settled tool. Applying the reasoning to one person’s longitudinal data is further out still, because the controlled knobs are absent and the equilibrium assumptions have to be argued case by case. Phillips repeats the critique from inside biology in the talk: that physical biology may amount to dotting i’s and crossing t’s in areas already understood. His answer is that “you never exactly know where the surprise is going to come from”. We think that is the right attitude to take toward one’s own data as well.

Questions people also ask

What is the physical biology of the cell? An approach that begins with order-of-magnitude estimates of sizes and rates and copy numbers, then uses statistical mechanics to build models with few parameters that can be tested against quantitative experiments.

What is the thermodynamic model of gene regulation? A model in which gene expression is proportional to the equilibrium probability that RNA polymerase occupies the promoter, computed from Boltzmann weights over the arrangements of regulatory proteins on the DNA 7 8.

Why can equilibrium physics be used on a living cell? Because a fast subprocess such as protein binding relaxes much more quickly than the slow processes that drive the cell, and on that separation of timescales its statistics are those of equilibrium 14.

Woolf Software builds longitudinal molecular profiles of individuals: whole-genome sequencing, RNA sequencing, proteomics, blood biomarkers, and continuous glucose data, integrated into one model of you. Build your profile.

Footnotes

  1. William B. Whitman, David C. Coleman, William J. Wiebe. Prokaryotes: The unseen majority. Proceedings of the National Academy of Sciences, 1998. https://doi.org/10.1073/pnas.95.12.6578

  2. Johannes Müller, Stefan Oehler, Benno Müller-Hill. Repression of lac Promoter as a Function of Distance, Phase and Quality of an Auxiliary lac Operator. Journal of Molecular Biology, 1996. https://doi.org/10.1006/jmbi.1996.0143

  3. Douglas E. Smith, Sander J. Tans, Steven B. Smith, et al. The bacteriophage φ29 portal motor can package DNA against a large internal force. Nature, 2001. https://doi.org/10.1038/35099581

  4. Paul Grayson, Lin Han, Tabita Winther, Rob Phillips. Real-time observations of single bacteriophage λ DNA ejections in vitro. Proceedings of the National Academy of Sciences, 2007. https://doi.org/10.1073/pnas.0703274104

  5. Ron Milo, Paul Jorgensen, Uri Moran, Griffin Weber, Michael Springer. BioNumbers—the database of key numbers in molecular and cell biology. Nucleic Acids Research, 2010. https://doi.org/10.1093/nar/gkp889 2

  6. Nitzan Rosenfeld, Jonathan W. Young, Uri Alon, Peter S. Swain, Michael B. Elowitz. Gene Regulation at the Single-Cell Level. Science, 2005. https://doi.org/10.1126/science.1106914

  7. G. K. Ackers, A. D. Johnson, M. A. Shea. Quantitative model for gene regulation by lambda phage repressor. Proceedings of the National Academy of Sciences, 1982. https://doi.org/10.1073/pnas.79.4.1129 2

  8. Lacramioara Bintu, Nicolas E. Buchler, Hernan G. Garcia, et al. Transcriptional regulation by the numbers: models. Current Opinion in Genetics & Development, 2005. https://doi.org/10.1016/j.gde.2005.02.007 2

  9. Hernan G. Garcia, Rob Phillips. Quantitative dissection of the simple repression input–output function. Proceedings of the National Academy of Sciences, 2011. https://doi.org/10.1073/pnas.1015616108 2

  10. Rob Phillips, Nathan M. Belliveau, Griffin Chure, et al. Figure 1 Theory Meets Figure 2 Experiments in the Study of Gene Expression. Annual Review of Biophysics, 2019. https://doi.org/10.1146/annurev-biophys-052118-115525 2

  11. Jacque Monod, Jeffries Wyman, Jean-Pierre Changeux. On the nature of allosteric transitions: A plausible model. Journal of Molecular Biology, 1965. https://doi.org/10.1016/S0022-2836(65)80285-6

  12. Manuel Razo-Mejia, Stephanie L. Barnes, Nathan M. Belliveau, et al. Tuning Transcriptional Regulation through Signaling: A Predictive Theory of Allosteric Induction. Cell Systems, 2018. https://doi.org/10.1016/j.cels.2018.02.004 2

  13. Griffin Chure, Manuel Razo-Mejia, Nathan M. Belliveau, et al. Predictive shifts in free energy couple mutations to their phenotypic consequences. Proceedings of the National Academy of Sciences, 2019. https://doi.org/10.1073/pnas.1907869116 2

  14. Rob Phillips. Napoleon Is in Equilibrium. Annual Review of Condensed Matter Physics, 2015. https://doi.org/10.1146/annurev-conmatphys-031214-014558 2