Back to blog
Bioprocess

How our AI cuts media waste by learning from every batch

by Marcus Delacroix

Abstract visualization of AI-driven bioprocess optimization cycle

Media costs are one of the larger variable expenses in a mammalian cell culture operation. For a 200 L bioreactor running bovine mammary epithelial cells, the nutrient cocktail including insulin, transferrin, growth factor fractions, and lipid supplements can run several hundred dollars per run at pilot scale. Get the feeding strategy wrong and you either starve the culture partway through or flush functional components before the cells have consumed them. Either way, the cost does not disappear.

The feedback loop we built at Opalia started from a specific frustration. Our initial runs used a conservative, literature-derived feeding schedule and consistently finished with significant residual glucose and amino acid concentrations in the spent media. We were paying for nutrients the cells never used, and protein yield was lower than it needed to be because we were not hitting peak secretion periods with the right substrate availability at the right time.

What "learning from a batch" actually means in practice

The phrase "AI optimization" gets applied to a wide range of things in bioprocess, most of them involving dashboards and alert thresholds rather than actual parameter adjustment. What we mean is narrower. At the end of each production batch, we extract HPLC purity data, protein yield measurements in mg/L, spent media metabolite analysis, and the complete time-series record of dissolved oxygen, pH, glucose concentration, and temperature deviations. These become inputs to a Bayesian optimization model that treats each subsequent batch's parameters as a sequential decision problem.

The model does not converge on a fixed recipe. It maintains a probabilistic belief about which parameter regions are likely to improve yield while staying within quality constraints, and each batch updates those beliefs. This differs from manual parameter tuning or running a fixed factorial design. We are not exhausting the parameter space in grid fashion, we are searching it efficiently with uncertainty quantification guiding where to look next.

The four variables we actually control

Four main variables change between batches: media composition (primarily glucose and glutamine concentrations and feed timing relative to cell density readings), gas flow mix (CO2 and O2 governing pH and dissolved oxygen setpoints), temperature profile (we run a temperature shift protocol toward the end of the secretion phase), and feeding schedule (bolus addition timing relative to metabolic state).

We do not adjust the base cell line conditions, bioreactor geometry, or downstream processing between batches. Those are held constant. The optimization scope is deliberately narrow: these four variables control the nutrient environment that bovine mammary epithelial cells experience during the secretion phase, and secretion rate is acutely sensitive to that environment.

What we learned early is that the relationship between glucose concentration and lactic acid accumulation is not linear. Above roughly 15 mM glucose, the cells produce lactic acid at a rate that shifts pH faster than our CO2 buffer can correct without introducing osmolality stress. Below 4 mM, secretion rates drop measurably within six hours. The feeding schedule the model has converged on holds the culture in a narrower glucose band than our initial protocol, using more frequent smaller boluses rather than the larger twice-daily additions we started with.

What the data shows after twelve batches

Across our 12 completed pilot production batches at 200 L scale, protein yield improved 31% from batch 1 to batch 12, measured in mg of beta-Casein per liter of harvest. Batch-to-batch variation, expressed as coefficient of variation across consecutive runs, sits at 2.7%, below our internal threshold of 3%.

Residual media waste has decreased substantially. Spent media glucose concentration in our early batches averaged near 9 mM at harvest, meaning roughly 40% of the glucose added was not consumed. By batch 10, that number was below 2.5 mM. Some residual is intentional as a buffer against unexpected metabolic shifts, but the gap between what we add and what the cells use has narrowed considerably.

The improvement trajectory was not linear. Batches 1 through 4 showed modest gains as the model accumulated data points. Batches 5 through 8 showed the steepest improvements, particularly in yield per liter. Batches 9 through 12 showed smaller incremental gains with tighter consistency. This pattern matches the expected behavior of a Bayesian optimization approach as model uncertainty about the optimal region decreases over successive experiments.

Where the model still underperforms

We are not claiming that the optimization loop has solved our media problem entirely. There are two areas where it reliably falls short of what we would want.

The first is cell line passage variation. Our bovine mammary cell line shows subtle phenotypic drift across high-passage numbers, and the model was not originally trained to account for this. A batch run at passage 35 behaves differently from a batch at passage 22, even under identical parameters. We have since added passage number as a covariate in the model inputs, but the interaction is still not well characterized and we manage it with conservative parameter constraints rather than active optimization.

The second is downstream interference. Our HPLC purity measurement at harvest captures total beta-Casein, but it does not capture the impact of media changes on downstream purification yield. A feeding strategy that produces higher harvest yield but leaves residual components that co-elute during chromatographic separation can result in lower final purity even if the upstream number looks better. We are working to incorporate downstream yield as a lagged objective in the optimization loop, but that data arrives two days after each batch closes, which complicates the sequential update timing.

We want to be clear on this point: the model is a production tool for a specific, bounded optimization problem. It is not a general bioprocess intelligence and we are skeptical of vendors who describe their tools that way. The value here comes from systematic data collection and principled search, not from the model having deep understanding of cell biology.

What this means for production economics at pilot scale

A 31% yield improvement on a fixed media spend means the effective cost per gram of protein has dropped materially, even at pilot scale. The absolute numbers are still not competitive with commodity dairy protein concentrate pricing, but that is not the right comparison frame for this stage.

What the optimization work provides is a tighter cost per pilot batch and the ability to quote formulation customers a consistent specification. If a food manufacturer is evaluating Opalia BC-1 for a cheese analog application, batch-to-batch protein purity variation forces reformulation between deliveries. Holding CV below 3% makes the formulation conversation considerably simpler.

The direction this pushes us toward is a production model where accumulated data allows predictive batch scheduling rather than reactive adjustment. We are not there yet after 12 batches, but the trajectory is clear and each additional run tightens the model further.