Platform
Three capabilities, deliberately built as one system.
Combinatorial chemistry makes the molecules. Pico-scale screening measures them. Machine learning decides what to make next. None of the three is useful without the other two.
Overview
Drug discovery is slow largely because the feedback loop between making a molecule and knowing whether it works is long. Our platform shortens that loop to days and then runs it continuously. Synthesis, assay and modelling are colocated in one facility in La Jolla, under one data model, so a result measured on Tuesday can change the library designed on Thursday.
The consequence is not a faster version of a conventional screen. It is a different kind of dataset: millions of measured structure–activity relationships, generated on demand, in the chemical space our programmes actually care about.
Combinatorial chemistry
Libraries are assembled from validated building blocks by parallel synthesis, with reaction conditions chosen for reproducibility at scale rather than for the single best yield. Each cycle is designed by the model against a set of constraints supplied by chemists: what will dissolve, what will be stable, what can be scaled for a later programme.
Design constraints
- Reaction classes restricted to those that reproduce reliably in plate format
- Building block sets curated for purity and availability at scale
- Medicinal chemistry filters applied before synthesis, not after
- Every compound registered with its synthetic route and QC data
Pico-scale activity-based screening
Screening runs at pico-scale in miniaturised formats, which reduces the material required per datapoint by orders of magnitude and makes it practical to measure very large libraries. Assays are activity-based: biochemical for direct target engagement, cell-based where cellular context changes the answer.
Because compounds are consumed at such small quantities, the same library can be screened against several targets and in several formats, which is what makes the resulting dataset useful for transfer learning rather than for a single programme.
Assay throughput per week
Biochemical readouts dominate volume; cell-based assays are run on the subset where context matters. Representative figures.
Machine learning
Models are trained on our own measurements. The task is not to predict activity in the abstract, but to rank the next library so that the experiments we run carry the most information per compound synthesised. We use ensembles, and we use their disagreement as the acquisition signal.
What we model
- Structure–activity relationships across targets within a family
- Selectivity against closely related off-targets
- Early ADME properties, to fail compounds before they are made
- Synthetic accessibility, to keep proposals makeable
Models are retrained on a weekly cadence as new measurements land. A model that cannot be validated against held-out wet-lab data is not used to prioritise synthesis.
The weekly loop
Monday: the current model ensemble proposes the next library against programme objectives. Tuesday to Thursday: synthesis and QC. Thursday and Friday: screening across biochemical and cell-based formats. Friday: measurements are ingested, models retrained, and the cycle begins again with a model that has seen the week's data.
Nothing about this loop is exotic individually. What is unusual is that it runs without a human in the middle of the data path, and that programme teams see the results the same week they were requested.
Data quality
A model trained on unreliable measurements learns the assay, not the chemistry. Every plate carries controls; every compound carries QC; every assay has a defined acceptance window, and results outside it are recorded as failures rather than quietly dropped.
- Intra-plate and inter-plate controls on every run
- Compound identity confirmed before screening
- Assay acceptance criteria fixed in advance and versioned
- Failed measurements retained, because the failure rate is itself data