Periplasmic single-domain antibody
E. coli BL21(DE3)
Loading…
Research
Research preview · August 2026
E. coli BL21(DE3)
E. coli DH5α
E. coli DH5α
E. coli BL21(DE3)
E. coli BL21(DE3)
E. coli DH5α
HEK293T
Human cell line
S. cerevisiae
Summary: This post sets out two evaluations of the GeneLoop agent, and it reports no results. The first asks whether the agent can carry a construct-design brief from an instruction to a persisted, independently reviewable artifact. The second asks whether its verifier agrees with an independent analysis of the same molecule — in both directions, on clean constructs as well as broken ones. We are publishing the briefs, the measurement definitions, and the protocol before running anything, so that the numbers which eventually fill these panels can be read against a design nobody adjusted after seeing them.
Designing a plasmid that expresses a protein is not, for most people who do it, a hard intellectual problem. It is an accounting problem carried out across a dozen tools: find the right sequence, confirm it is the right one, pick a backbone whose promoter and selection marker suit the host, choose enzymes that cut where you need and nowhere else, keep the reading frame from the ribosome binding site through the junction, and check afterwards that the file you built is the construct you described. Each step is tractable. Doing all of them, in order, without an error that only surfaces three weeks later at the bench, is what takes the time.
GeneLoop is built for the part of that work where a plausible answer is worth nothing. The agent grounds a brief in sourced parts and the host's actual capabilities, writes a plan the scientist reads and approves before anything is written, composes the construct by reference to explicit workspace paths, and then verifies the persisted file rather than its own account of what it did. Raw bases never enter its context at any point; it reads the molecule through analyses it authors and the server runs.
The first campaign tests that loop end to end on the nine briefs above. It measures the artifact — whether a construct exists, whether it is bound to a revision and a hash, whether an independent reading of the same file agrees with the agent's conclusions. It measures nothing about whether the construct works at the bench, because nothing has been to the bench.
Briefs attempted
Sessions that ran to a terminal state without operator intervention. A session that stalls counts as attempted and failed, not as excluded.
Persisted and verified
Constructs written to a workspace path, bound to a revision and sequence hash, and returned ready or partial by the verifier.
Independent concordance
Briefs where a blind BioPython analysis of the same persisted revision reached the agent's conclusion.
Three of the nine briefs are clean: a competent scientist would produce a working construct, and so should the agent. The interesting four carry a seeded defect, and the last two are holdouts written after the agent's prompt and tool surface were frozen. What follows is what each group is for.
Functional requirements. Asked to express a protein that needs something the host cannot supply — a chromophore, a disulfide-permissive compartment, a glycosylation the cytoplasm will not perform — an agent can produce a construct that is correct in every nucleotide and still cannot produce the phenotype requested. The failure we are testing for is not a wrong sequence. It is a real requirement written into a risks footnote instead of into the design. These briefs each carry one unmet requirement, and we record whether the agent redesigns to satisfy it, asks the scientist whether the apo product is acceptable, or ships the construct with a caveat attached.
Planned readout. A caveat is not a design decision; only a redesign or a recorded user choice counts.
Junction integrity. Cloning boundaries are where a construct that looks right stops working. Keeping one byte too few at a religated site shifts the reading frame from the upstream ribosome binding site, and the annotated coding sequence downstream is never translated — the feature map still renders correctly, the length is plausible, and nothing about the file announces the defect. We seed a boundary error and measure whether the agent reads the actual backbone rather than a remembered convention, anchors on the bytes it read, and proves the repair on the persisted file rather than on its own description of the edit.
Planned readout. Detection and repair are scored separately — an agent that notices the problem and cannot fix it has not built the construct.
Evidence over labels. An annotation is a claim someone made about a molecule, not a property of it. A feature named for a resistance marker may span a promoter and a ribosome binding site as well as the coding sequence, or may name the wrong gene entirely, and a duplicated element hides completely if one of its two copies carries a different label. This brief carries deliberately misleading annotations over honest bytes. What we are measuring is whether the agent's conclusions come from translating the sequence or from reading the label above it, and whether it names what the bytes actually are when the two disagree.
Planned readout. Reporting a construct as sound because its labels look right is a failure even when the construct is sound.
Recovery after failure. Most build failures in practice are recoverable: an enzyme that does not cut uniquely on this backbone, a boundary anchored to a site that appears twice, a source that resolves to the wrong revision. The useful behaviour is to re-ground on the file, take the remedy the error named, and make one corrected write — or, when the method itself is the blocker, change method. The failure mode is a truthful report that the build did not happen, delivered instead of the build. We inject one typed failure per session and record what the agent does next.
Planned readout. Re-issuing the same failing call, or stopping to ask the scientist a question the error already answered, counts as no recovery.
A passing agent run is not a biological result, and the report will not present it as one. The campaign records failures, partial verifications and unresolved claims as themselves. An evaluation that converts its own uncertainty into a success rate is measuring its authors rather than the system.
The second evaluation concerns reading rather than building. Given a persisted construct, GeneLoop's verifier translates coding regions, reconciles each annotation against what the bytes actually encode, checks copy number, tests whether a promoter is positioned to drive the transgene it is supposed to drive, and reports what it found. The question is whether that reading agrees with an independent one of the same file.
Agreement has to be measured in both directions, which is why the fixture set contains clean controls as well as known defects. A verifier that catches every seeded fault and also rejects correct constructs has not solved the problem; it has moved it. In practice the second failure is the more expensive one, because a scientist whose correct design is refused has no way to proceed and no way to tell a real objection from a detector that simply does not cover their construct.
Verifier agreement
Diagnostics matching an independent reading of the same revision, counted per claim rather than per construct.
False blocks
Correct constructs stopped by a detector that could not reach a verdict but reported one anyway.
False passes
Seeded defects that neither the agent's own read nor the verifier surfaced.
Most of the false blocks we have seen came from the same shape of mistake: a two-valued output over a three-valued reality. A detector that can only say pass or fail has to say fail when it cannot tell, and “I could not establish this” then reaches the scientist as “this is wrong.” Every finding therefore carries an evidence class, and its severity is derived from that class rather than set by hand.
The bytes positively disprove the claim — a premature stop inside the coding sequence, a frame that does not close, an element present twice where the design allows one.
Blocks. The construct is not deliverable as designed.
The detector could not reach a verdict. A promoter recogniser that knows only curated motifs has nothing to say about a synthetic response element, and its silence is not evidence of absence.
Ships with the uncertainty stated. The construct is never described as verified.
The finding is equally true of the molecule the build started from, so it describes the backbone rather than the work.
Recorded, never blocking.
A construct with unproven findings is still delivered, with the uncertainty stated in the report and attached to the claim it qualifies. It is never described as verified. That distinction is the thing the second evaluation is really testing: not whether the verifier is strict, but whether it is honest about the limits of what it checked.
The protocol is published here in full because a pre-registration that can be revised after the fact is not one. The order is load-bearing: the truth records and the metric definitions are written before the agent sees a brief, and the panels above are replaced only after the record behind them is public.
GeneLoop does not show raw DNA to the chat model, and that constraint shapes what an evaluation of it can publish. A nucleotide string is not tokenised in any biologically meaningful way for a language model: it cannot reliably count within one, locate a position, or hold a reading frame, and when it tries it produces confident, specific, wrong answers. The agent works from derived facts instead — coordinates, feature names, codon and residue identities, the deterministic consequence class of an edit, hashes — all computed server-side from the real sequence.
The evaluation records preserve that boundary. What we publish for each session is the brief, the plan, the tool-call ledger, the diagnostics, the sequence hash and the independent reading, which is everything needed to check the claimed result. What we do not publish is the constructs themselves as sequence, and no design in this campaign is selected for a property that would make its publication a problem. Briefs are screened before they enter the frozen set, and a brief that does not pass screening is removed rather than run and withheld.
If these panels fill in well, what they will show is narrow and worth having: that an agent can take a construct brief to a persisted artifact that an independent reading agrees with, and can tell the difference between a molecule it has disproved and one it merely could not check. That is a claim about a workflow. It is not a claim that the construct expresses, folds, or does anything at all in a cell, and no amount of agreement between two analyses of the same file will make it one.
The part we expect to be hardest is not design. It is the long tail of constructs a fixed set of detectors was never written for, where the honest verdict is that the evidence does not reach — and where saying so, clearly, beats guessing in either direction. We will publish what happens, including the runs that do not work.
Related research
How a reviewable plan, sourced parts and an explicit approval boundary make an agent's work auditable by the scientist who has to sign for it.
Why a construct report has to separate what the bytes disprove from what the detector could not reach, and what a two-valued verifier costs.
Turning cofactors, folding environment, topology and host capability into design requirements the agent has to satisfy or escalate.