Making Science Programmable.
Dynamical Systems is a research lab building AI that learns how to discover.
We make every experiment verifiable, replayable, and compounding.
Research
Dynamical-SDL-1: Measuring How Scientific Agents Learn from Physical Experiments
Dynamical-SDL-1 compiles a complete physical reaction map into controlled long-horizon campaigns that measure how efficiently scientific agents convert experimental evidence into better decisions. Across six systems and six campaign assignments, local evidence assimilation, held-out transfer, and adaptive experiment selection emerge as distinct capabilities.
Simulator-Verified Skill Acquisition for Scientific Instruments
Proprio gives an agent a persistent simulator loop to draft, execute, inspect, and repair an instrument operating skill, while an independent verifier it cannot change decides what enters the catalog. Verified feedback produced 14 non-regressive repairs from 18 paired drafts where blind retrying produced none, and one frozen protocol acquired, repaired, verified, and evolved skills across three external instrument families with zero invalid promotions. Verified in simulation. Hardware validation remains separate.
Can a Self-Driving-Lab Agent Tell When the Evidence Is Enough?
We turn the historical record of a lab into source-located replay tasks that measure evidence-boundary judgment before a self-driving lab is trusted to run on its own. Across six frontier models and 1,872 trajectories, agents reach a valid decision on 90% of runs and the reference-equivalent path on 72%, and no model clears the benchmark.
Scaling Test-Time Verification for Novel Materials
Crystal diffusion models encode property signals in their hidden states that they never use during sampling. Probe-gradient guidance steers unconditional generation toward target properties at test time, comparable to conditional models at over 50x the speed.
The Missing Layer in Autonomous Science
Multi-turn RL in verified campaign environments lifts hypothesis accuracy from 55.2% to 79.3% on held-out campaigns, surpassing GPT-5.4 with 3B active parameters.