← All recipes Every night
Overnight autoresearch
Go to bed with a benchmark. Wake up to a faster app and the lab notes to prove it.
An agent runs experiments against your benchmark all night, keeping only the changes that measurably win and pass every check.
How it runs
-
It starts
02:00, every night, a task starts on its own.
-
The agent works
Profiles, forms one hypothesis, makes the smallest change that tests it, and reruns the benchmark and your checks. Built-in CI runs in the same sandbox, so a red check is fixed or reverted in the same run.
-
You approve
Read the chart and the hypothesis table over coffee, then merge the wins.
The prompt
Paste it into the automation’s instructions. Each run starts a task from it.
Make our app faster under load. Measure with <your benchmark command>. Run it three times and take the median as the baseline before you change anything.
Then loop: profile to find where the time actually goes (don't guess), form one hypothesis, and make the smallest change that tests it. Rerun the benchmark and the project's checks. Keep the change only if the median improves by at least 2% and every check passes; otherwise revert it. One idea per commit.
Stop when you've beaten the baseline by 10%, or after 20 attempts, whichever comes first. Never weaken a test, a check or the benchmark itself to win.
Attach results.html: a chart of the benchmark across every attempt, and a table of each hypothesis, its measured effect and whether it was kept.
Inspired by Andrej Karpathy’s autoresearch loop, and the 53% speed-up Shopify’s Liquid team got from it.
Set it up
In your project, open Settings → Automations → New automation.
- Trigger
- Schedule: daily, 02:00
- Action
- Create task
- Start each task
- One-shot implementation