Biology
Cell and molecular biology, genomics and bioinformatics, structural biology, neuroscience, microbiology, and ecology.
Tools you might use
- Fiji / ImageJ
- CellProfiler
- Biopython
- BLAST
- Snakemake
- R / Bioconductor
- AlphaFold
- PyMOL
- GROMACS
Open call · Biology & materials science
We're AI researchers who teach language-model agents to use real scientific software. Now we want to learn from the people doing the science: how you work, where AI already helps, and where it falls short. If you're a graduate student in biology or materials science, come shape what we build and write the paper with us.
Our team includes researchers from
and collaborators at other universities
Why we need you
We built a test from 100 tasks based on real scientific software and ran eight frontier models on it. The best one solved one task in five. Life-science tasks were solved once in 120 attempts.
Most training data for AI agents comes from software engineering, so the workflows scientists actually run are barely covered. Closing that gap takes people who know what a correct result looks like in their field.
Who we're looking for
Master's or PhD, wet lab or computational, any level of AI experience. If you script, simulate, image, or analyze data as part of your research, you have exactly the knowledge we're missing.
Cell and molecular biology, genomics and bioinformatics, structural biology, neuroscience, microbiology, and ecology.
Tools you might use
Computational materials, characterization, polymers and soft matter, energy materials, and metals and ceramics.
Tools you might use
Working in a neighboring field, like chemistry, biophysics, or chemical engineering? We'd still love to hear from you.
Let's talk about your research
We want to understand research from the inside: what fills your week, which steps AI already speeds up, and which ones it gets wrong. Here is our starting picture of a typical project. Help us correct it.
Pick a stage of a typical project.
No preparation needed. These are the kinds of things we'd talk about in a first conversation.
Demos from our previous work
Each demo is a task from our research. An agent sees a few real runs of a scientific tool and writes a program that reproduces it. We then run that program on settings it has never seen and check every output against the real software.
Life & Molecular Sciences
The same steps as counting nuclei or measuring cell shape in your own fluorescence images.
Physical & Engineering Simulation
Meshing, boundary conditions, and solver settings: the setup behind any flow or transport simulation.
Spatial, Visual & 3D
Parametric CAD for 3D-printed sample holders, fixtures, and flow cells.
Earth & Space Sciences
Aperture photometry and background subtraction, the same math as quantifying spots or bands in noisy images.
Spatial, Visual & 3D
The color-space math behind calibrated imaging, colorimetric assays, and optical characterization.
Data & Document Workflows
Reading numbers out of scanned tables, the first step in mining property data from papers.
Physical & Engineering Simulation
Physical models with masses, inertias, and joint limits, as used to simulate lab robots before running them.
Spatial, Visual & 3D
A physics simulation whose outcome shifts with wind, constraints, and resolution.
Spatial, Visual & 3D
Design-rule checks on circuit boards, like the custom electronics inside lab instruments.
Software & Formal Systems
Recording exactly which software went into an environment, the backbone of reproducible analysis.
Plenty of research time goes to installing tools, managing data, and fixing broken environments. Our agents practice that as well. Each replay shows commands from the task's own reference solution.
HPC & clusters
Install the Flux Python package from the tarball at /app/flux-python-0.48.0rc6.tar.gz after verifying its integrity against the provided checksum file.
Research data management
Set up a DVC-tracked dataset under /app/dataset/ with two items: dataset/foo and dataset/test/0.
Python environments
In /app, diagnose and fix the environs URL parsing bug.
How we'd work together
Start small and go as far as you like. Every step helps, and you decide how involved you want to be.
An informal conversation about your research, your tools, and where AI helps or gets in the way.
Together we choose a real piece of your work: an analysis script, a simulation setup, or a data pipeline.
We package it as a task an agent can attempt, with checks you agree are scientifically right.
We measure today's agents, train better ones, and write up what we learn together.
Tell us how you work and where AI lets you down.
Help turn a piece of your research into a task with real checks.
Join the analysis and the writing of the paper.
Our research so far
The demos above come from this work. Each paper tackles one part of the problem: building tasks that check results reliably, making tasks harder over time, and turning an agent's successes into better training data.
Builds agent tasks directly from existing scientific software. The agent sees a few real runs and must reproduce the workflow, and hidden runs check the result. The tasks span six domains, including life and molecular sciences, physical and engineering simulation, and earth and space sciences.
Why it matters to you: this is how we'd turn your workflows into tasks.
Fine-tuning on these tasks, Terminal-Bench 2
Long-horizon terminal tasks
Automated checks agree with blinded human reviewers
Grows simple, verified tasks into long, multi-step ones, round after round. Every new task is checked in a fresh sandbox before it is kept, at roughly $0.05 per task.
Why it matters to you: agents learn to finish whole projects, not just single commands.
Same strong model, tasks from round 1 vs. round 15
Median reference solution length
Reinforcement learning on these tasks, Terminal-Bench Hard
Collects an agent's successful solutions from several different setups, then rewrites them into one consistent style the model can learn from. The rewritten data teaches the model far more than the raw successes do.
Why it matters to you: one good solution to your task can become many training examples.
Terminal-Bench 2, solved within three tries
Terminal-Bench Hard, solved within three tries
Terminal-Bench 4, solved within three tries
Questions
Anything else? Ask us directly. We're happy to talk it through.
AI researchers from the University of Maryland, the University of Georgia, the National University of Singapore, Indiana University, The Hong Kong Polytechnic University, and other universities. We build and train AI agents that work with real software, and we publish our results on arXiv.
No. We bring the AI side. What we need is your knowledge of how the science is done and how you know a result is right.
It's up to you. A single conversation already helps us. Contributing a workflow takes more time, and we'll plan it around your schedule.
No. Tasks can be built from open-source software with public or synthetic inputs. We'll agree together on anything that goes into a paper.
Yes. Collaborators who contribute workflows, tasks, or analysis can be co-authors on the resulting paper, following standard authorship norms.
Graduate students, Master's or PhD, in biology, materials science, and related fields. Not sure you fit? Write to us anyway.
Get involved
A few lines is plenty, and we'll reply to set up a conversation. Sending opens your email app with your answers filled in, so nothing is stored on this site.
Prefer plain email? Write to zli12321@terpmail.umd.edu and zl22754@uga.edu.
The agent sees public runs of the real software and writes a program that reproduces it. That program is then run on hidden settings of these parameters and must match the software's outputs, field by field.
Work with tools like these? Let's talksolve.sh