Open call · Biology & materials science

Let's build AI that actually helps your research.

We're AI researchers who teach language-model agents to use real scientific software. Now we want to learn from the people doing the science: how you work, where AI already helps, and where it falls short. If you're a graduate student in biology or materials science, come shape what we build and write the paper with us.

  • No ML background needed
  • Your workflows, your expertise
  • Contributors can join the paper

Our team includes researchers from

  • University of Maryland
  • University of Georgia
  • National University of Singapore
  • Indiana University
  • The Hong Kong Polytechnic University

and collaborators at other universities

Why we need you

Frontier AI agents still struggle with real science.

We built a test from 100 tasks based on real scientific software and ran eight frontier models on it. The best one solved one task in five. Life-science tasks were solved once in 120 attempts.

Most training data for AI agents comes from software engineering, so the workflows scientists actually run are barely covered. Closing that gap takes people who know what a correct result looks like in their field.

20%Success rate of the best of eight frontier models on our held-out scientific-software test
1 in 120Attempts that succeeded on life & molecular science tasks, across all eight models
98%Agreement between our automated checks and blinded human reviewers

Who we're looking for

Graduate students in biology and materials science.

Master's or PhD, wet lab or computational, any level of AI experience. If you script, simulate, image, or analyze data as part of your research, you have exactly the knowledge we're missing.

Biology

Cell and molecular biology, genomics and bioinformatics, structural biology, neuroscience, microbiology, and ecology.

Tools you might use

  • Fiji / ImageJ
  • CellProfiler
  • Biopython
  • BLAST
  • Snakemake
  • R / Bioconductor
  • AlphaFold
  • PyMOL
  • GROMACS

Materials science

Computational materials, characterization, polymers and soft matter, energy materials, and metals and ceramics.

Tools you might use

  • VASP
  • Quantum ESPRESSO
  • LAMMPS
  • ASE
  • pymatgen
  • Materials Project
  • GSAS-II
  • OVITO
  • COMSOL

Working in a neighboring field, like chemistry, biophysics, or chemical engineering? We'd still love to hear from you.

Let's talk about your research

Where does AI help you today, and where does it fall short?

We want to understand research from the inside: what fills your week, which steps AI already speeds up, and which ones it gets wrong. Here is our starting picture of a typical project. Help us correct it.

Pick a stage of a typical project.

Biology

Reading & planning

What it looks like
Keeping up with papers and preprints, comparing protocols, and planning the next experiment.
Where AI falls short
Confident summaries with missing or invented citations, and little sense of which protocol details actually matter.
What we'd build together
Tasks that check an agent's answers against the papers and protocols you rely on.

Designing experiments

What it looks like
Designing primers, guide RNAs, constructs, and controls, often across several separate tools.
Where AI falls short
Chaining specialized tools together and respecting organism- and lab-specific constraints.
What we'd build together
Design tasks scored by the same tools and rules you use to sanity-check a design.

Processing data

What it looks like
Sequencing pipelines, microscopy segmentation, and flow-cytometry or plate-reader data.
Where AI falls short
Installing tools, wiring up multi-step pipelines, and catching silent errors in the output.
What we'd build together
Pipelines rebuilt from your real runs, with checks on the numbers, not just on whether files exist.

Analysis & modeling

What it looks like
Statistics, structure prediction, docking, and molecular dynamics.
Where AI falls short
Choosing sensible methods and parameters, and noticing when a result is biologically implausible.
What we'd build together
Analyses checked on hidden test cases, so an agent has to get the method right, not just the format.

Figures & writing

What it looks like
Figures, methods sections, and notebooks that others can rerun.
Where AI falls short
Methods text that drifts from what the code actually did.
What we'd build together
Tasks where every figure and methods detail has to match the analysis that produced it.

Materials science

Reading & planning

What it looks like
Finding synthesis routes, processing conditions, and property data scattered across many papers.
Where AI falls short
Pulling values out of tables and figures without mixing up units, samples, or conditions.
What we'd build together
Extraction tasks checked against values curated by people who know the field.

Setting up simulations

What it looks like
Building DFT or molecular-dynamics inputs, running convergence tests, and writing cluster job scripts.
Where AI falls short
Physically sensible parameters, pseudopotential and force-field choices, and the quirks of each cluster.
What we'd build together
Simulation tasks checked by rerunning the real codes on inputs the agent never saw.

Characterization

What it looks like
Processing XRD patterns, electron-microscopy images, and spectra.
Where AI falls short
Peak fits and refinements that look reasonable but aren't physically right.
What we'd build together
Characterization pipelines rebuilt from real instrument data, with checks that respect measurement tolerances.

Screening & data

What it looks like
Querying materials databases, featurizing structures, and screening candidates.
Where AI falls short
Keeping structures, compositions, and units consistent from one tool to the next.
What we'd build together
Screening workflows checked against reference results from the codes you trust.

Figures & writing

What it looks like
Phase diagrams, band structures, and methods that others can reproduce.
Where AI falls short
Plots and text that drift from the underlying calculations.
What we'd build together
Tasks where every figure traces back to a verified calculation.

Questions we'd love to ask you

No preparation needed. These are the kinds of things we'd talk about in a first conversation.

  1. Which part of your week feels most like busywork?
  2. Where do you already use AI, and where have you stopped trusting it?
  3. What software, instruments, or pipelines does your work depend on?
  4. How do you check that a result is right before you believe it?
  5. Which task would you hand to a reliable assistant tomorrow?
  6. What would an AI agent need to show before you'd publish its result?

Demos from our previous work

Agents working with real scientific software.

Each demo is a task from our research. An agent sees a few real runs of a scientific tool and writes a program that reproduces it. We then run that program on settings it has never seen and check every output against the real software.

Microscopy label masks from public runs

Life & Molecular Sciences

Microscopy cell segmentation and counting

The same steps as counting nuclei or measuring cell shape in your own fluorescence images.

  • Biology
  • Microscopy
  • CellProfiler
  • scikit-image
  • Bio-Formats
OpenFOAM dam-break mesh drawn from the public polyMesh

Physical & Engineering Simulation

OpenFOAM 2-D dam-break CFD

Meshing, boundary conditions, and solver settings: the setup behind any flow or transport simulation.

  • Simulation
  • Fluid dynamics
  • OpenFOAM
Rendered fan-duct mesh from a public OpenSCAD run

Spatial, Visual & 3D

OpenSCAD / BOSL2 parametric fan duct

Parametric CAD for 3D-printed sample holders, fixtures, and flow cells.

  • Lab hardware
  • CAD
  • OpenSCAD
  • BOSL2
Astronomy FITS cutouts from public runs

Earth & Space Sciences

Astronomy FITS/WCS cutout and photometry

Aperture photometry and background subtraction, the same math as quantifying spots or bands in noisy images.

  • Imaging
  • Signal extraction
  • Astropy
  • Photutils
Color-pipeline previews from public runs

Spatial, Visual & 3D

Color-space and tone-curve pipeline

The color-space math behind calibrated imaging, colorimetric assays, and optical characterization.

  • Imaging
  • Optics
  • colour-science
  • ImageMagick
  • scikit-image
OCR page preview from a public run

Data & Document Workflows

Tesseract table-image OCR

Reading numbers out of scanned tables, the first step in mining property data from papers.

  • Literature
  • Data extraction
  • Pandoc
  • Poppler
  • Tesseract
  • QPDF
MuJoCo joint-position trajectories across public scenarios

Physical & Engineering Simulation

MuJoCo / URDF mass-inertia repair

Physical models with masses, inertias, and joint limits, as used to simulate lab robots before running them.

  • Lab automation
  • Physics
  • MuJoCo
  • URDF/MJCF tooling
Blender cloth scene rotation dial across public scenarios

Spatial, Visual & 3D

Blender cloth flag scene

A physics simulation whose outcome shifts with wind, constraints, and resolution.

  • Simulation
  • Physics
  • Blender
KiCad board with design-rule findings plotted at their positions

Spatial, Visual & 3D

KiCad LED-matrix driver PCB

Design-rule checks on circuit boards, like the custom electronics inside lab instruments.

  • Instrumentation
  • Electronics
  • KiCad
Container image layers and attestation from a public run

Software & Formal Systems

Container SBOM and supply-chain attestation

Recording exactly which software went into an environment, the backbone of reproducible analysis.

  • Reproducibility
  • Software
  • Hadolint
  • Syft
  • Skopeo

Everyday research computing, too

Plenty of research time goes to installing tools, managing data, and fixing broken environments. Our agents practice that as well. Each replay shows commands from the task's own reference solution.

HPC & clusters

Install an HPC scheduler's Python bindings, verified

Install the Flux Python package from the tarball at /app/flux-python-0.48.0rc6.tar.gz after verifying its integrity against the provided checksum file.

  • 171-line reference solution
  • Checked in a fresh sandbox

Research data management

Version a dataset with DVC and reconcile its history

Set up a DVC-tracked dataset under /app/dataset/ with two items: dataset/foo and dataset/test/0.

  • 313-line reference solution
  • Checked in a fresh sandbox

Python environments

Diagnose and fix a Python dependency bug

In /app, diagnose and fix the environs URL parsing bug.

  • 202-line reference solution
  • Checked in a fresh sandbox

How we'd work together

From one conversation to a shared paper.

Start small and go as far as you like. Every step helps, and you decide how involved you want to be.

  1. 1

    Talk

    An informal conversation about your research, your tools, and where AI helps or gets in the way.

  2. 2

    Pick a workflow

    Together we choose a real piece of your work: an analysis script, a simulation setup, or a data pipeline.

  3. 3

    Turn it into a test

    We package it as a task an agent can attempt, with checks you agree are scientifically right.

  4. 4

    Improve and publish

    We measure today's agents, train better ones, and write up what we learn together.

What you'd bring

  • How research in your field actually gets done
  • Real workflows: scripts, simulations, and pipelines
  • Judgment about what counts as a correct result
  • Feedback on what agents get right and wrong

What you'd get

  • Co-authorship on the resulting paper for substantial contributions
  • A first look at how today's AI agents handle your own workflows
  • A packaged, tested version of the workflow you contributed
  • Hands-on experience with how AI agents are trained and evaluated

Ways to take part

One conversation

Tell us how you work and where AI lets you down.

Contribute a workflow

Help turn a piece of your research into a task with real checks.

Co-author

Join the analysis and the writing of the paper.

Our research so far

Results from our recent papers.

The demos above come from this work. Each paper tackles one part of the problem: building tasks that check results reliably, making tasks harder over time, and turning an agent's successes into better training data.

Scientific softwarearXiv · October 2026

Self-Supervised Scaling of Terminal Environments for Scientific Domains

Zhongzhi Li, Yucheng Shi, Zongxia Li, et al.

Builds agent tasks directly from existing scientific software. The agent sees a few real runs and must reproduce the workflow, and hidden runs check the result. The tasks span six domains, including life and molecular sciences, physical and engineering simulation, and earth and space sciences.

Why it matters to you: this is how we'd turn your workflows into tasks.

Key results

Fine-tuning on these tasks, Terminal-Bench 2

Before 47.9%
After 53.6%

Long-horizon terminal tasks

Before 20.7%
After 27.7%

Automated checks agree with blinded human reviewers

98%

Harder tasksarXiv · August 2026

Recursive Synthesis for Long-Horizon Terminal Tasks

Zhongzhi Li, Yucheng Shi, Zongxia Li, et al.

Grows simple, verified tasks into long, multi-step ones, round after round. Every new task is checked in a fresh sandbox before it is kept, at roughly $0.05 per task.

Why it matters to you: agents learn to finish whole projects, not just single commands.

Key results

Same strong model, tasks from round 1 vs. round 15

Round 1 solved 90%
Round 15 solved 2.5%

Median reference solution length

67 lines
374 lines

Reinforcement learning on these tasks, Terminal-Bench Hard

Before 22.7%
After 32.0% (+41% relative)

Better training dataarXiv · October 2026

Scaling Trajectories for Complex Tasks through Recursive Self-Rewrite

Zongxia Li, Yucheng Shi, Zhongzhi Li, et al.

Collects an agent's successful solutions from several different setups, then rewrites them into one consistent style the model can learn from. The rewritten data teaches the model far more than the raw successes do.

Why it matters to you: one good solution to your task can become many training examples.

Key results

Terminal-Bench 2, solved within three tries

Before 57.0%
After 74.2%

Terminal-Bench Hard, solved within three tries

Before 39.0%
After 63.0%

Terminal-Bench 4, solved within three tries

Before 1.5%
After 9.1%

Questions

Before you write to us.

Anything else? Ask us directly. We're happy to talk it through.

Who are you?

AI researchers from the University of Maryland, the University of Georgia, the National University of Singapore, Indiana University, The Hong Kong Polytechnic University, and other universities. We build and train AI agents that work with real software, and we publish our results on arXiv.

Do I need machine-learning experience?

No. We bring the AI side. What we need is your knowledge of how the science is done and how you know a result is right.

How much time does it take?

It's up to you. A single conversation already helps us. Contributing a workflow takes more time, and we'll plan it around your schedule.

Do I have to share unpublished data?

No. Tasks can be built from open-source software with public or synthetic inputs. We'll agree together on anything that goes into a paper.

Can I be an author on the paper?

Yes. Collaborators who contribute workflows, tasks, or analysis can be co-authors on the resulting paper, following standard authorship norms.

Who can join?

Graduate students, Master's or PhD, in biology, materials science, and related fields. Not sure you fit? Write to us anyway.

Get involved

Tell us about your research.

A few lines is plenty, and we'll reply to set up a conversation. Sending opens your email app with your answers filled in, so nothing is stored on this site.

Prefer plain email? Write to zli12321@terpmail.umd.edu and zl22754@uga.edu.

Field
How would you like to take part?