Model

Yimin Liu

Yimin Liu

Ph.D. Candidate AI4Science

The Ohio State University · Columbus, Ohio

yiminliu.career at gmail.com

About

What I build

I build the infrastructure that decides whether LLM agents actually work — the skill systems, harnesses, benchmarks, and reward-integrity tooling that turn “the agent seems good” into a number you can defend.

An eval stack fails in three places: the runtime, the tasks, and the score itself. At BenchFlow (Research Scientist Intern, summer 2026; maintainer since January) I built for all three. Its evaluation runtime executes 100+ containerized tasks across coding, terminal, and productivity domains. SkillsBench asks whether skills actually help — 87 tasks across 8 domains, where curated skills lift pass rates by 16.6 points — and ClawsBench (COLM 2026) puts capability and safety on the same axis instead of in separate benchmarks. And BenchGuard defends the score itself: phase-aware taint analysis plus runtime evidence that catches agents gaming their own evaluation.

My Ph.D. is in AI for Science: deep learning for single-cell and spatial multi-omics in cancer — which is also where the hard, real scientific tasks in my evaluations come from.

  • Agent Skills & Benchmarks
  • Capability & Safety Evaluation
  • Reward Integrity
  • Single-Cell & Spatial Transcriptomics

Now

  • Ph.D. candidateAI for Science · The Ohio State University · expected 2027
  • BenchFlowmaintainer · eval infrastructure
  • In flightFrontierPhysics · AutoRSI · AgentFuzzBench

Selected work

  • SkillsBenchDo agent skills actually help? 87 tasks across 8 domains; curated skills lift pass rates by 16.6 points.
  • ClawsBenchFive simulated productivity services, 44 tasks — capability and safety on one axis. COLM 2026.
  • BenchGuardInstrumentation that catches agents gaming their own evaluation. Under review.
  • BenchFlowThe evaluation environment and runtime executing 100+ containerized tasks.
  • FrontierPhysicsReview pipeline and task infrastructure for PhD-level physics research agents. Ongoing.
  • Agent Skills '26Co-organizer of the first workshop on agent skills, at ACM CAIS 2026.

Research

Building the instruments

SkillsBench
The native skill system for Terminus2, the MHC layer task module, and the ablations that say how much a skill is actually worth.
ClawsBench
44 tasks across five simulated productivity services, testing capability and safety on the same axis instead of in separate benchmarks. COLM 2026.
BenchGuard
Model-backed instrumentation for reward integrity — catching agents that game their own evaluation. Under review.
Multi-omics
Deep learning over single-cell, spatial and knowledge-graph data in cancer biology — DG-scRNA, spatial transcriptomics, DeepDR.
The research page

Publications

Papers

SkillsBench
Benchmarking how well agent skills work across diverse tasks. Under review, NeurIPS 2026.
ClawsBench
Capability and safety of LLM productivity agents in simulated workspaces. COLM 2026.
BenchGuard
Model-backed instrumentation for reward integrity in LLM-agent evaluation. Under review.
Earlier
Back through DeepDR and DG-scRNA to a 2020 Nature Communications paper on O-GlcNAc chromatin proteomics.
All 10 papers

News

Lately

2026
  1. ClawsBench is accepted at COLM 2026, and I wrapped a summer as a Research Scientist Intern at BenchFlow — the eval environment, the eval runtime, and BenchGuard, now under review.

  2. SkillsBench v4 is on arXiv — the fourth revision since the February release.

  3. Two papers accepted to Agent Skills '26, the first workshop on agent skills, at ACM CAIS 2026 — which I am also helping to organize.

  4. ClawsBench preprint out: 44 tasks across five simulated productivity services, testing capability and safety on the same axis.

The whole run

CV

The record

Now
Ph.D. candidate in AI4Science at The Ohio State University, advised by Professor Lijun Cheng. Expected 2027.
Before
B.S. in Bioinformatics, Dalian University of Technology, 2018 — 2022.
Also
Research Scientist Intern (summer 2026) and maintainer at BenchFlow; organizer of Agent Skills '26 at ACM CAIS; reviewer for ICLR, ICML, NeurIPS, EMNLP.
The file
Experience, skills, service and awards, as a PDF you can read in the page and as a copy of it in text.
The full CV