Model
About
What I build
I build the infrastructure that decides whether LLM agents actually work — the skill systems, harnesses, benchmarks, and reward-integrity tooling that turn “the agent seems good” into a number you can defend.
An eval stack fails in three places: the runtime, the tasks, and the score itself. At BenchFlow (Research Scientist Intern, summer 2026; maintainer since January) I built for all three. Its evaluation runtime executes 100+ containerized tasks across coding, terminal, and productivity domains. SkillsBench asks whether skills actually help — 87 tasks across 8 domains, where curated skills lift pass rates by 16.6 points — and ClawsBench (COLM 2026) puts capability and safety on the same axis instead of in separate benchmarks. And BenchGuard defends the score itself: phase-aware taint analysis plus runtime evidence that catches agents gaming their own evaluation.
My Ph.D. is in AI for Science: deep learning for single-cell and spatial multi-omics in cancer — which is also where the hard, real scientific tasks in my evaluations come from.
Now
- Ph.D. candidateAI for Science · The Ohio State University · expected 2027
- BenchFlowmaintainer · eval infrastructure
- In flightFrontierPhysics · AutoRSI · AgentFuzzBench
Selected work
- SkillsBenchDo agent skills actually help? 87 tasks across 8 domains; curated skills lift pass rates by 16.6 points.
- ClawsBenchFive simulated productivity services, 44 tasks — capability and safety on one axis. COLM 2026.
- BenchGuardInstrumentation that catches agents gaming their own evaluation. Under review.
- BenchFlowThe evaluation environment and runtime executing 100+ containerized tasks.
- FrontierPhysicsReview pipeline and task infrastructure for PhD-level physics research agents. Ongoing.
- Agent Skills '26Co-organizer of the first workshop on agent skills, at ACM CAIS 2026.
Research
Building the instruments
- SkillsBench
- The 300-file migration to BenchFlow-native task packages, the paired statistical evaluation, and the analysis behind every figure.
- ClawsBench
- 44 tasks across five simulated productivity services, testing capability and safety on the same axis instead of in separate benchmarks. COLM 2026.
- BenchGuard
- Model-backed instrumentation for reward integrity — catching agents that game their own evaluation. Under review.
- Multi-omics
- Deep learning over single-cell, spatial and knowledge-graph data in cancer biology — DG-scRNA, spatial transcriptomics, DeepDR.
Publications
Papers
- SkillsBench
- Benchmarking how well agent skills work across diverse tasks. Under review, NeurIPS 2026.
- ClawsBench
- Capability and safety of LLM productivity agents in simulated workspaces. COLM 2026.
- BenchGuard
- Model-backed instrumentation for reward integrity in LLM-agent evaluation. Under review.
- Earlier
- Back through DeepDR and DG-scRNA to a 2020 Nature Communications paper on O-GlcNAc chromatin proteomics.
News
Lately
-
ClawsBench is accepted at COLM 2026, and I wrapped a summer as a Research Scientist Intern at BenchFlow — the eval environment, the eval runtime, and BenchGuard, now under review.
-
SkillsBench v4 is on arXiv — the fourth revision since the February release.
-
Two papers accepted to Agent Skills '26, the first workshop on agent skills, at ACM CAIS 2026 — which I am also helping to organize.
-
ClawsBench preprint out: 44 tasks across five simulated productivity services, testing capability and safety on the same axis.
CV
The record
- Now
- Ph.D. candidate in AI4Science at The Ohio State University, advised by Professor Lijun Cheng. Expected 2027.
- Before
- B.S. in Bioinformatics, Dalian University of Technology, 2018 — 2022.
- Also
- Research Scientist Intern (summer 2026) and maintainer at BenchFlow; organizer of Agent Skills '26 at ACM CAIS; reviewer for ICLR, ICML, NeurIPS, EMNLP.
- The file
- Experience, skills, service and awards, as a PDF you can read in the page and as a copy of it in text.