Model
About
What I build
I build the infrastructure that decides whether LLM agents actually work — the skill systems, harnesses, benchmarks, and reward-integrity tooling that turn “the agent seems good” into a number you can defend.
An eval stack fails in three places: the runtime, the tasks, and the score itself. At BenchFlow (Research Scientist Intern, summer 2026; maintainer since January) I built for all three. The BenchEval runtime executes 100+ containerized tasks across coding, terminal, and productivity domains. SkillsBench asks whether skills actually help — 87 tasks across 8 domains, where curated skills lift pass rates by 16.6 points — and ClawsBench (COLM 2026) puts capability and safety on the same axis instead of in separate benchmarks. And BenchGuard defends the score itself: phase-aware taint analysis plus runtime evidence that catches agents gaming their own evaluation.
My Ph.D. is in AI for Science: deep learning for single-cell and spatial multi-omics in cancer — which is also where the hard, real scientific tasks in my evaluations come from.
Research
Building the instruments
- SkillsBench
- The native skill system for Terminus2, the MHC layer task module, and the ablations that say how much a skill is actually worth.
- ClawsBench
- 44 tasks across five simulated productivity services, testing capability and safety on the same axis instead of in separate benchmarks. COLM 2026.
- BenchGuard
- Model-backed instrumentation for reward integrity — catching agents that game their own evaluation. Under review at USENIX Security '27.
- Multi-omics
- Deep learning over single-cell, spatial and knowledge-graph data in cancer biology — DG-scRNA, spatial transcriptomics, DeepDR.
Publications
Papers
- ClawsBench
- Capability and safety of LLM productivity agents in simulated workspaces. COLM 2026.
- BenchGuard
- Model-backed instrumentation for reward integrity in LLM-agent evaluation. Under review, USENIX Security 2027.
- Agent Skills '26
- What keeps agent skills from being reusable? Evidence from 138K SKILL.md files. Workshop paper, ACM CAIS.
- Earlier
- Back through DeepDR and DG-scRNA to a 2020 Nature Communications paper on O-GlcNAc chromatin proteomics.
News
Lately
-
ClawsBench is accepted at COLM 2026, and I wrapped a summer as a Research Scientist Intern at BenchFlow — the eval environment, the BenchEval runtime, and BenchGuard, now under review at USENIX Security '27.
-
Two papers accepted to Agent Skills '26, the first workshop on agent skills, at ACM CAIS 2026 — which I am also helping to organize.
-
ClawsBench preprint out: 44 tasks across five simulated productivity services, testing capability and safety on the same axis.
CV
The record
- Now
- Ph.D. candidate in AI4Science at The Ohio State University, advised by Professor Lijun Cheng. Expected 2027.
- Before
- B.S. in Bioinformatics, Dalian University of Technology, 2018 — 2022.
- Also
- Research Scientist Intern (summer 2026) and maintainer at BenchFlow; organizer of Agent Skills '26 at ACM CAIS; reviewer for ICLR, ICML, NeurIPS, EMNLP.
- The file
- Experience, skills, service and awards, as a PDF you can read in the page and as a copy of it in text.