News

Lately

Log

2026
  1. ClawsBench is accepted at COLM 2026, and BenchGuard — model-backed instrumentation for reward integrity in agent evaluation — is under review. That closes out a summer as a Research Scientist Intern at BenchFlow.

  2. SkillsMetric is on arXiv — mapping where static analysis stops being able to tell a useful agent skill from a hostile one.

  3. SkillsBench v4 lands on arXiv — the fourth revision since the February release.

  4. Two papers accepted to Agent Skills '26, the first workshop on agent skills, at ACM CAIS 2026 — which I am also helping to organize.

  5. Started the summer as a Research Scientist Intern at BenchFlow in San Francisco — the eval environment, the evaluation runtime, and the benchmarks that run on it.

  6. ClawsBench preprint out: 44 tasks across five simulated productivity services, testing capability and safety on the same axis.

  7. SkillsBench v1 out on arXiv (v2 and v3 follow in March). I built the native skill system for Terminus2 and the ablation design behind the results.

  8. Joined BenchFlow as an open-source contributor and maintainer.

2025
  1. DeepDR preprint out with Shuting Jin's group at WUST — 15+ deep learning models over a knowledge graph of 5.9M edges.

Education

Education

  • The Ohio State University Aug 2022 — present

    Ph.D. in AI4Science · Columbus, OH

    Advisor: Professor Lijun Cheng

  • Dalian University of Technology Sep 2018 — Jun 2022

    B.S. in Bioinformatics · Dalian, CN