News
Lately
Log
-
ClawsBench is accepted at COLM 2026, and BenchGuard — model-backed instrumentation for reward integrity in agent evaluation — is under review. That closes out a summer as a Research Scientist Intern at BenchFlow.
-
SkillsMetric is on arXiv — mapping where static analysis stops being able to tell a useful agent skill from a hostile one.
-
SkillsBench v4 lands on arXiv — the fourth revision since the February release.
-
Two papers accepted to Agent Skills '26, the first workshop on agent skills, at ACM CAIS 2026 — which I am also helping to organize.
-
Started the summer as a Research Scientist Intern at BenchFlow in San Francisco — the eval environment, the evaluation runtime, and the benchmarks that run on it.
-
ClawsBench preprint out: 44 tasks across five simulated productivity services, testing capability and safety on the same axis.
-
SkillsBench v1 out on arXiv (v2 and v3 follow in March). I built the native skill system for Terminus2 and the ablation design behind the results.
-
Joined BenchFlow as an open-source contributor and maintainer.
-
DeepDR preprint out with Shuting Jin's group at WUST — 15+ deep learning models over a knowledge graph of 5.9M edges.
Education
Education
-
The Ohio State University Aug 2022 — present
Ph.D. in AI4Science · Columbus, OH
Advisor: Professor Lijun Cheng
-
Dalian University of Technology Sep 2018 — Jun 2022
B.S. in Bioinformatics · Dalian, CN