CV
Yimin Liu
Ph.D. Candidate in AI4Science, The Ohio State University. Columbus, Ohio · yiminliu.career at gmail.com
Sheet
The transcription below carries the same content for anything that cannot render the sheet.
Education
Education
-
The Ohio State University Aug 2022 — present
Ph.D. in AI4Science · Columbus, OH
Advisor: Professor Lijun Cheng
-
Dalian University of Technology Sep 2018 — Jun 2022
B.S. in Bioinformatics · Dalian, CN
GPA 86.4 / 100
Experience
Experience
-
Research Scientist Intern & Open-Source Contributor Jan 2026 — present
BenchFlow · San Francisco, CA
Research Scientist Intern: May — Aug 2026
- Evaluation infrastructure — built the BenchFlow evaluation environment and runtime powering benchmark suites including SkillsBench, AgentFuzzBench, and FrontierPhysics (100+ containerized tasks) across coding, terminal, and productivity domains.
- BenchGuard (under review, USENIX Security 2027) — designed a model-backed instrumentation layer combining phase-aware static taint analysis with runtime infrastructure-side evidence to detect reward hacking; built BenchGuard Trajectories, a human-labeled corpus of 456 adjudicated trajectories from 31K+ public agent runs, reaching 96% detection accuracy and lifting full-chain recall from 23–94% to 77–100%.
- ClawsBench (COLM 2026) — built a benchmark simulating five productivity services (Gmail, Slack, Google Calendar, Google Docs, Google Drive) with 44 tasks testing both capability and safety of LLM agents.
- SkillsBench (87 tasks / 8 domains / 18 model–harness configs; under review, NeurIPS 2026) — implemented the native skill system for Terminus2, where curated skills lift mean pass rate from 33.9% to 50.5% (+16.6 pp); developed the mHC layer task, a credentialed A100 GPU task implementing manifold-constrained hyper-connections in nanoGPT bundling 3 reusable skills; built the analysis pipeline from raw trajectories to paper figures and led code review.
- FrontierPhysics (ongoing) — building the contribution review pipeline and task infrastructure (Docker environments, automated verifiers, rubric-based LLM-as-judge reviewer agents) for an open benchmark of PhD-level physics research tasks.
-
Independent Research Collaborator Jun 2026 — present
AutoRSI & AgentFuzzBench · Remote
Collaborators: Yuxuan Jiang (AutoRSI); Security Lab, UC Davis (AgentFuzzBench)
- AutoRSI (manuscript in preparation) — a general framework that automatically strengthens a naive base method across domains via iterative propose–implement–evaluate–reflect cycles; building the RSI loop pipeline and the automatic construction and refinement of the improvement-plan pool.
- AgentFuzzBench — a benchmark for end-to-end agentic fuzzing (build, harnessing, fuzzing, crash triage); built the uniform task-package pipeline across 8 exact-pinned OSS targets and packaged it as an externally hosted BenchFlow benchmark with committed AFL++ oracle evidence.
-
Research Collaborator Sep 2025 — Nov 2025
Wuhan University of Science and Technology · Wuhan, CN (remote)
Collaborator: Shuting Jin, School of Computer Science and Technology
- DeepDR — built the drug-repositioning web platform: 15+ deep learning models over a knowledge graph of 5.9M edges and 107 relationship types, with visualizations explaining why each drug is recommended.
-
Graduate Research Assistant Mar 2023 — present
The Ohio State University · Columbus, OH
Advisor: Professor Lijun Cheng
- Developed DG-scRNA, a graph-convolutional network annotating cell types from scRNA-seq data — 97.2% accuracy on 150K+ cells (first-author manuscript).
- Integrated scRNA-seq and spatial transcriptomics across 25 thyroid cancer samples; identified 50+ ligand–receptor interactions and two novel cell subtypes.
- Benchmarked CNN, RNN, and Transformer architectures for CRISPR gRNA on-target prediction over 50K+ curated sequences.
-
Research Assistant Aug 2021 — Oct 2021
Harvard Medical School · Boston, MA
Advisor: Martin Hemberg
- Built an automated mutation-calling pipeline (GATK, R) over 50K+ cells from public scRNA-seq datasets; clustered on SNV information and expression matrices separately to predict variant–phenotype relationships in cancer.
-
Research Assistant Mar 2021 — Aug 2021
Beijing Institute of Genomics, Chinese Academy of Sciences · Beijing, CN
Advisor: Mingkun Li
- Built a pipeline to optimize COVID-19 hybrid capture probes; located low-efficiency regions and redesigned them, improving sensitivity by 25%.
-
Research Assistant Nov 2019 — Feb 2021
Dalian University of Technology · Dalian, CN
Advisor: Yubo Liu
- Undergraduate research on O-GlcNAc chromatin regulation in breast cancer (proteomics, ChIP-seq, RNA-seq); co-authored Nature Communications 2020, Cell Biology International 2021, and BBA — General Subjects 2021.
Skills
Skills
- LLM & agent evaluation
- Benchmark design, eval harnesses & runtimes, reward integrity, LLM-as-judge, trajectory scoring
- Programming & ML
- Python, PyTorch (GCNs, Transformers), NumPy/Pandas/scanpy, R (Seurat, Tidyverse), Bash, JavaScript
- Bioinformatics & genomics
- scRNA-seq, spatial transcriptomics, drug repositioning, cellranger, samtools, STAR, bwa
- Web, cloud & infrastructure
- Full-stack development, REST APIs, Docker, HPC, Slurm, Git, Linux/Unix
Service
Service
- Workshop Organizer — Agent Skills '26, The First Workshop on Agent Skills, ACM CAIS2026
- Conference Reviewer — ICLR2024 — 2026
- Conference Reviewer — ICML2024 — 2025
- Conference Reviewer — NeurIPS2024 — 2025
- Conference Reviewer — EMNLP2026
- Journal Reviewer — PLOS Computational Biology—