CV

Yimin Liu

Ph.D. Candidate in AI4Science, The Ohio State University. Columbus, Ohio · yiminliu.career at gmail.com

2027 Ph.D. expected

Sheet

This browser will not display the PDF inline. Download it instead.

The transcription below carries the same content for anything that cannot render the sheet.

Education

Education

  • The Ohio State University Aug 2022 — present

    Ph.D. in AI4Science · Columbus, OH

    Advisor: Professor Lijun Cheng

  • Dalian University of Technology Sep 2018 — Jun 2022

    B.S. in Bioinformatics · Dalian, CN

    GPA 86.4 / 100

Experience

Experience

  • Research Scientist Intern & Open-Source Contributor Jan 2026 — present

    BenchFlow · San Francisco, CA

    Research Scientist Intern: May — Aug 2026

    • Evaluation infrastructure — built the BenchFlow evaluation environment and runtime powering benchmark suites including SkillsBench, AgentFuzzBench, and FrontierPhysics (100+ containerized tasks) across coding, terminal, and productivity domains.
    • BenchGuard (under review, USENIX Security 2027) — designed a model-backed instrumentation layer combining phase-aware static taint analysis with runtime infrastructure-side evidence to detect reward hacking; built BenchGuard Trajectories, a human-labeled corpus of 456 adjudicated trajectories from 31K+ public agent runs, reaching 96% detection accuracy and lifting full-chain recall from 23–94% to 77–100%.
    • ClawsBench (COLM 2026) — built a benchmark simulating five productivity services (Gmail, Slack, Google Calendar, Google Docs, Google Drive) with 44 tasks testing both capability and safety of LLM agents.
    • SkillsBench (87 tasks / 8 domains / 18 model–harness configs; under review, NeurIPS 2026) — implemented the native skill system for Terminus2, where curated skills lift mean pass rate from 33.9% to 50.5% (+16.6 pp); developed the mHC layer task, a credentialed A100 GPU task implementing manifold-constrained hyper-connections in nanoGPT bundling 3 reusable skills; built the analysis pipeline from raw trajectories to paper figures and led code review.
    • FrontierPhysics (ongoing) — building the contribution review pipeline and task infrastructure (Docker environments, automated verifiers, rubric-based LLM-as-judge reviewer agents) for an open benchmark of PhD-level physics research tasks.
  • Independent Research Collaborator Jun 2026 — present

    AutoRSI & AgentFuzzBench · Remote

    Collaborators: Yuxuan Jiang (AutoRSI); Security Lab, UC Davis (AgentFuzzBench)

    • AutoRSI (manuscript in preparation) — a general framework that automatically strengthens a naive base method across domains via iterative propose–implement–evaluate–reflect cycles; building the RSI loop pipeline and the automatic construction and refinement of the improvement-plan pool.
    • AgentFuzzBench — a benchmark for end-to-end agentic fuzzing (build, harnessing, fuzzing, crash triage); built the uniform task-package pipeline across 8 exact-pinned OSS targets and packaged it as an externally hosted BenchFlow benchmark with committed AFL++ oracle evidence.
  • Research Collaborator Sep 2025 — Nov 2025

    Wuhan University of Science and Technology · Wuhan, CN (remote)

    Collaborator: Shuting Jin, School of Computer Science and Technology

    • DeepDR — built the drug-repositioning web platform: 15+ deep learning models over a knowledge graph of 5.9M edges and 107 relationship types, with visualizations explaining why each drug is recommended.
  • Graduate Research Assistant Mar 2023 — present

    The Ohio State University · Columbus, OH

    Advisor: Professor Lijun Cheng

    • Developed DG-scRNA, a graph-convolutional network annotating cell types from scRNA-seq data — 97.2% accuracy on 150K+ cells (first-author manuscript).
    • Integrated scRNA-seq and spatial transcriptomics across 25 thyroid cancer samples; identified 50+ ligand–receptor interactions and two novel cell subtypes.
    • Benchmarked CNN, RNN, and Transformer architectures for CRISPR gRNA on-target prediction over 50K+ curated sequences.
  • Research Assistant Aug 2021 — Oct 2021

    Harvard Medical School · Boston, MA

    Advisor: Martin Hemberg

    • Built an automated mutation-calling pipeline (GATK, R) over 50K+ cells from public scRNA-seq datasets; clustered on SNV information and expression matrices separately to predict variant–phenotype relationships in cancer.
  • Research Assistant Mar 2021 — Aug 2021

    Beijing Institute of Genomics, Chinese Academy of Sciences · Beijing, CN

    Advisor: Mingkun Li

    • Built a pipeline to optimize COVID-19 hybrid capture probes; located low-efficiency regions and redesigned them, improving sensitivity by 25%.
  • Research Assistant Nov 2019 — Feb 2021

    Dalian University of Technology · Dalian, CN

    Advisor: Yubo Liu

    • Undergraduate research on O-GlcNAc chromatin regulation in breast cancer (proteomics, ChIP-seq, RNA-seq); co-authored Nature Communications 2020, Cell Biology International 2021, and BBA — General Subjects 2021.

Skills

Skills

LLM & agent evaluation
Benchmark design, eval harnesses & runtimes, reward integrity, LLM-as-judge, trajectory scoring
Programming & ML
Python, PyTorch (GCNs, Transformers), NumPy/Pandas/scanpy, R (Seurat, Tidyverse), Bash, JavaScript
Bioinformatics & genomics
scRNA-seq, spatial transcriptomics, drug repositioning, cellranger, samtools, STAR, bwa
Web, cloud & infrastructure
Full-stack development, REST APIs, Docker, HPC, Slurm, Git, Linux/Unix

Service

Service

  • Workshop Organizer — Agent Skills '26, The First Workshop on Agent Skills, ACM CAIS2026
  • Conference Reviewer — ICLR2024 — 2026
  • Conference Reviewer — ICML2024 — 2025
  • Conference Reviewer — NeurIPS2024 — 2025
  • Conference Reviewer — EMNLP2026
  • Journal Reviewer — PLOS Computational Biology