Huanmi Tan
谭欢秘, pronounced as /'hwɑnmi tæn/

I am a software engineer at AWS, on the Web Intelligence team within the Agentic AI organization. We build local and remote web browsing agents that perceive and act on live web interfaces to carry out real user tasks, along with the tooling and MCP integrations that connect them to external systems.

I received my Master's degree in Software Engineering from SCS of Carnegie Mellon University, and my Bachelor's degree in Software Engineering from Tongji University. During my Bachelor's, I interned at Shanghai AI Lab and ByteDance AI Lab.

My work sits between language models and the systems they act on. Current interests:

  • LLM agents: web agents, browser automation, tool use and MCP
  • Evaluation: automated benchmark construction and failure analysis
  • Code generation: correctness, execution feedback, and optimization
  • Reasoning: logical and constraint-based problem solving in LLMs

Building agents in production offers a vantage point that is hard to reach from benchmarks alone: real websites change, break, and fail in ways that frozen evaluation environments never surface. I am always glad to talk about agent evaluation, and I am open to collaboration.


Education
  • Carnegie Mellon University
    Carnegie Mellon University
    Master's in Software Engineering
    Aug. 2023 - Dec. 2024
  • Tongji University
    Tongji University
    B.E. in Software Engineering
    Sep. 2019 - Jul. 2023
  • North Carolina State University
    North Carolina State University
    GEARS Research Program
    Jun. 2022 - Aug. 2022
Experience
  • Amazon Web Services
    Amazon Web Services
    SDE | Web Intelligence, Amazon Quick
    Jun. 2025 - Present
  • ByteDance AI Lab
    ByteDance AI Lab
    Intern | Volcano Machine Translation, AI Lab
    Mar. 2023 - Aug. 2023
  • Shanghai AI Lab
    Shanghai AI Lab
    Research Intern | AI for Imaging Group
    Nov. 2022 - Feb. 2023
News
2026
One of our papers got accepted to COLM 2026!
Jul 08
I got 80+ citations 🎉
Jul 01
2025
I proudly became a cat lady! Introducing my baby girl, Juzai.
Sep 24
One of our papers got accepted to EMNLP main. See you in Suzhou, China!
Aug 20
I started a full time position at AWS!
Jun 09
2024
I received my Master's degree from CMU!
Dec 18
Selected Publications (view all )
SuperCoder: Assembly Program Superoptimization with Large Language Models

Anjiang Wei, Tarun Suresh, Huanmi Tan, Yinglun Xu, Gagandeep Singh, Ke Wang, Alex Aiken

COLM 2026

We build the first large-scale benchmark for LLM-based assembly superoptimization (8,072 real-world programs), then fine-tune a model with reinforcement learning plus best-of-N sampling and iterative refinement — reaching 95.0% correctness and a 1.46x average speedup over gcc -O3.

VeriCoder: Enhancing LLM-Based RTL Code Generation through Functional Correctness Validation

Anjiang Wei, Huanmi Tan, Tarun Suresh, Daniel Mendoza, Thiago SFX Teixeira, Ke Wang, Caroline Trippel, Alex Aiken

NuerIPS 2025 Fourth Workshop on Deep Learning for Code 2025

We build a 125,000+ example RTL dataset via feedback-directed refinement — iteratively fixing designs and tests against simulation results, rather than relying on syntactic checks alone — then fine-tune a code model on it, improving over prior work by up to 71.7% on VerilogEval.

VeriCoder: Enhancing LLM-Based RTL Code Generation through Functional Correctness Validation
SATBench: Benchmarking LLMs’ Logical Reasoning via Automated Puzzle Generation from SAT Formulas

Anjiang Wei, Yuheng Wu, Yingjia Wan, Tarun Suresh, Huanmi Tan, Zhanke Zhou, Sanmi Koyejo, Ke Wang, Alex Aiken

EMNLP Main 2025

SATBench automatically converts Boolean satisfiability formulas into natural-language logic puzzles, targeting the search-based reasoning prior benchmarks miss by focusing on rule-based inference. Even the strongest model we test, o4-mini, scores only 65.0% on hard puzzles, barely above chance.

Improving Assembly Code Performance with Large Language Models via Reinforcement Learning

Anjiang Wei, Tarun Suresh, Huanmi Tan, Yinglun Xu, Gagandeep Singh, Ke Wang, Alex Aiken

NeurIPS 2025 Fourth Workshop on Deep Learning for Code 2025

Early workshop version of our COLM 2026 paper: we introduce an 8,072-program benchmark and train an LLM with PPO reinforcement learning, rewarding both test-case correctness and speedup over gcc -O3, reaching a 1.47x speedup while outperforming 20 larger baseline models.

All publications
Honors & Awards
  • Second Prize of Citi Cup (Fintech Innovation Application Competition)
    2022
  • Academic Excellence Scholarship of Tongji University (Top 10%)
    2022
  • Third Prize in Mobile Application Innovation Competition
    2021
  • First Prize of East China Region in Mobile Application Innovation Competition of CCCC
    2021
  • Academic Excellence Scholarship of Tongji University (Top 10%)
    2021
  • Academic Excellence Scholarship of Tongji University (Top 10%)
    2020
Visitor Map