About Me
I am a Research Scientist in the ByteDance Seed Infrastructure Team. I received my Ph.D. and M.S. in Electrical and Computer Engineering from The University of Texas at Austin, and my B.S. in Computer Science from Shanghai Jiao Tong University.
Before joining ByteDance, I was a Graduate Researcher in UT Austin's Circuit Research Lab, where I worked on system, architecture, and circuit co-design for AI and combinatorial optimization accelerators.
My research interests span AI infrastructure, computer architecture, accelerator design, compute-in-memory, and system-circuit co-design. I am particularly interested in building efficient systems for large-scale AI training and inference, emerging graphics and rendering workloads, and combinatorial optimization.
News
- I joined the ByteDance Seed Infrastructure Team as a Research Scientist.
- Our paper, Charon: A Unified and Fine-Grained Simulator for Large-Scale LLM Training and Inference, will be presented at MLSys 2026. See you in Bellevue!
- Our paper, TACC: A 42.9 TOPS/W/mm² Ternary/INT/FP LLM Accelerator with Lossless Weight Compression, will be presented at CICC 2026. See you in Seattle!
- Our paper, GSAcc: Accelerate 3D Gaussian Splatting via Depth Speculation and Gaussian-Centric Rasterization, will be presented at DAC 2025. See you in San Francisco!
Selected Publications
View all publications
Charon: A Unified and Fine-Grained Simulator for Large-Scale LLM Training and Inference
Charon is a unified, modular simulator for both large-scale LLM training and inference. Its operator-level graph model captures computation, communication, and parallelism choices with under 5.35% prediction error across evaluated systems, while enabling substantially cheaper design exploration than cluster profiling.

TACC: A 42.9 TOPS/W/mm² Ternary/INT/FP LLM Accelerator with Lossless Weight Compression
TACC combines lossless ternary-weight compression with decompress-free execution, symmetric Gray coding, and reconfigurable LUT tensor cores. Fabricated in 16 nm CMOS, it supports ternary, integer, FP8, and BF16 computation while reaching 42.9 TOPS/W/mm².

GSAcc: Accelerate 3D Gaussian Splatting via Depth Speculation and Gaussian-Centric Rasterization
GSAcc restructures 3D Gaussian Splatting around depth speculation and a Gaussian-centric dataflow so preprocessing, sorting, and rasterization can overlap without storing large intermediate results. Dedicated sorting and rasterization units deliver up to 2.3× better PPA and 2.9× lower energy than the prior GSCore accelerator.

CILP: An Arbitrary-Bit Precision All-Digital Compute-in-Memory Solver for Integer Linear Programming Problems
CILP repurposes foundry 8T SRAM for in-memory constraint checking, objective evaluation, and variable updates in integer linear programs. Its all-digital, arbitrary-precision compute unit preserves exact feasibility while reducing data movement, measured cycles, and energy relative to software solvers.

Vecim: A 289.13 GOPS/W RISC-V Vector Co-Processor with Compute-in-Memory Vector Register File for Efficient High-Performance Computing
Vecim turns a RISC-V vector register file into an all-digital compute-in-memory engine for INT8, BF16, and FP16 multiply-add operations. The measured 65 nm chip reduces register-file data movement and sustains up to 289.13 GOPS/W while retaining a programmable vector architecture.

Snap-SAT: A One-Shot Energy-Performance-Aware All-Digital Compute-in-Memory Solver for Large-Scale Hard Boolean Satisfiability Problems
Snap-SAT maps Boolean clauses and variables into reusable SRAM so clause evaluation and local variable updates occur directly in memory. The measured all-digital 65 nm prototype scales to hard SAT instances while providing large speed and energy gains over CPU and mobile software solvers.

A 118 GOPS/mm² 3D eDRAM TensorCore Architecture for Large-Scale Matrix Multiplication
This architecture repurposes monolithic 3D eDRAM as a dense, reconfigurable matrix-multiplication fabric. A bit-stationary TensorCore dataflow supports INT8 and BF16 workloads, reaching 118 GOPS/mm² and improving compute density on NeRF and LLaMA-7B evaluations.
