Portrait of Mingyu Sung
1-bit → 8-bit

Mingyu Sung

AI Research Engineer — Efficient GenAI Inference (Quantization · Caching · Acceleration)

Nota AI, Platform Team (Quantizer) · South Korea

40 citations/h-index 4/7 published, 5 under review/14 merged PRs across 4 repos/as of 2026-07-20

AI Research Engineer at Nota AI (Platform Team, Quantizer part). Ph.D. in Artificial Intelligence (KNU, 2026). Works on efficient generative-model inference: model quantization, diffusion/LLM cache acceleration, CUDA kernel optimization, split computing, and training-free KV-cache compression. Core contributor to the Nunchaku CUDA acceleration engine (ICLR 2025 Spotlight); 14 merged PRs across Nunchaku, ComfyUI-nunchaku, ByteDance DreamO, and Hugging Face Transformers.

Experience

2026-02 — present
AI Research Engineer — Platform Team, Quantizer Part · Nota AI
  • LLM/VLM quantization platform: GPTQ, AWQ, SmoothQuant, QuaRot/SpinQuant pipelines; NVFP4 and GGUF export paths; MoE all-expert calibration; KV-cache quantization (verifiable via merged internal PR record).
  • Model enablement & evaluation: EXAONE 4.0 / MoE, Qwen3 / 3.5(-VL), Phi-3 (LongRoPE), gpt-oss, embedding/reranker and video-diffusion (Wan, Cosmos) evaluation harnesses.
  • K-EXAONE 236B (MoE) optimization for FuriosaAI's data-center NPU, with LG AI Research — ~71% model-size reduction at ~99.2% accuracy retention (GPQA 79.80, IFBench 68.98, AIME25 88.57); announced 2026-06-30.

Open source14 merged PRs

Nunchaku

nunchaku-ai/nunchaku  ·  Core Contributor  ·  9 merged  ·  2025 – 2026-02
CUDA acceleration engine for 4-bit neural networks (ICLR 2025 Spotlight, 3.9k+ stars)
  • Cache optimizations — V2 FBCaching (#621), double FB cache, TeaCache batch processing (#601) — for 2–5x additional speedup
  • IP-Adapter (XLabs flux-ip-adapter-v2) support (#418)
  • Fixes: LoRA key mismatch (#557), ControlNet (#360, #452), Sana (#380), offload segfault (#440)
  • Forward-pass tensor-op cleanup (#491)

ComfyUI-nunchaku

nunchaku-ai/ComfyUI-nunchaku  ·  Contributor  ·  2 merged  ·  2025
ComfyUI integration for the Nunchaku engine
  • IP-Adapter support for ComfyUI (#305)
  • Cache-mechanism file separation (#474)

DreamO

bytedance/DreamO  ·  Contributor  ·  1 merged  ·  2025
ByteDance's unified image-customization framework
  • Dynamic FBCache / DoubleFBCache support for Nunchaku engine integration (#104)

Transformers

huggingface/transformers  ·  Contributor  ·  2 merged  ·  2026
Hugging Face Transformers library
  • Fixed float16 overflow in Gemma4 vision pooler (#46277)
  • Fixed Gemma `sliding_window` being halved on every config save/reload, a silent accuracy regression for EmbeddingGemma (#47940)

Publications7 published

2025
H2-Cache: A Novel Hierarchical Dual-Stage Cache for High-Performance Acceleration of Generative Diffusion ModelsIEEE Open Journal of the Computer Society · M. Sung, I.-M. Kim, S. Yun, J.-M. Kang
A Novel VLM-Guided Diffusion Model for Remote Sensing Image Super-ResolutionIEEE Geoscience and Remote Sensing Letters · M. Sung, M.-G. Gong, S.-J. Ham, I.-M. Kim, S. Yun, J.-M. Kang cited ×2 Best Paper Award, KNU-EE Research Congress
DeCo-MeSC: Deep Compression-Based Memory-Constrained Split Computing Framework for Cooperative Inference of Neural NetworkIEEE Transactions on Vehicular Technology · M. Sung, V. Palakonda, I.-M. Kim, S. Yun, J.-M. Kang cited ×6
Generative Diffusion Model-Based Deep Learning Framework for Remaining Useful Life PredictionIEEE Internet of Things Journal · S. Ha, M. Sung, F. Saeed, S. Yun, I.-M. Kim, J.-M. Kang cited ×7 Co-first author
2024
Entropy-based sampling for efficient training of deep learning on CNC machining datasetElectronics Letters · M. Sung, C. Park, S. Ha, M. Ha, H. Lee, J. Kim, J.-M. Kang cited ×1
2023
2020

Preprints5 under review

Research

2026
StreamPRS
Single-context-prefill approximation of KVzip's reconstruction-based KV-cache eviction; contributions framed as Factorization / Method / System. (Scope-limited per claims ledger — do not overstate.)

Projects

2026 · Quantization Engineer
K-EXAONE 236B MoE Optimization (Nota AI × LG AI Research × FuriosaAI)
Optimized the 236B-parameter MoE model for FuriosaAI's data-center NPU via targeted precision analysis of degradation-prone sections — ~71% size reduction with ~99.2% accuracy retention (GPQA 79.80 / IFBench 68.98 / AIME25 88.57).
2024 · Lead Developer
LLM Inference Optimization (Samsung Challenge)
Optimized pipeline for Phi-3 using Torch-TensorRT conversion and memory-aware dynamic batching to accelerate LLM inference.
2023 · Lead Developer
Advanced CNC Machine Tool Diagnosis & Prognosis (KERI)
Deep learning models for tool wear classification, anomaly detection, and RUL prognosis.
2023 · AI Optimization Engineer
AI-based Intelligent CCTV (ABB Project)
Optimized an LSTR behavior-detection model for edge devices via Knowledge Distillation (TransKD) and Deep Compression — 25x model compression.
2022 · Participating Researcher
Global Basic Research Laboratory (NRF Project)
Deep learning-based channel estimation (CNN denoising, self-supervised learning) for IRS-aided communication systems.
2022 · Lead Developer
ML-based Sensor Data Analysis (JS System Project)
ML-based sensor feature-importance extraction and model performance optimization.

Education

2021-08 — 2026-02
Ph.D., Artificial Intelligence · Kyungpook National University (KNU)
Thesis: Fast and Trustworthy Super-Resolution with VLM-Guided Diffusion Model
2019-09 — 2021-08
M.S., Computer Science · Kyungpook National University (KNU)
Thesis: Probabilistic Classification Method of Spiking Neural Network Based on Multi-Labeling of Neurons
2013-03 — 2019-08
B.S., Computer Science · Kyungpook National University (KNU)

Awards

Skills

Research
model quantizationKV-cache compressiondiffusion/LLM inference accelerationsplit computingsuper-resolutionlong-context evaluation
Programming
PythonC++CUDA CMATLAB
Frameworks & Tools
PyTorchHuggingFace TransformersTensorRTTritonflash-attnDockerGitLaTeX
Methodology
preregistrationclaims ledgerfreeze discipline