2026-02 — present
AI Research Engineer — Platform Team, Quantizer Part · Nota AI
- LLM/VLM quantization platform: GPTQ, AWQ, SmoothQuant, QuaRot/SpinQuant pipelines; NVFP4 and GGUF export paths; MoE all-expert calibration; KV-cache quantization (verifiable via merged internal PR record).
- Model enablement & evaluation: EXAONE 4.0 / MoE, Qwen3 / 3.5(-VL), Phi-3 (LongRoPE), gpt-oss, embedding/reranker and video-diffusion (Wan, Cosmos) evaluation harnesses.
- K-EXAONE 236B (MoE) optimization for FuriosaAI's data-center NPU, with LG AI Research — ~71% model-size reduction at ~99.2% accuracy retention (GPQA 79.80, IFBench 68.98, AIME25 88.57); announced 2026-06-30.