About Me
👋 Hi, I’m Wen-Chi Tsai (Vicky)
I’m an MS student at Carnegie Mellon University (MIIS — Master of Intelligent Information Systems), with a CS degree from National Central University (NCU, Taiwan).
I build practical AI systems that connect LLMs, retrieval, evaluation, and real engineering workflows — from LLM infrastructure serving 200+ engineers at ASUS, to compressed models running on factory edge devices, to AI products used by real customers.
I’m seeking Summer 2027 internships in AI systems, LLM infrastructure, ML systems, and applied AI.
🏢 Experience
Software Engineering Intern — ASUS, Open Cloud Infrastructure Software R&D
Shipped production LLM infrastructure used daily by 200+ engineers across 5 product teams:
- RAG assistant indexing 25+ heterogeneous sources — cut support tickets by 25% at 85% answer accuracy
- Context-aware retrieval pipeline — improved AI code-suggestion relevance by 30%
- LLM-powered issue triage — reduced median triage time by 38%, with 80% agreement with human reviewers
- TypeScript E2E testing framework — raised coverage from 45% to 63%, cutting manual regression effort by 50%
Undergraduate Researcher — NCU Wireless Ad-Hoc & Sensor Networks Lab
- LLM compression: distilled Llama-3.2-3B to 1B via LoRA/QLoRA knowledge distillation — 67% fewer parameters, 2× faster inference, 95%+ teacher performance retained — for factory edge deployment
- UAV inspection: dual-strategy path planner achieving 97.3% surface coverage and 99.7% collision-free rate with 5× planning-time reduction — published at Ubi-Media 2026
AI Research Intern — AI Post
- Technical articles on AI agents and LLM development, reaching 50,000+ readers
- Taught 3 workshops with 200+ participants (4.8/5 satisfaction)
🚀 Products I’ve Shipped
- RoomMuse — AI interior-design configurator with AR room scanning, floor-plan OCR, and a building-code-aware spatial validation engine; secured $20K Phase 1 development funding from Singapore furniture company MOZU
- Museum artifact recognition app (AIGO) — YOLOv8 + GAN pipeline at 94% accuracy, Dockerized on AWS with 75% p95 latency reduction; iOS app rated 4.6/5 by 500+ visitors
🔬 Focus Areas
- LLM & Retrieval Systems: RAG, context-aware retrieval, vector search (LanceDB), agent workflows, LLM evaluation
- ML Systems: knowledge distillation, LoRA/QLoRA, ONNX, inference optimization, edge deployment
- Infrastructure: FastAPI, Docker, AWS, PostgreSQL, CI/CD, load testing
🎧 Outside of Tech
I enjoy concerts, photography, and following the work of artists like NewJeans, Mamamoo, and aespa.
Good music always helps me think more clearly while coding.
📬 Connect With Me
- LinkedIn: https://www.linkedin.com/in/wen-chi-tsai/
- GitHub: https://github.com/vicky0619
- Website (Portfolio): https://vicky0619.github.io/
- Email: vicky46586038@gmail.com