C Kaustubh
I build reliable distributed systems, ML pipelines, and performance-oriented developer tools with Rust, Go, Python, CUDA, and PyTorch.
Now
AIML Intern at AVEVA (Schneider Electric), building production LLM agents for industrial software.
Research
Uncertainty-aware RAG, RL post-training for ASR fairness, and video anomaly detection.
Open Source
LanceDB, Firecracker, and core contributor to Vektori.
Focus
LLM agents in production, ML systems, distributed infrastructure, and low-level performance.
Proof
Production AI agent work at AVEVA, Deployfest 2026 1st runner-up, Agentathon '26 winner, and Google DeepMind Bangalore Hackathon finalist.
Open To
Fall 2026 internships in ML, MLOps, systems, backend engineering, and applied AI R&D.
Featured: qwen_cuda
From-scratch CUDA C++ inference engine for Qwen3-8B with no frameworks: GGUF parsing, custom kernels, GQA with KV cache, flash attention, and CUDA graph capture in ~1600 lines. Runs at ~37 tok/s on an RTX 4090 at ~17.1 GB peak VRAM.