EP.01 第1话
Ashish Kumar
- 日本語アシシュ・クマール
- 中文阿希什·库马尔
- РусскийАшиш Кумар
- DeutschAschisch Kumar
- 한국어아시시 쿠마르
AI Systems & Research Engineer
I build AI end to end, from the model all the way down to the chip it runs on.
- AI
- Language modelsDiffusionVideoSpeechMultimodalWorld modelsAgentsRobotics
- Systems
- GPU kernelsDistributed trainingInferenceFPGAsHW/SW codesign
Field map 分野 领域
Things I can't stop thinking about
Keep scrolling and the whole map slides past. Language models first; everything else is where they're headed.
About 自己紹介 关于我
Who's behind all this?
Hey, I'm Ashish, an undergrad (class of 2029) who got hooked on machine learning and never really recovered. What I enjoy most is building AI systems from the ground up. Give me a paper and an empty folder and I'll happily spend a week turning its equations into working layers, then another week finding out why the loss curve looks wrong. Most of that energy goes into large language models: tokenizers, pretraining runs, fine-tuning, evaluation, and the serving code that finally lets you chat with the thing you trained. I like knowing every part of the machine instead of trusting a black box, and I like showing people the numbers, including the unflattering ones. It's slower than stitching existing libraries together, but it's the only way I've found to actually understand what's going on inside. Honestly, watching a model I built from nothing string its first coherent sentence together is still the best feeling I know.
Language models sit in the middle, but they're not the whole picture. One eye stays on generative media: diffusion and flow models for images and video, speech and text-to-speech, and multimodal models that look and read at the same time. The other stays on world models, agents and reasoning, the pieces that feel closest to the big question of how we get to general intelligence. Underneath all of it is the machinery. I write CUDA kernels to see where the milliseconds go, I think about memory far more than is healthy, and I'm teaching myself digital design so that one day the model and the chip can be designed together. Robots, embodied AI and AR/VR are where I'd love all of this to land, somewhere a model has to deal with a body, a camera and the messy real world. It's a long list, and I have no intention of making it shorter.
Principles やり方 原则
How I like to build
Four rules I keep coming back to, whether it's a model, a kernel or a web app.
- 其の一
From the paper, not the repo
I implement architectures from the published papers instead of adapting released code. Slower, and I learn far more.
- 其の二
Measure, don't assume
Every approximation gets scored against the exact answer, and every baseline goes through the same harness instead of being quoted.
- 其の三
All the way down
From the model to the CUDA kernel to the logic gate. I like seeing the hardware underneath too.
- 其の四
Ship something usable
A chat app for a model, a demo for a library. People should be able to try the thing.
Mini game ゲーム 小游戏
Beat my transformer at tic-tac-toe
Win and I'll build your next idea for free. It scores every square, shows you where it's looking, then moves. It almost never loses.
Your move. You are ✕, the model is ◯.
Where the model is looking
- You
- 0
- Draws
- 0
- Model
- 0
Buy me a coffee 応援 请我喝咖啡
I run on coffee and rented GPUs
Everything here is trained on my own budget. If something helped you or made you smile, you can chip in for the next cup, the next GPU hour or a whole training run.
Buy me a coffeeSay hi 一緒に作ろう 一起做吧
Humans, aliens and fellow model-trainers welcome.
I'm always up for research ideas, collaborations and good questions.