About 自己紹介 关于我

こんにちは, I'm Ashish

  • 日本語アシシュ・クマール
  • 中文阿希什·库马尔
  • РусскийАшиш Кумар
  • DeutschAschisch Kumar
  • 한국어아시시 쿠마르
第一章

How it started

I wrote my first Hello, World! at eleven and I've been poking at computers ever since. For years that meant whatever looked fun that week: little Python games, desktop tools, bots for Telegram and Discord, browser extensions, and a long run of web apps built mostly to learn how the pieces fit together. None of them changed the world, but together they taught me the unglamorous half of programming. Read the error message slowly. Ship before it feels ready. Delete code without taking it personally.

Somewhere in the middle of all that I trained my first neural network, watched the loss go down, and quietly stopped caring about almost everything else. Those older experiments still live in the archive, and I'm fond of them the way you're fond of old school notebooks.

第二章

What I work on now

These days I build language models end to end, and I mean every step. I clean the data, train the tokenizer, write the attention and expert layers from the paper rather than from someone else's repository, pretrain, fine-tune, and then score the result with exactly the same harness I use for the models I compare against. The last step is my favourite: wrapping it in a small app so a real person can talk to it.

Doing this on a student budget turns out to be a great teacher. When every GPU hour comes out of your own pocket, you learn to read the profiler, to trust small experiments before big ones, and to never launch a run you haven't thought through. Mixture-of-Experts is my current obsession. Sending each token to a handful of specialists is such a simple idea, and there are so many interesting ways for it to go wrong.

第三章

All the way down

Models don't run on vibes, they run on silicon, and I want to understand that layer too. I write CUDA to find out where the time really goes, I've spent more evenings than I'd admit staring at memory access patterns, and lately I've been learning Verilog one logic gate at a time. The destination is FPGAs and hardware–software codesign: shaping the model and the chip around each other instead of forcing one to fit the other.

The same curiosity pulls me towards robotics, embodied AI and AR/VR. That's where a model finally has to cope with a body, a camera and a world that doesn't come with labels, instead of a tidy benchmark.

第四章

What I believe

I care a lot about how we get to AGI, and I think the honest route is many careful, reproducible steps rather than a few loud ones. So I check approximations against exact answers, I report the numbers that didn't go my way, and I share weights, datasets and training logs whenever I can.

I also think research should end in something people can use. A model nobody can try is only half finished, and a result nobody can reproduce is mostly a rumour.

第五章

Right now

I'm an undergrad graduating in 2029, reading more papers than I can implement and implementing more than I probably should. If you're working on language models, generative media, world models or the systems underneath them, I'd love to hear about it. Ideas, collaborations and hard questions are all welcome.