Show HN: NanoEuler – GPT-2 scale model in pure C/CUDA from scratch

53 points · 7 visible top comments · 2026-06-28 19:38:14 UTC

Hi everyone,

I started working on nanoeuler after the ban of anthropic's fable because my ambition and dream is to work in the AI field in anthropic. The two interesting reasons that led me to create nanoeuler were (1) interfacing with llm does not mean understanding how they are composed and (2), working on llm with a very low-level layer to understand the correlation between parameters and data and growth of the model and how the GPU works and how some layers can be optimized.

So I started working on it with a research aspect by making nanoeuler grow more and more but doing one step after another starting from Shakespeare.txt and understanding what a text generation model understands at 23 million parameters. For example, nanoeuler at that number had understood that Name: started a line and wrote that line with sense.

I wrote everything in CUDA because I wanted to not use any intermediary between the model in training and inference and what it had to do. Then the use of SFT and much more, even if in small ways, were really useful to understand the various step to make an llm like a chatbot.Any feedback, help, or suggestions are absolutely welcome!

Comments

Chu4eeno · 2026-06-28 19:44:21 UTC

Very weird coding style, did you run astyle --style=python on C code?

Also, your LLM left a comment in the cuda source that it is untested, does the cuda stuff work?

vforno · 2026-06-28 20:45:56 UTC

yes yes tested on a 4070 ti 16gb everything worked without problems!

dang · 2026-06-28 21:26:31 UTC

> Very weird coding style, did you run astyle --style=python on C code?

I'm sure you mean it in a more curious way but this type of comment on a Show HN often comes across as too harshy/snarky/dismissive for what we want here (see https://news.ycombinator.com/showhn.html).

gaflo · 2026-06-29 14:08:21 UTC

Consider adding a rule that an author must disclose (in their own words) for what parts and to what extent LLMs have been used to assist their project.

dang · 2026-06-29 16:35:55 UTC

I don't think it would help: most people wouldn't know about it, of those who did many wouldn't conform, and there would be no reliable way to enforce it.

gaflo · 2026-06-30 15:45:13 UTC

You're probably right, I guess I'm too optimistic.

bArray · 2026-06-28 22:04:16 UTC

Not sure, but the code is quite dense and lacking in comments. `nanoeuler` & `nanoeuler_check` is itself the binary checked straight into git with the `.log` file? All of the commit messages are "Add files via upload" and happened in quick succession.

I suspect this is LLM generated, which is cool, but shouldn't then have the claim "forward and backward passes are written and verified by hand" unless it is true.

Regarding the data, old texts from Gutenberg probably lowers the performance - especially as many texts are on purpose whimsical. Shakespeare for example made up words to be theatrical. You have a mix of different old English styles in the corpus - it's a terrible way to learn modern English. I had some success using .ZIM data archives from Kiwix as a source, you should get a more stable output using that data.

vforno · 2026-06-28 22:09:23 UTC

Hi, the uploads are one after the other because it was a long, step-by-step research project where I tested the code on another machine. I admit that I'm slowly making up for the commits on all the projects. For Gutenberg and Shakespeare, I admit that they were the best tests I could do, but I'll always improve!

andai · 2026-06-29 04:19:55 UTC

I haven't tested NanoEuler yet but Gutenberg is awesome. Maybe a matter of taste but I like it much better than modern English.

ericb · 2026-06-28 22:02:52 UTC

How long was it trained for? How many tokens?

vforno · 2026-06-28 22:06:53 UTC

Hi, a couple of hours, not too much! Including sft!

tdesilva · 2026-06-29 00:06:06 UTC

Mentioning neural ODE doesn't make sense here, as this is unrelated. Basically any implementation of transformer uses residuals, but you're not really training a neural ODE here.

Also consider getting rid of the em-dashes. I don't know if you mostly vibe-coded this or not, but the README is pretty clearly AI generated.

vforno · 2026-06-29 05:23:12 UTC

Hi, thanks for the comment. Nanoeuler is starting as a study and research project that will obviously improve over time. I'll do my best to make the readme and other things more readable. Thank you very much.

isatty · 2026-06-29 02:19:27 UTC

I'm genuinely curious how much of this is LLM generated?

vforno · 2026-06-29 05:20:53 UTC

Most part of trasformer and sft!

ali_chherawalla · 2026-06-29 10:31:43 UTC

this is super interesting. Looking forward to trying this out!

vforno · 2026-06-29 10:52:55 UTC

Really thanks If you need any help or have any questions I'm here.

AndReics · 2026-06-29 10:56:49 UTC

Wow that's really cool i'll definitely check it out! have played around with machine learning algorithms built from scratch in c / cuda too, but once i hit the cuda part of it i kinda just left it to the side. i'm curious how did you use CUDA to optimize the matrix multiplications? how optimized is training, does it take much longer then using pytorch?

vforno · 2026-06-29 12:00:21 UTC

Hi, in nanoeuler I use cuBLAS (NVIDIA's super-optimized library) for all matrix multiplications, with the tensor cores in TF32 mode. It's the same thing PyTorch uses underneath, so it's very fast. What I've optimized (and will improve even more) and written by hand are the kernels for the other parts (like FlashAttention, which gave a nice 3x speedup), while I've delegated the large matrices to cuBLAS. Training the 116M model on a 4070 runs well and in reasonable times. Compared to PyTorch, it's a bit slower (probably 1.5-2.5x), but nothing dramatic, especially considering it's all done from scratch without a framework and there are no other optimizations that would make it faster. I'm working on it.

novaRom · 2026-06-29 13:00:43 UTC

Do you have a guess why your code is so much slower than torch? I didn't look, but there must be no reason to have 2x slower code esp. for a simple grid of FMAs.

vforno · 2026-06-29 13:11:13 UTC

Yes, because it has many separate kernels instead of aggressive merges like PyTorch (with Torch Compile). Each pass (norm, matmul, residual, RoPE, etc.) launches its own kernel, which increases launch overhead and memory traffic. CuBLAS helps, but it's not enough to compensate.

AndReics · 2026-06-29 13:30:59 UTC

i see really cool, where i failed was trying to build my own matrix operations library, it was just too much, but using cuBLAS definitely helps, i'll look into the custom kernels you wrote they seem interesting!

did you build the backprop yourself? it is a really cool project to build and i think you can agree that it teaches you a lot of how LLMS and machine learning in general works.

vforno · 2026-06-29 13:35:57 UTC

Absolutely yes! With nanoeuler I learned so much by testing every little detail of the project. Every little part you see has been tested and proven several times so that it could be understood and worked.

memezed · 2026-06-29 19:37:38 UTC

Amazing, i’ll try to learn from your code. Thanks!

vforno · 2026-06-29 20:08:26 UTC

Thanks to you! I will update the model later to make it more and more optimized but you will see everything you need in the readme.