[{"content":"I\u0026rsquo;ve been a software engineer for most of my career.\nBut with Physical AI turning a corner with general models able to one-shot new tasks, robotics seems like the next frontier. Like how I jumped into the deep end of the pool with iOS programming 16 years ago, I know the best way to get started is to get my feet wet. So here I am, training a policy and trying to build my first experiment!\nData is scarce in the robotics world. The way data is used today for most VLA models is to pretrain on generic egocentric data, and posttrain using teleoperation data and RL. This of course, is changing as we speak with In-Context Learning models.\nHowever, the underlying notion is still the same: useful data is scarce and important.\nAs I read more about egocentric videos and policy training, I learned that wristview videos provide a secondary viewpoint to help train policies better (Hsu et al., 2022). I started thinking: could we provide more useful data to policies without having to redo much work?\nScouring the internet led me to the WARPED paper. Though the paper\u0026rsquo;s not been accepted yet, it did seem very promising, and I was keen on replicating parts with a minimal setup. This would also come with a few differences: I\u0026rsquo;d be using cheaper cameras (Arducams and my iPhone), a much smaller desk, and no roboarm.\nSo the question I set out to answer: can I take video from a camera on my head, and produce video from a camera that was never on my wrist?\nLeft: egocentric camera. Right: generated wrist video.\nToday, I\u0026rsquo;ll share my process, my code, and lessons learned. Fair warning on the second half: I did eventually build an experiment to test the videos, and the answer was not the one I expected. Working out what it was actually telling me took longer than building the pipeline did.\nWhy even build this Research has shown that wrist videos provide an useful alternate viewpoint, and improves policy training by up to 30 percentage points on certain tasks (For instance to move an aluminium can, a wrist camera increases success from 43.3% to 73.3%).\nWhy are wrist videos important? Human head-mounted egocentric video doesn\u0026rsquo;t always have the right view. Think about this: there are times when the hands leave the frame to do something else, or the object is too small to be in view of the egocentric cam.\nWhich means that if we want to add wrist video, we either collect real wrist videos, or try to generate them from a room scan.\nWhy is collecting real wrist videos hard? It\u0026rsquo;s expensive. Even the cheapest Arducam cameras are roughly $45 each (aka $90 for both wrists). Multiplied by your number of field collectors, this number starts to add up fast if you\u0026rsquo;re collecting thousands or millions of hours of video.\nCould we generate synthetic wrist videos from egocentric video alone, with nothing else? This is hard, and honestly, not advisable. A head camera a meter and a half away just doesn\u0026rsquo;t capture enough of the working area up close to invent a view from 25 cm.\nWhat can we do instead then to synthesize wrist videos? We can scan the room and working area. The more data here the better. We can generate an approximate room and scene reconstruction. A room scan costs as little as 3 minutes and is only needed once.\nOf course, the drawback for this approach is that you can\u0026rsquo;t easily run this pipeline on open-source egocentric data like Egocentric-100k), since there aren\u0026rsquo;t accompanying room scans.\nProcess The equipment I used was simple: an iPhone 16 Pro with the Blackmagic Cam iOS app on my head, and an Arducam for the wrist camera.\nAfter several tries, I had to shift the wrist camera to my bicep, about 45 cm from my fingertips, pointing down my arm at my hand. The Arducam\u0026rsquo;s lens was too narrow to be useful any closer. While I got a much better viewpoint this way, this choice also cost me later. More on that below.\nI then filmed two things:\nA complete scan of the room. I spent about 4 minutes filming this scan from a variety of angles, from both far away (1.5m) and close up (15cm close.). Samples from my room scan. Lighting is slightly different here, which is another problem I had to control for.\nDemos on a simple task. In my case, the task was a simple lift. I grasped a cube, lifted it about 15 cm, held for a second, put it back, then moved the hand away. I filmed 60 takes across 6 sessions (10 takes per session), knowing that I would lose some to the pipeline. I used whistles to sync the clocks between both the wrist and ego cams. If I\u0026rsquo;d filmed each individual take separately, this would have taken exponentially more time. Rendering pipeline My pipeline runs in six stages, turning raw egocentric video + the room scan into a synthetic wrist-camera video. As a reminder, the code is open source and lives here.\n1. Ingestion. First, we read camera intrinsics off them. We sample the scan at 6 fps and the demos at 20 fps. This provided enough frames to remove duplicates. We also throw away blurry frames at this step. (code pointer)\n2. Scene reconstruction. Now we need to know what the room looks like in 3D. Running COLMAP over the scan frames for structure from motion (SfM) gives us a camera pose for every frame and a sparse point cloud of the room.\nThe reconstruction has no units of its own, so we use the 100mm x 100mm ArUco marker on the desk to set the scale. We then train a 3D Gaussian splat on those poses, and use this splat to render the scenes later. (code pointer)\nRequires GPU. Since gsplat is CUDA-only, I had to rent GPUs (4090s) from RunPod for this (shoutout to them here for how easy it was to use! It worked with Claude Code really well).\nThe Gaussian splat rendered from three points along a wrist trajectory. At this point, there\u0026rsquo;s no gripper or object drawn yet. But you can already start to see the fidelity of the reconstruction, which is pretty cool.\n3. Localize viewpoint. Currently, we know what the room looks like, however, we still do not know where I was standing in it during each demo. To tackle this, we have to pull 10 most similar scan frames for every frame within the demo, match them with SuperPoint and LightGlue, and solve the pose with PnP.\nThis means we end up with a head camera position and orientation for every frame of every take, in the same coordinates as the room. (code pointer)\n4. Estimate hand and object. With the camera placed, we can now locate the things that move. We track the hand with WiLoR, which returns 3D landmarks. (I picked WiLoR over HaMeR since I ran into some python build issues.)\nFor the object we use two models in sequence to build the reconstruction. Grounding DINO takes a text prompt and returns a rough box with a confidence score. Then SAM 2 turns that box into an exact per-pixel mask and tracks it through the rest of the clip. (code pointer)\n5. Retarget. In this step we turn the hand pose into a two-finger gripper, and build a trajectory out of it. We take the thumb-to-index distance and use it as the jaw width. We cap that at about 8.5 cm, which is roughly what a parallel gripper opens to. Then we resample onto a 15 Hz control rate, so each row of the dataset is something a policy could be asked to output. (code pointer)\nStage 4. WiLoR\u0026rsquo;s 21 hand landmarks in orange, SAM 2\u0026rsquo;s per-pixel object mask is also shown in green and tracked across the clip.\n6. Render wrist video. Finally we put the virtual wrist camera on that trajectory. We render the splat through it frame by frame, then composite the gripper and the object back in. (code pointer)\nThe finished wrist view: splat rendered through the virtual wrist camera, with the gripper and object composited in. Again, was running into some lighting issues here, but we managed to figure it out at the end.\nWhat building this process taught me I made so many mistakes building this experiment, so I hope this helps anyone who is trying to attempt something similar.\n1. Use the right camera for the job I started with the camera strapped to my wrist, but unfortunately it didn\u0026rsquo;t give the pipeline enough to work with. At that range the cube fills almost the whole frame.\nI got either too much hand, or too little hand. That left nothing for the matcher to anchor on, and nothing for the splat to draw.\nThe same task from both mounts. On the wrist, the cube fills the frame. You cannot see my hand, the desk, or anything the room scan recorded. From the bicep, the hand, the cube and the desk are all in shot which made it much easier to reconcile frames.\nTo fix this, I first looked at how other people tackle this issue. The answer lay in the hardware. Most wrist cameras in these setups are fisheye, or at least have a wide angle lens. I unfortunately didn\u0026rsquo;t have that option with the Arducam that I ordered, and I didn\u0026rsquo;t want to return a camera I had already used so much.\nInstead of widening the lens, I moved the camera back onto my bicep and let the extra distance do the same job.\nThis came back to bite me in the experiment. The real camera sat at 45 cm, while the rendered one sat at the 25 cm standoff that UMI uses.\nSo when I compared the two later on, I was partly comparing two camera positions rather than two ways of making an image. The renders weren\u0026rsquo;t useless, but this made the comparison much muddier than it needed to be.\n2. Don\u0026rsquo;t become a message bus between yourself and your AI agent. In Stage 6, when I trained a splat and rendered a few frames out of it, I only got noise.\nThe number people use here is PSNR (peak signal-to-noise ratio). It compares a rendered image against the real photo it\u0026rsquo;s supposed to reproduce, and higher is better. A working splat sits around 30 dB, while mine came out at 7.21 dB, which is roughly what you\u0026rsquo;d score by rendering television static. 😅\nSo I did the obvious thing and went hunting for the corruption. Was the file damaged on the way down from the pod? Had the training set drifted out of sync with the scan? I pointed Claude Code at each of these in turn and let it dig. This went on for days, and used a fair number of tokens.\nBut it turns out, the splat had been fine the whole time.\nThe actual problem was in the renderer, which I\u0026rsquo;d written earlier on my Mac (bad idea - don\u0026rsquo;t try to do this without a proper GPU). To draw a splat you chop the image into small tiles and work out, for each tile, which Gaussians land on it.\nMy version had a fixed ceiling on how many Gaussians one tile could hold, and when more than that showed up it quietly dropped the rest, up to 11M of them. I put the same file back on the GPU, rendered it with gsplat instead of my code, and it scored 30.46 dB. (code pointer)\nBut the renderer isn\u0026rsquo;t really the lesson here. The lesson is what I was doing while all that was happening.\nSomewhere in those days I stopped being the person making decisions and turned into a message bus, handing hypotheses to Claude Code and handing results back, without once stopping to ask whether we were even looking in the right place. It happened gradually enough that I didn\u0026rsquo;t notice it happening. Anyways, hard lesson learned. You always need a human in the loop.\n3. Put gates before the expensive step A Gaussian splat\u0026rsquo;s only built out of the frames you feed it. If you try to ask it for a view from somewhere I never stood, it tries to fills the gap in with random pixels. And sometimes, the result looks convincing if you don\u0026rsquo;t look at it properly.\nThree frames from the same clip. All three are extrapolated: the typical frame sits 15.05 cm from the nearest place the scan stood, against a limit of 15.00. The first two are obviously wrong. The one on the right has a desk, a towel, the marker and the cube in it, and is slightly more promising.\nTo prevent this, we introduced a gate to check these renders. This check walks through the frames of a rendered wrist video, finds the nearest spot the phone actually stood during the scan, and measures how far apart they are.\nIf the typical frame sits more than 15 cm away from anywhere the scan visited, I treat immediately stop the run. (This 15 cm is a threshold I chose, since most grippers are rendered at 25cm..) Past that 15cm, the renders would look bad to me, which would be enough to stop spending GPU time on them.\nHere\u0026rsquo;s the sample code:\n# Gate where I measure gap from scan floor MAX_VIEWPOINT_GAP_M = 0.15 MAX_FRACTION_BELOW_SCAN_FLOOR = 0.10 gaps = np.linalg.norm(render[:, None, :] - scan[None, :, :], axis=2).min(axis=1) scan_floor = float(_heights(scan, plane_normal, plane_offset).min()) render_heights = _heights(render, plane_normal, plane_offset) below = float((render_heights \u0026lt; scan_floor).mean()) (code pointer)\nI initially placed the check after rendering the frames, which was really expensive. Rent the pod, wait for it to provision, train the splat for 40 minutes, render frames. And then the check finally runs, right at the end, after all the GPU money is spent.\nI eventually brought a cheaper step to Stage 1 that measures whether the scan can constrain geometry. This was much better and caught most of the invalid frames way ahead of time. (code pointer)\n4. Rich texture is not necessarily useful. Structure from motion works by picking out small distinctive spots in each frame, called keypoints, and then finding the same spot again in the next frame. Two frames that share enough matched keypoints can be placed relative to each other, and that\u0026rsquo;s how you get camera positions out of a video at all.\nMy bare wooden desk didn\u0026rsquo;t give it much to work with, and the matcher kept complaining. So I went looking for a surface with more going on, and landed on.. my kitchen towel (see pic below). There\u0026rsquo;s dogs, pumpkins, a woven weave. Plenty to look at\u0026hellip; should be easy to build key points in the splat, right?\nThe kitchen towel I shot the demos on. Yes, it\u0026rsquo;s a fall-themed towel.\nWell, while keypoints (marked points to measure across frames) per frame went up 25 to 40%, which felt like a win for about ten minutes, the number of actual matches had barely moved at all.\nSince every dog on that towel looks like every other dog on that towel, the matcher would pick up a spot in one frame, go looking for it in the next, and find forty equally good candidates. Because it picked wrong a lot, about half my matches came out geometrically wrong.\nAnd so back to the bare desk I went. There were fewer keypoints to match this time, and it was still blurry in places, but at least the matcher could tell one bit of wood from another.\nWell, is the rendered video any good? Now I had the generated videos. Were they actually any good? Let\u0026rsquo;s take a look at a few of them.\nOne of the renders. Hey, not bad! We see it moving from right to left.\nOK, maybe not that great. The cube has flown away. But the splat is still pretty good.\nWith so many rendered videos, the obvious test was to hold the two side by side. I had a real wrist video and a rendered one of the same reach, so I\u0026rsquo;d measure the difference between them. Simple, and straightforward, right..?\nWell, not at all. To line those two pictures up you need to know where the real wrist camera was, to about a millimeter, and recovering that from video turns out to be a gigantic problem. I didn\u0026rsquo;t know that just yet, but my naive self went off to build this.\nRemember how I talked about shifting the wrist camera? This happens here.\nEarly demos when the camera was still attached to my wrist. Reconstruction from this angle proved much harder than I expected because well, it really wasn\u0026rsquo;t capturing anything. And thus it was even harder trying to match rendered videos with this angle.\nFirst, I wanted to ask the wrist cam \u0026ldquo;which part of the room is this?\u0026rdquo; and get a position back. In practice only 28% of those points gave me a useful position. Most of them ended up placing my camera inside a wall. 😅\nSo I tried a different route.. I knew I wanted to have the ego and wrist cameras match somehow in order to know where the wrist camera was. This time, I figured I\u0026rsquo;d give the camera its own printed ArUco marker (a different ID, of course) and wrap it around the strap.\nAttempt two: a second marker taped around the strap. Wrapped around a curve, it read in 2 frames out of 1384.\nHowever, what I hadn\u0026rsquo;t thought about is that the marker needs to be flat. Wrapping it around a curved strap made things worse instead of better. Only 2 out of 1300+ frames registered.\nI spent hours on trying to get this work before I realised I shouldn\u0026rsquo;t be doing it at all. Looking at the WARPED paper again, they never compare pixels either. They report task success and nothing else. So I stopped trying to measure the render, and started measuring what a policy does with these videos instead.\nOne caveat with this approach: the WARPED paper does have a robot arm, and at this point, I still don\u0026rsquo;t. I wanted to push the limits of what we could do without one.\nThe experiment I ended up building I had 53 demonstrations after throwing out a few bad takes. I kept 43 for training and held out ten for scoring.\nWith these, I trained three policies. Same demos, same everything, except which cameras each one got to see:\nA: head camera only B: head camera plus the rendered wrist video C: head camera plus the real wrist video How do you score a policy without a robot? You ask it to guess. Each policy looks at the current frame and predicts where the gripper goes over the next eight steps at 15 Hz, about half a second of movement. I then measure how far that guess is from what my hand actually did, in millimeters. The lower the better.\nI also ran three baselines to know what \u0026ldquo;bad\u0026rdquo; looks like. One repeats whatever the last move was. One always guesses the average move. One is a policy with no camera at all, just the hand position.\nHow many times do you train each one? Five. I learned this the hard way. If you train a policy once, you get one number, and you have no idea how much of that number is the policy and how much is the random seed. So each of the three policies was trained five times with five different seeds, for 10,000 steps each, with a learning rate that decays to zero. I did fifteen runs, five rented 4090s, about four hours, $14.47.\nI also wrote the decision rule down before the first pod started, so I couldn\u0026rsquo;t talk myself into anything afterwards. A difference between two policies only counts if its mean is bigger than twice its standard error, and the sign agrees on at least four of the five seeds.\nResults Let\u0026rsquo;s take a look.\nPolicy s1 s2 s3 s4 s5 Mean sd Repeat my last move 3.72 mm Always guess the average move 6.14 mm No camera, just hand position 6.63 mm A Head camera 3.64 3.74 3.65 3.78 3.81 3.72 mm 0.08 B Head + rendered wrist 3.90 3.86 3.86 3.80 3.75 3.84 mm 0.06 C Head + real wrist 3.77 3.68 3.72 3.70 3.61 3.70 mm 0.06 Now, the three policies of one seed share the same data order and the same GPU, so the fair comparison is seed by seed:\nQuestion Difference Mean Same sign Verdict Does the real wrist camera help? A − C +0.03 mm 3 of 5 not established Does the rendered wrist hurt? A − B −0.11 mm 4 of 5 not established Real wrist against rendered wrist C − B −0.14 mm 5 of 5 established So, does the real wrist camera help? No. A minus C is 0.03 mm, and which one wins flips from seed to seed. The seed-to-seed spread is only 0.06 to 0.08 mm, so this isn\u0026rsquo;t an effect hiding in noise, because there\u0026rsquo;s no effect to hide.\nThat also means the number I originally set out to measure doesn\u0026rsquo;t exist. I\u0026rsquo;d planned to report how much of the real camera\u0026rsquo;s benefit the render recovered:\nrecovery = (A error - B error) / (A error - C error) Thus the fraction has nothing to divide by. The real camera never beat the head camera, so there was no gap for the render to close. I\u0026rsquo;m not reporting a recovery number, because there isn\u0026rsquo;t one.\nDoes the rendered wrist hurt? Probably a little, but not enough for me to say so. Four of the five seeds (not all) say yes, and the average is 0.11 mm worse.\nIs the real wrist better than the rendered one? Yes, by 0.14 mm, on every single seed. This is the only established result in the whole experiment. However, before you get excited, remember that neither of them beats the head camera alone.\nWhy is the render worse? I don\u0026rsquo;t know for sure. My best guess is colour, not geometry. After normalisation, the rendered wrist frames sit much further from the ImageNet statistics my encoder expects than the real wrist frames do, and that was flagged before the runs started. A camera placement problem would look different from this.\nWhat about the top row? Every policy ends up level with \u0026ldquo;repeat my last move\u0026rdquo;. Head camera 3.72, real wrist 3.70, the dumb rule 3.72. The real-wrist policy edges below it on four of five seeds, but only just. My take is that lifting a cube on an empty desk is a very repeatable movement, so a rule that copies the last step is hard to beat. Pouring a cup of coffee or wiping a plate would probably be a different story.\nOne thing I didn\u0026rsquo;t vary above is the encoder, and that\u0026rsquo;s because I\u0026rsquo;d already tested it. Swapping the head-camera encoder from random weights to R3M, which is pretrained on egocentric video, cut error by about 1 mm in an earlier run, from 6.56 to 5.61 mm. That\u0026rsquo;s ten times the spread between any of the cameras, so every policy above uses R3M. I would definitely try a different encoder next time as well (see lesson #5 below.)\nIn total, training cost me about $18 and 24 hours of rented GPU time.\nWhat I\u0026rsquo;d do differently next time After two weeks of building, I know how to transform egocentric to wrist views, but there\u0026rsquo;s much more work to be done to translate this to be useful for a policy. Here\u0026rsquo;s what I\u0026rsquo;d do differently in the future:\nMeasure task success instead of millimeters. A robot that picks up the cube 8 times out of 10 is a result anyone can read. A policy that\u0026rsquo;s 0.14 mm better at guessing the next half second is not. WARPED only reports task success, and now I understand why.\nPick a task that actually needs a wrist view. Lifting a cube on an empty desk can be solved from the head camera alone, and the numbers say so. If I want to detect a wrist-view effect, I need a task where losing the wrist view actually hurts. Something with occlusion, like reaching into a drawer.\nRun several seeds and train to convergence before comparing anything. With one seed you can\u0026rsquo;t tell an effect from luck. Five seeds and a cosine learning rate got my noise floor down to 0.07 mm, which is small enough to trust a 0.1 mm difference. Get that floor first, then compare cameras.\nFix the camera geometry before running anything. My real camera sat on my bicep at 45 cm because my Arducam wasn\u0026rsquo;t wide angle. The rendered one sat at the standard 25 cm. So B against C was partly comparing two camera positions, not two ways of making an image. A fisheye lens at the right distance would remove that entirely. Something like this would probably work.\nTry other pretrained encoders. R3M produced the largest change I saw short of training longer, so I\u0026rsquo;d push on it. VIP is the obvious next one: a ResNet50 trained on the same Ego4D footage with a completely different training objective, so it separates \u0026ldquo;egocentric video helps\u0026rdquo; from \u0026ldquo;R3M\u0026rsquo;s specific method helps\u0026rdquo;. ImageNet ResNet50 as the control tells me whether I\u0026rsquo;m measuring pretraining or just capacity. CLIP ViT-B/16 is what WARPED actually used, and that\u0026rsquo;s a transformer, not a ResNet, so it won\u0026rsquo;t drop into my harness at all without custom encoder code.\nAugment the rendered demos, since that\u0026rsquo;s the whole point of rendering. WARPED turns 30 demos into 300 by re-rendering each one with the object moved, retextured, and the camera perturbed. I rendered each demo exactly once since I didn\u0026rsquo;t want to spend more time than I had already.\nA different policy (e.g. pretrained VLAs) is perhaps worth a look. I used a diffusion policy because WARPED did. ACT is cheaper to train and might separate the arms differently. Fine-tuning a pretrained VLA instead of training a policy from scratch is the direction the field is actually going, and Ego-Pi fine-tunes one on egocentric human data directly. Some options here are SmolVLA, which already lives inside lerobot so it would drop into my harness with the least work, and then also π0 through openpi, and OpenVLA.\nAll in all this was a really fun experiment. I learned a ton: WiLoR, Grounding DINO, Gaussian splat training, and a lot about where checks belong. I hope this helps shed some light on turning egocentric into wrist videos, and hopefully you don\u0026rsquo;t make the same mistakes I did.\nThank you to Harry Freeman, author of the WARPED paper for answering a lot of my noob questions.\nReferences Freeman, H. et al. WARPED: Wrist-Aligned Rendering for Robot Policy Learning from Egocentric Human Demonstrations. arXiv:2604.10809. arXiv · author\u0026rsquo;s PDF\nHsu, K., Kim, M. J., Rafailov, R., Wu, J. and Finn, C. Vision-Based Manipulators Need to Also See from Their Hands. ICLR 2022. arXiv:2203.12677. arXiv. Why a wrist view is worth having at all.\nKim, M. J., Wu, J. and Finn, C. Giving Robots a Hand: Learning Generalizable Manipulation with Eye-in-Hand Human Video Demonstrations. arXiv:2307.05959. arXiv. A real camera on a human forearm, the route I did not take.\nQian, Z. et al. WristWorld: Generating Wrist-Views via 4D World Models for Robotic Manipulation. arXiv:2510.07313. arXiv. Generated wrist views from third-person robot video.\nKim, J. W. et al. Ego-Pi: VLA Fine-Tuning for Ego-Centric Human and Robot Data. arXiv:2606.08107. arXiv.\nMandlekar, A. et al. What Matters in Learning from Offline Human Demonstrations for Robot Manipulation. CoRL 2021. arXiv:2108.03298. arXiv · robomimic study. The source of the wrist-camera success-rate numbers I quote.\nSchönberger, J. L. and Frahm, J.-M. Structure-from-Motion Revisited. CVPR 2016. paper · COLMAP\nKerbl, B., Kopanas, G., Leimkühler, T. and Drettakis, G. 3D Gaussian Splatting for Real-Time Radiance Field Rendering. SIGGRAPH 2023. arXiv:2308.04079. arXiv · project site\nGarrido-Jurado, S. et al. Automatic generation and detection of highly reliable fiducial markers under occlusion. Pattern Recognition, 2014. doi · OpenCV tutorial. The ArUco markers.\nDeTone, D., Malisiewicz, T. and Rabinovich, A. SuperPoint: Self-Supervised Interest Point Detection and Description. CVPRW 2018. arXiv:1712.07629. arXiv · code\nLindenberger, P., Sarlin, P.-E. and Pollefeys, M. LightGlue: Local Feature Matching at Light Speed. ICCV 2023. arXiv:2306.13643. arXiv · code\nPotamias, R. A. et al. WiLoR: End-to-end 3D Hand Localization and Reconstruction in-the-wild. arXiv:2409.12259. arXiv · code\nPavlakos, G. et al. Reconstructing Hands in 3D with Transformers. CVPR 2024. arXiv:2312.05251. arXiv · code. HaMeR, which I tried first.\nLiu, S. et al. Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection. arXiv:2303.05499. arXiv · code\nRavi, N. et al. SAM 2: Segment Anything in Images and Videos. arXiv:2408.00714. arXiv · code\nYang, L. et al. Depth Anything V2. arXiv:2406.09414. arXiv · code\nNair, S., Rajeswaran, A., Kumar, V., Finn, C. and Gupta, A. R3M: A Universal Visual Representation for Robot Manipulation. CoRL 2022. arXiv:2203.12601. arXiv · code. The pretrained encoder that produced the largest change I measured.\nMa, Y. J. et al. VIP: Towards Universal Visual Reward and Representation via Value-Implicit Pre-Training. ICLR 2023. arXiv:2210.00030. arXiv · code. The encoder I would try next.\nChi, C. et al. Diffusion Policy: Visuomotor Policy Learning via Action Diffusion. RSS 2023. arXiv:2303.04137. arXiv · project site. The policy I trained.\nChi, C. et al. Universal Manipulation Interface: In-The-Wild Robot Teaching Without In-The-Wild Robots. RSS 2024. arXiv:2402.10329. arXiv · project site. Where the 0.25 m camera standoff comes from.\nLeRobot, Hugging Face. The dataset format and the policy training.\ngsplat, Nerfstudio. The CUDA splat trainer and rasteriser.\n","permalink":"https://sidwyn.com/posts/ego2wrist/","summary":"\u003cp\u003eI\u0026rsquo;ve been a software engineer for most of my career.\u003c/p\u003e\n\u003cp\u003eBut with Physical AI \u003ca href=\"https://itcanthink.substack.com/p/in-context-learning-results-hint\"\u003eturning a corner\u003c/a\u003e with general models able to one-shot new tasks, robotics seems like the next frontier. Like how I jumped into the deep end of the pool with iOS programming 16 years ago, I know the best way to get started is to get my feet wet. So here I am, training a policy and trying to build my first experiment!\u003c/p\u003e","title":"ego2wrist: Turning egocentric videos into wrist videos."},{"content":" This was initially published on Path to Staff, but I bring it here as I write more about learning AI.\nWelcome back to Path to Staff! This is the final part of the Unpacking AI series, where we cover Training, Post-Training and Inference, and watch it all come together.\nI also crosspost this to my personal Substack where you can subscribe if you are interested in more technical posts like these. After this piece, we will return to regular programming on career growth on Path to Staff.\nHere’s the series thus far:\nThe Hardware Behind AI. Where we talked about hardware, including transistors, fabricators, and most famously the memory-compute bottleneck (why RAM is so expensive!). Model Architecture. Where we covered the “Attention is All You Need” paper, talked about tokens, embeddings, attention, and covered the transformer block. This is a must-pre-read if you haven’t read it. You’ll need to understand this in order for this piece to make sense. Training, Post-Training, and Inference. (This essay.) In this piece, we dig into how the numbers inside a model get set, refined and finally return an answer to you. I originally sketched this as five parts, but as I’ve gone through this, I’ve realized it makes more sense to combine the last three into one post. So here we go.\nA model is just a pile of numbers When we look at a model, say ChatGPT 5.6, industry estimates put frontier models like it in the trillions of parameters. So for this article, let’s work with a model that has a trillion parameters. These trillion random weights sit in memory and are all randomly initialized at the start.\nNow when I say weights, I also include both weights and biases. These trillion weights and biases control connection strengths and trigger neuron thresholds respectively.\nWhen ChatGPT 5.6 is first created with these randomized numbers, if you ask it anything, it is going to return pure gibberish to you. That’s because all of these numbers were randomly initialized.\nNow, the life of these numbers has three phases:\nPretraining: where we set the weights and biases. Post-training: where we refine the weights, adjust biases. Inference: where we read the weights and biases to answer a user’s prompt. We’re going to cover each of them in detail. Strap in!\nPretraining What is pretraining? Pretraining is how these weights go from random values to their first real values.\nUsing thousands of GPUs over months, the model reads the entire Internet corpus. As it reads this data, it blindfolds itself by hiding the next word. It then asks itself, “Hey, what’s the next word?” And based on whether it gets it right or wrong, it adjusts its weights.\nThis happens over and over again until its loss function plateaus, or the training budget runs out.\nBy the end of this pretraining, the model learns human language. A variety of concepts compress into a fixed number of weights, which include:\nLanguage: grammar, vocabulary, sentence structure World knowledge: what’s the capital of X? Relationships: how do ideas and themes connect? Why do we need pretraining? Before LLMs were popular, AI models were supervised. This means that humans labeled examples by hand. A human would label a review as positive and hand it over to the machines to learn. As noble as this work was, you could never label enough examples to cover all of human language. It was impossible.\nInstead, researchers took another spin on this. Pretraining is self-supervised. This means that the text it reads is also its own answer key.\nWhen the model hides the next word for itself to guess, no humans are required. And best of all, now the entire internet is training data.\nThe pretraining loop Here’s how pretraining works in detail.\nInitialize random weights. Forward pass: the model predicts the next token in the sequence by calculating self-attention. Grade the prediction: using cross-entropy loss, produce a grade by taking the negative log of the probability the model got it correct. Backpropagate: figure out which of the 1 trillion weights to nudge and in which direction. Defining cross-entropy loss What is cross-entropy loss? Let’s go back to the example “The cat sat on the _”.\nAs the model guesses the next word, it assigns a confidence probability to every word in its vocabulary. The loss looks only at the confidence it gave the correct word.\nThis is cross-entropy loss because it measures how different the model’s guessed probabilities are from the absolute certainty of the true answer.\nIf the model was 90% sure the next word was “mat”, the penalty is small, say maybe 0.1. And vice versa: if it gave the word “mat” a 1% probability, the penalty is large. This penalty is then handed over to backpropagation, which refines the weights for the next round.\nHow backpropagation works Now for the fun part. We have calculated the cross-entropy loss and have 1 trillion weights to tune. Which weights do we tune?\nThis is what backpropagation is about. It walks backward through the network, using the calculus chain rule to assign each weight a gradient, which measures how much that specific weight contributed to the error.\nIt then sweeps backwards layer by layer.\nThis sweep costs about twice the compute of the forward pass, because every layer has to do two matrix multiplications (matmuls).\nGradients for its own weights. This tells the layer how it should change. It takes just one matmul, where the incoming gradient is multiplied by the activations that we saved earlier during the forward pass. Gradients to pass downstream. This tells the layer how much blame to forward to the layer below it. We need a second matmul here, multiplying the incoming gradient by the transpose of the layer’s own weights. As it moves backward through the hidden layers, backpropagation works out how a tweak ripples all the way through to the final error.\nUpdating weights with gradient descent With backpropagation, every weight now knows its share of the blame.\nTo then update the weights themselves, we need to utilize an optimization algorithm called gradient descent. This is where we take a small step in the direction the gradients point.\nHow much of a step do we take? The learning rate tells us so. This is a hyperparameter (aka tunable by the ML engineer), and is one of the most important dials in all of training.\nModern training runs don’t keep learning rates fixed either. They follow a schedule, usually a cosine decay curve that starts small, peaks, then comes back down to around 10% of its maximum by the end of the run.\nWhy this rollercoaster schedule, you might ask? Because the model’s needs change drastically over the run.\nAt the start, the freshly randomized weights produce chaotic gradients, which means big steps this early would destabilize the entire network. So we warm up slowly.\nOnce things stabilize, the learning rate peaks and the model takes massive strides across the loss landscape, rapidly picking up broad concepts. Then near the end, a high learning rate would just violently overshoot the target, so we shrink the steps down and let the weights settle into the lowest point of error.\nThe Chinchilla Paper Now that we know how to modify these weights with backpropagation and gradient descent, let’s go back to the data the model first ingests. How much data exactly is needed to train a model?\nIn 2020, OpenAI’s Kaplan paper proposed the first scaling laws, which said that as compute grows, you should make the model vastly larger while only slightly increasing the data it sees.\nIn other words, pile on as many parameters as you can. The industry started rushing towards enormous models trained on relatively tiny datasets. That’s why you see models like OpenAI’s original GPT-3 (175 billion parameters) or Meta’s OPT-175B trained on just 300 billion tokens. Even Google’s massive PaLM at 540B parameters was trained on only 780 billion tokens. By Kaplan’s math, these were supposedly correct, since the data should not increase as much as the parameters do.\nHowever, things changed two years later. In 2022, DeepMind suspected these gigantic models were severely undertrained, so they ran the experiment properly this time, building 400 different baseline models across a wide range of sizes and data pools. Most importantly, they built their own 70B model, Chinchilla.\nChinchilla soundly beat much larger models like the 280B Gopher (DeepMind’s own model) and the 540B PaLM, with just a fraction of the parameters.\nTheir refined math gave the industry a new rule of thumb:\nThe Chinchilla-Optimal Rule says that for compute-optimal training, you should feed the model roughly 20 tokens for every 1 parameter (e.g., a 70B model needs about 1.4 trillion tokens).\nPost-training So we’ve spent months and (likely) a small fortune pretraining a model. Now what do we have? An elite next-word predictor.\nAsk it “What’s the capital of France?” and it might reply with “What’s the capital of Germany? What’s the capital of Italy?”\nNow that’s a perfectly reasonable continuation if all you’ve ever done is read the internet, but pretty useless if you actually wanted an answer. A base model that’s pretrained will happily autocomplete your prompt or ramble on endlessly.\nPost-training is where it learns to become an assistant that follows instructions, tells good answers from bad ones, and aligns with human preferences.\nA standard post-training pipeline has two phases, Supervised Fine-Tuning (SFT) and Reinforcement Learning from Human Feedback (RLHF). There’s also an emerging third act which we’ll get to shortly.\nPhase 1: Supervised fine-tuning (SFT) SFT is teaching by example. And it is an important first step before we can move on to other RL methods (like RLHF below).\nWe take the raw base model and fine-tune it on a bunch of input/output pairs. These pairs teach it a specific behavior, tone or task.\nFor instance, you can teach it to be a customer service assistant. Or to be a calculator.\nThe loss calculation mechanics in SFT should look familiar, because it’s exactly the same next-token prediction loss as pretraining: using cross-entropy loss. However this time, the loss is only calculated on the response, and not the prompt. This is because we don’t want it to memorize the prompts (or the questions themselves).\nPhase 2: Reinforcement learning from human feedback (RLHF) Now, SFT teaches the model formatting and conversational behavior. However, the model still struggles with fuzzier ideals like truthfulness or being helpful.\nThat’s where RLHF comes in.\nInstead of being fed one perfect answer, the model generates several candidate responses to a prompt. Human annotators then rank these outputs from best to worst.\nAs a ChatGPT user, you’ve probably done this labeling yourself without even realizing it.\nScreenshot of an older version of ChatGPT, taken from Reddit.\nUsing reinforcement learning, the model receives a mathematical reward for high-ranking answers and a penalty for the bad ones. Of course, this is also one of the most expensive steps since it requires humans to tune the models. But now, after we’ve covered all the major themes, all 1 trillion weights are actively optimizing for human preference, tone and safety.\nThe RLHF loop where the model writes, humans rank, and then weights get updated.\nWithin RLHF, there are a few techniques. In a nutshell these are:\nRejection sampling: generate several responses, then keep only the winning samples and fine-tune on those.\nProximal policy optimization: introduce a penalty (known as the Kullback-Leibler penalty) which is multiplied by a coefficient and subtracted from the reward when tuning the new model. This keeps the new model from drifting too far away from the original.\nDirect preference optimization: directly optimize the language model using the chosen and rejected responses, without training any reward model at all. Traditional RLHF trains a reward model and then runs PPO on top of it, while DPO goes straight from the preference data to the aligned model.\nThe next frontier: RL on verifiable rewards Now, the classic version of RLHF has a massive bottleneck.\nHuman preference is subjective, expensive to scale, and often rewards text that merely sounds plausible over text that is actually correct.\nSo the industry is shifting toward verifiable rewards, which are problems where the ground truth can be checked by a computer program instead of a human eye.\nFor example, instead of a human reading an essay, a sandbox environment tests whether the model’s generated code compiles and passes all unit tests, or whether its step-by-step math lands on the correct final equation. Once you replace human judges with code executors and math verifiers, models can self-correct, “think” longer via reasoning loops, and scale their capabilities far beyond the limits of human annotation.\nMeasuring model helpfulness and safety using evals As models progress through pretraining and post-training, we need objective benchmarks to measure how capable they actually are. You have likely seen the scorecard tables in release notes from frontier AI companies.\nThese are standardized lists of verified benchmarks, or “evals”, and they roughly fall into these buckets:\nKnowledge \u0026amp; Reasoning: everything from broad general knowledge across dozens of subjects to graduate-level logic puzzles designed by PhDs (MMLU / GPQA). Math \u0026amp; Logic: ranges from multi-step grade school word problems to competition-level algebra and calculus (GSM8K / MATH). Coding \u0026amp; Engineering: measures whether the model can write code that passes unit tests, and navigate complex multi-file repositories to fix real bugs (HumanEval / SWE-bench). Safety \u0026amp; Alignment: how reliably the model catches and refuses dangerous requests, such as generating hate speech, planning cyberattacks, or building weapons (Red-Teaming / Refusal Evals). Companies are creating new evals constantly, and the big labs even have entire eval divisions now. When an eval saturates (models start hitting 95%+), it gets hardened or updated.\nTo keep things fair, it’s also important that benchmarks don’t leak into training sets, which is why the trajectory is moving towards private, frequently updated tasks.\nInference Alright, we’re done with pretraining and post-training. It’s time to move on to the last part of this article. If you’ve made it this far, congrats!\nWe’ve finished training the weights, so now how do we turn them into answering machines?\nEnter inference. Inference happens when you type a prompt and the model “infers” it to return an answer. There are two major phases:\nPrefill: processes the entire prompt in parallel, performing a matmul across all the tokens and storing the Keys and Values in the KV cache. We covered this in the previous essay. This phase is compute bound (buy the best GPUs!) Decode: generates the response one token at a time, and every single token requires reloading the model weights from HBM. This phase is memory bound. (Remember the memory wall from Chapter 1?) Now since there are two phases with two different bottlenecks, this means there are two obvious metrics to track.\nPrefill is measured by TTFT (time to first token), while decode is measured by TPOT (time per output token). Sites like Artificial Analysis track these numbers across providers.\nOptimizing Inference Let’s see how to optimize inference. We’ll talk about two key technologies, batching and PagedAttention. As software engineers, we are already used to these concepts in system design!\nBatching Hauling 140GB of weights across the memory bus just to emit one token for a single prompt is extremely inefficient. What’s in play right now is continuous batching, where a scheduler processes all the active sequences together. As soon as a sequence emits an END token, a new queued request fills that slot. Check out BentoML’s animation if you want to see how this works.\nIn practice, OpenAI and Anthropic both have Batch APIs (OpenAI’s, Anthropic’s) that do offline batch inference, which combine latency-insensitive jobs. This reduces cost for both themselves and their customers.\nPagedAttention As the industry converged on continuous batching, a new problem popped up.\nEach request’s KV cache grows like crazy, and at admission time we have no idea whether a response will cost us 10 tokens or 10,000. As such, a lot of KV memory ended up fragmented and over-reserved.\nRemember defragmentation in your PCs? Well, the same solution can also be applied to LLMs.\nOperating systems solved this exact problem decades ago when they invented virtual memory. They chopped application data into fixed-size “pages” and mapped them to scattered pieces of physical memory (called “frames”) using a centralized page table.\nThe vLLM team looked at the KV cache and realized it behaved exactly like a running program’s memory footprint. It’s dynamic, unpredictable, and rapidly growing.\nSo instead of treating the KV cache as a giant rigid block that has to sit together in VRAM, they built a virtual memory manager for the GPU, where:\nTokens became bytes: fixed chunks of tokens (usually 16 per block) get grouped into pages. The KV cache became RAM: instead of allocating memory upfront based on a guess of the maximum response length, the system hands out a new 16-token physical block only when the model has actually generated enough text to fill the current one. The lookup table: just like an OS page table, a logical manager maps a user’s sequence to these scattered physical blocks across the GPU’s memory. Further inference optimization There’s a lot more to cover in inference optimization, and the BentoML inference handbook I linked actually does a really good job here.\nThree optimizations you should know are prefix caching, speculative decoding (having a small draft model race ahead), and prefill-decode disaggregation.\nPrefix caching. Requests share prompt prefixes constantly, whether it’s the system prompt, few-shot examples, or earlier turns of a chat. So we cache their KV pages once and reuse them, skipping that slice of prefill entirely. This is also why your API bill has cache-read line items at a fraction of the input price. Speculative decoding. A small draft model races ahead and proposes several tokens, then the big model verifies the whole run in one parallel pass and keeps the longest correct streak. Since verification preserves the exact output distribution, you get more or less identical quality at lower latency. An example is using Gemma-2-2B as a proposer and Gemma-2-9B as a verifier. Prefill-decode disaggregation. Remember how prefill is compute-hungry while decode is bandwidth-hungry? If you mix them on one GPU, a whale of a prompt will stall everyone else’s token stream during its prefill. The fix is simple here: run the two phases on separate pools that are sized independently, while handing off the KV cache between them. Hardware is now being designed around this split, like NVIDIA’s Rubin CPX, a prefill-specialized chip. Inference chip startups are also evolving like crazy in this space. Look at Groq, Etched or MatX. These companies are raising a lot of money to take on the giants (NVIDIA and Google) with their own inference chips. Definitely an interesting space to watch.\nTLDR Let’s wrap this up. Hopefully this article made sense. If not, feel free to leave your comments below.\nHere are 8 things to remember from today:\nA model starts as a pile of random numbers. Weights and biases, randomly initialized, that return pure gibberish until they’re trained. Pretraining sets the weights. The model guesses the next word across the entire internet, grades itself with cross-entropy loss, and repeats until the loss plateaus or the budget runs out. Backpropagation is an error assignment scheme. One backward sweep tells all 1 trillion weights how much they contributed to the error, at about twice the compute of the forward pass. Chinchilla says 20 tokens per parameter. Before this rule of thumb, the industry was training gigantic models that were severely undertrained. Post-training turns a predictor into an assistant. SFT teaches by example, RLHF optimizes for human preference, and RL on verifiable rewards is the next frontier. Evals keep score. Standardized benchmarks for knowledge, math, coding and safety, with the trajectory moving towards private, frequently updated tasks. Inference has two phases with two bottlenecks. Prefill is compute bound and measured by TTFT, while decode is memory bound and measured by TPOT. Inference optimization is a systems game. Continuous batching, PagedAttention, prefix caching and speculative decoding all squeeze more tokens out of the same GPUs. Thanks for following along on this Unpacking AI series!\nIf you like this sort of technical content, subscribe to my other Substack here. For Path to Staff, we’ll be returning to regular programming on career growth. Stay tuned.\n","permalink":"https://sidwyn.com/posts/random-numbers-learns-to-talk/","summary":"\u003cblockquote\u003e\n\u003cp\u003e\u003cem\u003eThis was initially published on \u003ca href=\"https://www.pathtostaff.com/p/how-a-pile-of-random-numbers-learns\"\u003ePath to Staff\u003c/a\u003e, but I bring it here as I write more about learning AI.\u003c/em\u003e\u003c/p\u003e\u003c/blockquote\u003e\n\u003cp\u003eWelcome back to Path to Staff! This is the final part of the Unpacking AI series, where we cover Training, Post-Training and Inference, and watch it all come together.\u003c/p\u003e\n\u003cp\u003eI also crosspost this to \u003ca href=\"https://sidwyn.com/\"\u003emy personal Substack\u003c/a\u003e where you can subscribe if you are interested in more technical posts like these. After this piece, we will return to regular programming on career growth on Path to Staff.\u003c/p\u003e","title":"How a Pile of Random Numbers Learns to Talk"},{"content":" This was initially published on Path to Staff, but I\u0026rsquo;m bringing it here as I start to write more about learning AI.\nWelcome back to Path to Staff! This series is a little different from our usual programming. In this series, we\u0026rsquo;re covering LLMs and AI in-depth.\nAs an engineer, I never really had the time to understand AI\u0026rsquo;s internals. But I\u0026rsquo;ve spent the past few weeks doing deep research to unpack it all.\nAs a reminder, this is Part Two of a five-part series:\nThe Hardware Behind AI – And How It\u0026rsquo;s Programmed. Transistors, semiconductors, and fabricators. Learn about the big players (TSMC, Nvidia, ASML). The memory-compute bottleneck. And all the acronyms you always wondered about (TPU, ASIC, FPGA, CUDA, etc.) Model Architecture. (We are here.) Learn about what models are made of. We\u0026rsquo;ll cover the paper that started it all (\u0026ldquo;Attention is All You Need\u0026rdquo;), plus talk about transformers and diffusion models. Training. The meat of teaching a model. How does pretraining work? What goes into it (backpropagation, optimizers, loss functions)? What scaling laws should we weigh before we kick off an expensive training run (up to hundreds of millions of dollars)? Post-Training \u0026amp; Alignment. How does one guide a model once it\u0026rsquo;s been taught? How do we apply safety? How do we benchmark and know the model got better? How do we evaluate a model\u0026rsquo;s performance? Inference, Serving and Agents. This might be the most familiar topic, since it\u0026rsquo;s closest to you as an AI user. How does a model output its token and serve the result to you (SSE)? How do systems stay fair and fast? What tools are available (MCP, RAG, tool use) and how do agents work? Recurrent Neural Networks are No Longer Popular I remember taking CS:188 by Pieter Abbeel at Berkeley, learning about Recurrent Neural Networks (RNNs) in 2014. I wish I\u0026rsquo;d paid more attention in class, but it was still a good foundation for working on this chapter.\nTo my surprise, RNNs are less popular today. And there\u0026rsquo;s a good reason for that.\nIf you haven\u0026rsquo;t studied neural networks before, you should know that a neural network looks at an input, guesses what it is, then immediately forgets it. The simplest form is a feed-forward neural net, where the data goes through and outputs as a layer at the end.\nRecurrent neural networks take this one step further, with a built-in memory loop. It looks at the first word, jots it down in a hidden notebook (or layer), and then reads the second and jots it down again. This repeats over and over again. The downside is that it has terrible long-term memory, and would only remember what was most recently fed.\nAn RNN reads \u0026ldquo;The cat sat on the mat\u0026rdquo; one word at a time - its memory of \u0026ldquo;The\u0026rdquo; has almost faded by \u0026ldquo;mat\u0026rdquo;.\nUpgraded versions of RNNs called LSTM (Long Short-Term Memory) and GRUs (Gated Recurrent Units) were implemented to highlight important memories, but there were still two key problems with these recurrent neural networks.\nThey were:\nSequential bottleneck: You can\u0026rsquo;t compute the next step in a recurrent neural network till the current operation is complete. This means there\u0026rsquo;s no way to parallelize training. A sequence of length N takes N sequential steps no matter how much hardware you throw at it. Training time was bound by the length of the chain, not the size of the chain. Long-range decay: information from early tokens fades before it reaches later ones, as seen in the image above. Training an RNN means backpropagating the signal through every step, multiplying by the same weights each time. So over a long sequence the gradient vanishes to zero or explodes, and the model can\u0026rsquo;t learn that an early token should shape a much later one. Enter Transformers In 2014, Bahdanau et al. added a different concept to neural networks, called attention. Rather than force the whole input through one fixed summary, let the model look back at every input token and weight the ones that matter for the current step. Simplified, when generating a new token, the model looks back at all the previous tokens to generate a vector for each of them.\nIt worked well, and took an important step toward the transformer. However, it was still bolted onto a recurrent network, so the sequential bottleneck remained.\nThree years later, Vaswani et al. made a large leap: keep the attention, drop the recurrence entirely. Their insight was that self-attention, where every token computes its relationship to every other token through dot products, solves both problems (sequential bottleneck \u0026amp; long-range decay) at once.\nWith matrix multiplication happening over all pairs simultaneously, the whole sequence can now be processed in parallel rather than step by step. And with no recurrent chain, there\u0026rsquo;s no information left to decay.\nThe same five tokens, two ways to mix them. The RNN passes information hand-to-hand down a chain, so it runs one step at a time and the earliest signal fades by the end. Self-attention connects every token to every other at once — fully parallel, with nothing left to decay.\nEight Google researchers laid this out in a paper titled \u0026ldquo;Attention Is All You Need\u0026rdquo;. The architecture they proposed within is called the transformer.\nAt first, they built it to improve machine translation, but they already saw further, noting in the paper that the same architecture should extend to other tasks. They were not wrong.\nAn Overview of Transformers A transformer is a neural network that processes sequences.\nAn input goes in, passes through N identical blocks and a prediction head converts it into an output.\nA prediction head is a translator at the end of the model that turns the model\u0026rsquo;s numbers into a score for every word. These raw scores are called logits.\nThere are different types of transformers. Let\u0026rsquo;s first look at an example of translating a sentence. You have an input that gets read, which gets translated to an output.\nSay you want to translate \u0026ldquo;The cat sat on the mat\u0026rdquo; to Chinese. This goes from \u0026ldquo;The cat sat on the mat\u0026rdquo; to \u0026ldquo;猫坐在垫子上\u0026rdquo;.\nIn this case, we are using an encoder-decoder architecture. The encoder reads the input, while the decoder generates output.\nHere\u0026rsquo;s what the architecture looks like: Now remember that the authors were solving machine translation. However, the AI field took this architecture apart and found that the halves were independently useful! Google used the encoder, trained it to understand text, and got BERT (Bidirectional Encoder Representations from Transformers), released in 2018.\nBy 2019, it was rolled into the ranking backend of Google Search. It was able to understand each word in context with all other words simultaneously. It was incredibly good at understanding context, and so most researchers would use it for sentiment analysis, or for named entity recognition.\nGoogle also worked on decoders, but it worked towards Generating Wikipedia by Summarizing Long Sequences. Then Alec Radford at OpenAI took the same move, but in a different, important direction: combining the decoder-only transformer with his generative-pretraining thesis.\nIn Improving Language Understanding by Generative Pre-Training, Alec et al. argue that raw text is everywhere, but text that\u0026rsquo;s been labeled for a specific task (like \u0026ldquo;this review is positive\u0026rdquo;) is rare and expensive, which limits models that can only learn from labeled examples.\nTo fix this, let a model learn language broadly by reading a huge pile of unlabeled text, then give it a small amount of labeled data to specialize on a particular task. That two-step recipe produces big improvements.\nEnter GPT These two steps are known as:\nGenerative pre-training (GPT, anyone?) - read unlabeled text Fine-tuning - specialize on a particular text This generative pre-training is the first step of an AI model, where it is trained on a massive general dataset to learn general-purpose representations and patterns (think grammar, language, reasoning).\nAlec, together with Karthik Narasimhan, Tim Salimans and Ilya Sutskever, came up with GPT-1 in June 2018, a decoder-only transformer.\nWhile Google was doubling down on BERT, OpenAI doubled down on Generative Pre-Training (GPT), pushing out GPT-2 in 2019 with the same architecture but roughly 10x the parameters and data. Now, the model could start doing tasks zero-shot from prompting, without any fine-tuning. GPT-3 in 2020 scaled another 100x.\nWhat happened to the fine-tuning step? By GPT-3, the model showed that it mostly didn\u0026rsquo;t need it. A big enough model was smart enough to understand most tasks from a prompt alone. More work was put into post-training, which we will cover in a future chapter.\nEncoder Only vs Encoder-Decoder Models Now, it is important to understand why GPTs (decoder-only transformers) became much more popular than say the encoder-decoder or the encoder model.\nThree reasons:\nA decoder trains on predicting the next token. This requires no labels and task-specific setup. This meant that the entire internet corpus becomes training data. The objective was universal. You could use it for translation (e.g. \u0026ldquo;continue this text\u0026rdquo;), summarization (\u0026ldquo;summarize this text\u0026rdquo;), etc. All of these reduced to continuation. On the other hand, BERT could never be prompted this way. Why it beat encoder-decoder models: There was just no input sometimes. Inputs work for translation, but there\u0026rsquo;s no clear input in a multi-turn conversation. Having both an encoder and a decoder also meant cross-attention added more surface area to the model. This also resulted in more ways for a training run to go wrong. If you want to see how far the decoder family tree has branched — LLaMA, Mistral, DeepSeek, Qwen, and the architectural deltas between them — Sebastian Raschka maintains an excellent visual gallery of LLM architectures.\nFrom here on, we\u0026rsquo;ll zoom into the decoder-only stack - the architecture behind GPT-style models - and walk through what happens inside it, one component at a time.\nTurning Text into Numbers In this section, we\u0026rsquo;ll also learn why the model keeps getting \u0026ldquo;how many r\u0026rsquo;s in strawberry\u0026rdquo; wrong. Yes, it has to do with tokenization!\nIf a transformer operates on sequences of vectors, how do we turn text into numbers? How does \u0026ldquo;The cat sat\u0026rdquo; turn into vectors? This is also known as tokenization, because we are breaking down a sentence into individual tokens.\nWhat are simple ways to break up a sentence into tokens? Here\u0026rsquo;s a few:\nWords: \u0026ldquo;The\u0026rdquo;, \u0026ldquo;cat\u0026rdquo;, \u0026ldquo;sat\u0026rdquo; are all valid words in the dictionary. However, unknown words have to be represented by \u0026ldquo;\u0026lt; UNK \u0026gt;\u0026rdquo;, where UNK stands for unknown. And when we decode and encode that word, we lose precision. Characters. Every character gets represented, however it is a lot longer than what we need. Subwords. Can we break up a word like \u0026ldquo;misrepresented\u0026rdquo; as \u0026ldquo;mis\u0026rdquo;, \u0026ldquo;represent\u0026rdquo; and \u0026ldquo;ed\u0026rdquo;, to individual tokens? This is the key idea behind Byte-Pair Encoding, which was actually invented in 1994. BPE was invented by Philip Gage for data compression, and was brought to NLP in 2016 by Sennrich et al. for machine translation.\nLet\u0026rsquo;s take a look: BPE on our sentence: after three merges, \u0026ldquo;at\u0026rdquo;, \u0026ldquo;th\u0026rdquo; and \u0026ldquo;the\u0026rdquo; have joined the vocabulary. Real tokenizers keep going to ~32k-200k.\nThe steps are simple:\nStart with a base vocabulary of individual characters. Count every adjacent pair of tokens Find the most frequent pair, and merge it into a new token, adding it to the vocabulary. Repeat step 2-3 until the vocabulary hits your target size (which is a model hyperparameter, typically around 32-200k). Now, BPE has evolved a little more than that into byte-level BPE. In GPT-2, BPE is run over raw bytes, so that the base vocab is exactly 256 symbols and every possible string (across languages) is just some sequence of bytes.\nWhere BPE falters There are a couple of places where BPE falters. Here\u0026rsquo;s three of them:\nToken spelling: \u0026ldquo;How many r\u0026rsquo;s are there in strawberry?\u0026rdquo; This is hard for a model because it could have broken out \u0026ldquo;straw\u0026rdquo; and \u0026ldquo;berry\u0026rdquo; into two different tokens, say token 2516 and 426. There would be no way for the model to know the spelling of token 426. How this was fixed: Reasoning (model actually decomposes the word), memorization (remembering that strawberry has 3 \u0026lsquo;r\u0026rsquo;s), and tool use (model writes code to count the number - do you see ChatGPT or Claude sometimes suddenly writing code to answer a math problem? Well, this is why.)\nArithmetic: Asking a model to do math is hard because digits are arbitrarily chunked. This was fixed by splitting digits into their own tokens.\nMultilingual cost inequality: While in English tokens cost roughly 4 characters each, other languages may cost more tokens per character. This is definitely trickier.\nHow this was fixed: Well, it\u0026rsquo;s still ongoing, but OpenAI\u0026rsquo;s GPT-4o tokenizer (o200k) doubled the vocabulary to 200k tokens and dramatically reduced token counts for non-English languages, letting them fit more words into the mix.\nThere\u0026rsquo;s also been new advances in this space, such as ByT5 and the Byte Latent Transformer (Meta, 2024), which operate on raw bytes, learning to group bytes into variable-size patches. This is an interesting frontier, but for the sake of this article, will be left as extra reading.\nFrom Token IDs to Meaning: Embeddings So tokenization gives us integers. But a token ID like 1024 is just a row number — it doesn\u0026rsquo;t mean anything by itself.\nThe thing that gives it meaning is a giant lookup table called the embedding matrix: one row per vocabulary entry, where each row is a long vector of numbers.\nThe length of that row is the model\u0026rsquo;s hidden size - in many 7B-class models, that\u0026rsquo;s 4,096 numbers per token (our running example uses the original paper\u0026rsquo;s 512). That\u0026rsquo;s a crazy large number, but if you think about it, it could actually encode a lot of information.\nCan a model encode too much information? Yes, in the sense of waste. A wider hidden size gives each token more room to store features, but past a point you get diminishing returns, where extra dimensions end up redundant or near-empty.\nLet\u0026rsquo;s work through an example.\nWhen the tokenizer hands the model the ID for \u0026ldquo;cat\u0026rdquo;, the model looks up that row and works with the vector instead. Here\u0026rsquo;s the magical part\u0026hellip; nobody designs these vectors!\nThey\u0026rsquo;re learned during training, and semantically similar tokens end up close together in space - \u0026ldquo;cat\u0026rdquo; lands near \u0026ldquo;kitten\u0026rdquo;, \u0026ldquo;mat\u0026rdquo; near \u0026ldquo;rug\u0026rdquo;. The geometry even supports arithmetic: the famous example is king − man + woman ≈ queen.\nIn this image, we\u0026rsquo;ve squashed our vector size to 2-dimensions. Take note that there are up to 4,096 dimensions as mentioned above.\nHowever, there\u0026rsquo;s one thing the embedding does not carry: position within a sentence. The vector for \u0026ldquo;cat\u0026rdquo; is identical whether \u0026ldquo;cat\u0026rdquo; is the first word of your prompt or the fiftieth.\nPositional Encoding Now in order for encoders and decoders to understand order in a sentence, we have positional encoding. This is placed at the start of the Encoder and Decoder stacks, immediately after raw tokens are converted into word embeddings, and before the Attention layers.\nWhere positional encoding sits in the stack - and what gets added to \u0026ldquo;cat\u0026rdquo;.\nNow for our same example, we need to distinguish between \u0026ldquo;The cat sat on the mat\u0026rdquo;. and \u0026ldquo;The mat sat on the cat\u0026rdquo;. To do this, we need each word\u0026rsquo;s position! For instance:\nPosition 0: \u0026ldquo;The\u0026rdquo; Position 1: \u0026ldquo;cat\u0026rdquo; Position 2: \u0026ldquo;sat\u0026rdquo; Position 3: \u0026ldquo;on\u0026rdquo; Position 4: \u0026ldquo;the\u0026rdquo; Position 5: \u0026ldquo;mat\u0026rdquo; The model handles this by generating a unique positional vector for each index and merging it with the word vector.\nPositional Vector: A 512-length vector is generated specifically for Position 1. We used sine and cosine waves in 2017 since they would fluctuate strictly between -1 and 1 and would not overpower the actual meaning of words. Word Vector: The word \u0026ldquo;cat\u0026rdquo; is converted into a 512-length vector representing its semantic meaning. Limitations of Sine \u0026amp; Cosine Waves However, there were limits to these sine and cosine waves for positional vectors. There were three key reasons:\nCorruption of word semantics: The sinusoidal method adds position directly into a token\u0026rsquo;s meaning. We\u0026rsquo;d rather have a way to encode position into a vector\u0026rsquo;s angle, preserving its length and not distorting the original features. Generalized sequence lengths: Positions are limited by the inputs. E.g. if you trained an LLM on 4000 token limits, it would not understand positioning for slots 4001 and beyond. Relative text shifts: Sometimes the same grammatical phrase occurs at different places in a text. The math with positional encoding leaves behind absolute room numbers that change the final score based on where the phrase sits in the document. Often all we want is the relative position, not the absolute one. We look at Rotary Positional Encoding (RoPE, proposed by Su et al. in 2021) to solve this. It gets a little deep here, so I\u0026rsquo;ll try to simplify: Instead of changing a word\u0026rsquo;s vector by adding position numbers to it, RoPE injects order by physically rotating the vectors in a mathematical space based on their position index.\nRoPE turns position into rotation: the hand\u0026rsquo;s length (meaning) never changes, only its angle (position).\nThink of each word as a clock hand. The hand\u0026rsquo;s length is the word\u0026rsquo;s meaning which never changes. Its angle is the position: \u0026ldquo;cat\u0026rdquo; at position 1 gets rotated a little, \u0026ldquo;mat\u0026rdquo; at position 5 gets rotated more.\nNow this way, we have a much better way to represent token positions.\nIs Attention All We Really Need? In 2017, eight researchers at Google were racing to finish a paper for NIPS, the field\u0026rsquo;s biggest conference, before the deadline. The work was done. The architecture that would reorganize all of AI was sitting in the draft. What they didn\u0026rsquo;t have was a title.\nThen, Llion Jones, the Welsh researcher among them, threw one out almost as a joke: after the Beatles\u0026rsquo; \u0026ldquo;All You Need Is Love,\u0026rdquo; why not \u0026ldquo;Attention Is All You Need\u0026rdquo;? He didn\u0026rsquo;t expect the team to keep it, but — they kept it!\nAnd the joke turned out to be the most honest summary. Where everyone else was bolting attention onto recurrent networks, this paper deleted the recurrence and kept only attention.\nHeads up: There\u0026rsquo;s going to be a lot of math in the next few sections. Take your time to digest them, and use your favorite AI to break it down for you. I recommend understanding this formula in-depth because it is the basis of today\u0026rsquo;s LLM models.\nLet\u0026rsquo;s now take a look at the key paper: Attention Is All You Need. Most importantly, look at this formula in Section 3.2.1: Scaled Dot-Product Attention:\nI know this formula looks scary if you haven\u0026rsquo;t seen it before, but let\u0026rsquo;s walk through it together with an example.\nThe key thing to note is that instead of RNNs reading one word at a time, we are going to have a sentence where each token looks at every other token.\nBack to our favorite example: The cat sat on the mat. \u0026ldquo;sat\u0026rdquo; attends to every token. If you look at those with the highest scores, it\u0026rsquo;s what\u0026rsquo;s relevant to \u0026ldquo;sat\u0026rdquo;. For instance: who sat? \u0026ldquo;cat\u0026rdquo; (0.42). Sat where? \u0026ldquo;mat\u0026rdquo; (0.30).\nNow each token gets broken down into three vectors:\nQuery: What am I looking for from other tokens? Key: What can I give to other tokens looking at me? Value: What gets passed along when I match with other tokens? Remember matmuls from the first chapter? Here\u0026rsquo;s where they come into play. Every token\u0026rsquo;s Query is going to be dot-producted with the Key of each other token, to produce a score. These are all values that have been initialized to random weights.\nThis calculation is QKT, where KT just means the K matrix transposed (flipped along its diagonal). We do this to produce a similarity score.\nStep 1: \u0026ldquo;sat\u0026rdquo;\u0026rsquo;s Query dot-producted with every Key, then divided by √dk.\nWe then divide by √dk. This refers to the key dimension (e.g. if the Key is [0.1, -0.5, 0.9, 0.2], the dimension is 4, and the sqrt of that is 2).\nThe reason we do this is that dot products of bigger vectors produce bigger numbers: the variance of the dot product of two random dk-dimensional vectors grows to dk, so dividing by sqrt(dk) brings the scores back to a sane scale before they hit the softmax.\nThen we need to convert these scores to probabilities.\nTo do this, we use the softmax function. Softmax exponentiates each score and normalizes.\nThe reason for the exponent is to eliminate negative numbers, since e raised to any power is always positive.\nStep 2: softmax exponentiates and normalizes - each row becomes a probability distribution.\nNow last but not least, we have to multiply the weights by the Value matrix. This focuses the model on the most compatible and valuable keys, producing a context-aware representation of each token.\nStep 3: each token\u0026rsquo;s Value is scaled by its attention weight and summed - \u0026ldquo;cat\u0026rdquo; and \u0026ldquo;mat\u0026rdquo; dominate the blend that becomes the new vector for \u0026ldquo;sat\u0026rdquo;.\nThe best part of all is that this is all happening in parallel, which is why GPUs have become so popular.\nMulti-Head Attention Now to take this even more parallel, we can split up the vector dimension into multiple smaller vectors. Each of these smaller vectors is a head. In the paper, the model\u0026rsquo;s total vector dimension was set to 512, split into 8 heads of 64 dimensions each. Each head learns to look for something different: one might track which noun a word refers to, another might track syntax.\nThe 512-dim vector is split into 8 heads of 64 dims each. The heads run in parallel, then their outputs are concatenated back into a single 512-dim vector.\nThat 64 turns out to be a very hardware-friendly number. Remember GPU warps from the first chapter? A warp is a group of 32 threads executing in lockstep, so a 64-dimensional head divides cleanly across warps. These days, AI engineers choose the head dimension - usually 64, 128 or 256 - to align with the GPU\u0026rsquo;s Tensor Cores.\nWhat do thousands of heads actually learn? Anthropic\u0026rsquo;s interpretability team found induction heads - heads that spot the pattern \u0026ldquo;A B \u0026hellip; A\u0026rdquo; in your prompt and predict B comes next. They\u0026rsquo;re one of the clearest known mechanisms behind in-context learning, where the model picks up a pattern from your prompt and continues it.\nAn induction head spots that the current token already appeared earlier, then predicts whatever followed it last time - \u0026ldquo;A B \u0026hellip; A\u0026rdquo; → B.\nMasking Matrices Now in the last part of Section 3.2.3 in the paper, there\u0026rsquo;s a key sentence here: \u0026ldquo;We implement this inside of scaled dot-product attention by masking out (setting to −∞) all values in the input of the softmax which correspond to illegal connections.\u0026rdquo;\nWe have another step here: to apply a mask matrix (kind of like a mask in Photoshop).\nThis ensures that only past and present words can be seen, and that the model cannot cheat by looking ahead during training. This is also known as autoregression in statistics. Why The Context Window Matters Here\u0026rsquo;s a crucial property of attention: its computational cost scales quadratically, O(N²), with the sequence length. E.g. if you double your prompt from 100 to 200 tokens, the compute and memory for attention quadruples.\nThis is why context windows exist at all. Every new token has to attend to every token before it, and the model has to keep all those Keys and Values around in memory (the \u0026ldquo;KV cache\u0026rdquo;).\nWith RoPE, there\u0026rsquo;s a second ceiling: positional coordinates only stretch as far as you\u0026rsquo;ve trained them. A model trained on 4,000-token sequences has simply never seen a rotation angle for position 50,000. So the context window is bounded by both compute and the positional encoding\u0026rsquo;s training range.\nOne more wrinkle: even within the window, attention isn\u0026rsquo;t spread evenly. Researchers documented a \u0026ldquo;lost in the middle\u0026rdquo; effect - models use information at the start and end of long prompts far more reliably than information buried in the middle. This is why prompt tips like \u0026ldquo;put the important context first\u0026rdquo; or \u0026ldquo;repeat the key instruction at the end\u0026rdquo; work well.\nThe Other Important Blocks Let\u0026rsquo;s go back to the transformer architecture. Attention gets all the headlines (and the paper title), but if you look at the diagram, there are still several other important blocks.\nAdd \u0026amp; Norm This pair shows up after every attention unit and every feed-forward unit, and it does two jobs.\nAdding: the block\u0026rsquo;s output (the residual) is added back to its input. This is to make sure that gradients don\u0026rsquo;t vanish halfway through a deep stack (just like how they did in RNNs). Think of it as a gradient highway: even if one layer learns nothing useful, the signal can flow straight through the addition unharmed. This trick came from ResNet in computer vision, and it\u0026rsquo;s a big part of why we can train models hundreds of layers deep.\nNormalization: after all that adding, we rescale each token\u0026rsquo;s vector to mean 0 and variance 1, to make sure we don\u0026rsquo;t blow numbers out of proportion:\nThe formula, if you were interested, is: $$\\text{LayerNorm}(x) = \\gamma \\cdot \\frac{x - \\mu}{\\sigma} + \\beta$$ Reading it left to right: $x$ is the token\u0026rsquo;s incoming vector. $\\mu$ is the mean of its values and $\\sigma$ their standard deviation, both computed per token across the vector\u0026rsquo;s own dimensions (not across the batch). Subtracting $\\mu$ then dividing by $\\sigma$ re-centers the vector to mean 0 and variance 1.\nBut forcing every vector to look statistically identical is too rigid, so we hand back two learned knobs: $\\gamma$ (gamma) rescales the spread and $\\beta$ (beta) shifts the center, and the model figures out the right values for each during training.\nModern models like Llama use a leaner cousin called RMSNorm (skip subtracting the mean, just divide by the magnitude), and apply it before each unit rather than after - \u0026ldquo;pre-norm\u0026rdquo; - because it trains more stably.\nFeed-Forward Network In this unit, each token digests what happened during attention. It takes its current vector and blows it up to 4x the width, before squashing it back down:\nThe FFN blows each token\u0026rsquo;s vector up 4x, bends it, and squashes it back. No token-to-token communication.\nWhy do we do this? When we blow up the vector, we can analyze it across many more dimensions. Think of a detective who empties a cramped case folder out across a huge table: in the original 512 dimensions, clues are stacked on top of each other and hard to tell apart, but spread across 2048 dimensions each pattern gets its own bit of room.\nNow, there is no token-to-token communication here. This all happens within one token.\nThat GELU (Gaussian Error Linear Unit) in the middle is the activation function. Modern models like Llama, Mistral and PaLM have since moved on to SwiGLU, a different variant.\nHere\u0026rsquo;s the part that surprised me: this boring block is where most of the model\u0026rsquo;s parameters live.\nFor the original 512-dimension transformer, $W_1$ is 512x2048 and $W_2$ is 2048x512 - about 2.1M parameters per layer, roughly double what attention uses.\nIn modern LLMs, the feed-forward layers hold around two-thirds of all weights. There\u0026rsquo;s research suggesting they act as key-value memories - this is plausibly where \u0026ldquo;Paris is the capital of France\u0026rdquo; is stored.\nFormula for those who are keen: $$\\text{FFN}(x) = \\text{GELU}(xW_1 + b_1),W_2 + b_2$$\nThe Final Linear Layer After the final block, each token\u0026rsquo;s vector gets multiplied by one last matrix (the \u0026ldquo;LM head\u0026rdquo;), producing one score per vocabulary entry - if your tokenizer has 200k tokens, that\u0026rsquo;s 200k scores.\nThese scores are called logits, and a final softmax turns them into next-token probabilities. We\u0026rsquo;ll look at this in-depth in the next essay.\nThe Attention Formula Let\u0026rsquo;s wrap up. How do we write Attention in code?\nWell, good news. The entire attention formula fits in about a dozen lines of NumPy. Andrej Karpathy also covers this in a longer video, but here it is:\nimport numpy as np tokens = [\u0026#34;The\u0026#34;, \u0026#34;cat\u0026#34;, \u0026#34;sat\u0026#34;, \u0026#34;on\u0026#34;, \u0026#34;the\u0026#34;, \u0026#34;mat\u0026#34;] d_k = 4 np.random.seed(0) E = np.random.randn(len(tokens), d_k) # pretend embeddings Wq, Wk, Wv = (np.random.randn(d_k, d_k) for _ in range(3)) Q, K, V = E @ Wq, E @ Wk, E @ Wv scores = Q @ K.T / np.sqrt(d_k) # QKᵀ / √dk mask = np.triu(np.ones_like(scores), k=1) # hide the future scores = np.where(mask == 1, -np.inf, scores) weights = np.exp(scores) / np.exp(scores).sum(-1, keepdims=True) # softmax output = weights @ V # weighted sum of Values print(np.round(weights, 2)) # each row: where that token is looking That\u0026rsquo;s the transformer!\nNext time you watch tokens stream out of Claude, you\u0026rsquo;ll know what\u0026rsquo;s happening underneath: bytes merging into tokens, vectors rotating by position, and every word attending to every word that came before it.\nExtra Reading Section! Because the entire essay is so long, I\u0026rsquo;ve left the following in an appendix here. I know most of you won\u0026rsquo;t get to these, but they are very useful for understanding modern architecture.\nGrouped-Query Attention Back to Attention for now. When the model generates the token 5,001, it needs the Keys and Values of all 5,000 tokens before it. Recomputing them every step would be madness, so they\u0026rsquo;re cached - per layer, per head. At long contexts, this cache can eat more VRAM than the model weights themselves.\nThe modern fix is Grouped-Query Attention (GQA): instead of every head keeping its own Keys and Values, groups of query heads share one K/V head. Llama 2 70B runs 64 query heads against just 8 shared K/V heads; Mistral 7B does 32 against 8. You keep almost all the quality of full multi-head attention with a fraction of the cache - which is why nearly every open model since 2023 ships with it.\nFlashAttention We\u0026rsquo;ll look into ways to optimize this through FlashAttention, invented by Tri Dao et al. in 2022. Remember the memory bottleneck that we explored last chapter? GPUs have a lightning-fast but tiny SRAM, and a big but slow VRAM (HBM).\nThe naive implementation computes the full N×N score table for \u0026ldquo;The cat sat on the mat\u0026rdquo;, writes it out to slow VRAM, reads it back in to do the softmax, writes it out again\u0026hellip; Now, the matrix isn\u0026rsquo;t the bottleneck. The memory traffic is. How do we solve this?\nWith FlashAttention. The trick here is to use tiling. The sentence gets cut into small blocks, and each Streaming Multiprocessor (SM) grabs a tile - say, the Queries for \u0026ldquo;The cat\u0026rdquo; against the Keys for \u0026ldquo;the mat\u0026rdquo; - and computes that entire piece inside its local SRAM without ever writing the intermediate table out to VRAM. To keep the math correct, it accumulates running statistics locally and applies the softmax normalization exactly once at the end of each row (this is called the online softmax).\nThe result is the exact same answer as standard attention, just with far fewer round trips to slow memory. No approximation, just better traffic management.\nThere\u0026rsquo;s also been further headway in this space, which I will leave as further reading:\nLinear Attention \u0026amp; State Space Models (Mamba): Architectures trying to replace the attention formula completely to achieve (O(N)) linear scaling. RoPE Scaling (YaRN \u0026amp; RoPE Interpolation): The mathematical tricks used to compress and stretch coordinate systems, letting a model trained on 4,000 tokens seamlessly read 128,000 tokens. Mixture of Experts (MoE) There\u0026rsquo;s a newer technique that turns a dense transformer block into a sparse network. This increases the total parameter count by hundreds of billions of weights without increasing the FLOPs spent per token. How do we do this?\nWe swap out the single Feed-Forward Network for a fleet of them - the \u0026ldquo;experts\u0026rdquo; - plus a tiny router in front.\nThe MoE router scores all 8 experts and wakes only the top 2 - the rest of the parameters sleep.\nWhen a token vector enters the MoE layer, the router multiplies the token vector by its own weight matrix to score every expert, then routes the token to the top-k highest-scoring experts. Only those experts run; their outputs are summed, weighted by the router\u0026rsquo;s scores. Not all the experts are activated at the same time - that\u0026rsquo;s the entire trick.\nThe math here:\n$$y = \\sum_{i ,\\in, \\text{TopK}} g_i(x), E_i(x), \\qquad g(x) = \\text{softmax}(x \\cdot W_g)$$\nSo when \u0026ldquo;cat\u0026rdquo; passes through, maybe expert 3 (which has drifted towards noun-ish, animal-ish patterns during training) and expert 7 fire, while the other six sleep. You get a model with trillions of potential parameters where each token only pays for a fraction of them.\nThe idea dates back to Shazeer et al. in 2017 (yes, the same Noam Shazeer from Attention Is All You Need - and yes this guy just moved to OpenAI this week.), scaled up by Google\u0026rsquo;s Switch Transformer. Today, Mixtral 8x7B (8 experts, top-2: 47B total parameters, only 13B active per token) and DeepSeek-V3 (256 experts plus 1 shared, top-8: 671B total, 37B active) are leveraging this.\nOne catch worth knowing: routers can collapse, learning to send everything to a couple of favorite experts while the rest atrophy. The fix is a load-balancing loss during training that nudges tokens to spread out evenly.\n","permalink":"https://sidwyn.com/posts/inside-an-llm/","summary":"\u003cblockquote\u003e\n\u003cp\u003e\u003cem\u003eThis was initially published on \u003ca href=\"https://www.pathtostaff.com/p/everything-a-senior-engineer-needs\"\u003ePath to Staff\u003c/a\u003e, but I\u0026rsquo;m bringing it here as I start to write more about learning AI.\u003c/em\u003e\u003c/p\u003e\u003c/blockquote\u003e\n\u003cp\u003eWelcome back to Path to Staff! This series is a little different from our usual programming. In this series, we\u0026rsquo;re covering LLMs and AI in-depth.\u003c/p\u003e\n\u003cp\u003eAs an engineer, I never really had the time to understand AI\u0026rsquo;s internals. But I\u0026rsquo;ve spent the past few weeks doing deep research to unpack it all.\u003c/p\u003e","title":"Everything a Senior Engineer Needs to Know About What's Inside an LLM"},{"content":" This was initially published on Path to Staff, but I\u0026rsquo;m bringing it here as I start to write more about learning AI.\nWelcome back to Path to Staff. I recently left Meta for personal reasons (not laid off!), and have found much more time to write. This means learning as much as I can and distilling what I learn into these articles.\nNow back to the topic: As an engineer who never really interfaced much with AI, I realized I didn\u0026rsquo;t know much about it at all. Since I’ve more time now, I\u0026rsquo;ve spent the past few weeks diving really deep to understand AI from the bottom up. And I want to share those learnings with you today.\nThis deep dive series is a five-part series:\nThe Hardware Behind AI – And How It\u0026rsquo;s Programmed. Transistors, semiconductors, and fabricators. Learn about the big players (TSMC, Nvidia, ASML). The memory-compute bottleneck. And all the acronyms you always wondered about (TPU, ASIC, FPGA, CUDA, etc.) Data \u0026amp; Model Architecture. Learn about how models are made of. We\u0026rsquo;ll cover the paper that started it all (\u0026ldquo;Attention is All You Need\u0026rdquo;), plus talk about transformers and diffusion models. And of course, we\u0026rsquo;ll cover how training data is prepared for these models (what sources? how is the data decontaminated and filtered?) Training. The meat of teaching a model. How does pretraining work? What goes into it (backpropagation, optimizers, loss functions)? What scaling laws before we kick off an expensive training run (up to hundreds of millions $)? Post-Training \u0026amp; Alignment. How does one guide a model once it\u0026rsquo;s been taught? How do we apply safety? How do we benchmark and know the model got better? How do we evaluate a model\u0026rsquo;s performance? Inference, Serving and Agents. This might be the most familiar topic, since it\u0026rsquo;s closest to you as an AI user. How does a model output its token and serve the result to you (SSE)? How do systems stay fair and fast? What tools are available (MCP, RAG, tool use) and how do agents work? Over the course of this series, I expect the syllabus to change as I learn more about AI. I also welcome questions in the comments, since this will help me tweak it. If there are several acronyms, I also list their definition at the top of the section.\nQuick note: As much as I love using AI in my work, most of these will be written by hand, with light editing by AI. If you\u0026rsquo;re interested in reviewing drafts, please also reply to this post and let me know!\nTransistors And Their Importance At a glance\nAcronyms Definition EUV Extreme Ultraviolet - type of lithography (manufacturing) used to print intricate scale onto chips ASML Advanced Semiconductor Materials Lithography - Dutch company that makes EUV machines TSMC Taiwan Semiconductor Manufacturing Company, leading semiconductor factory based in Taiwan To understand Artificial Intelligence we first have to go to the core of it. AI runs off GPU chips. These chips are made from transistors which are manufactured using EUV machines.\nLet\u0026rsquo;s break each of these down, starting with a transistor.\nA transistor is a semiconductor device that controls the flow of electricity. It uses a small electrical signal at one terminal to control a much larger current. It either (1) boosts a signal, or (2) decides whether a current can pass. In other words, it acts as either an amplifier or switch. A semiconductor, most commonly silicon, is a material that conducts electricity only under certain conditions. Its conductivity can be modified by adding impurities.\nWho designs these chips? Nvidia and AMD are the biggest players when it comes to designing chips. Not far behind is Google, Amazon and now Meta when it comes to chip design. However, these companies operate as \u0026ldquo;fabless\u0026rdquo; designers. That means they only build the architecture, while outsourcing the physical production which costs up to $20B for a foundry. These foundries require tens of billions a year in capital spending to stay current, a cost that most of these companies are not interested in.\nWho then, makes these transistors? TSMC (Taiwan Semiconductor Manufacturing Company) does. But they require special machines, called extreme ultraviolet (EUV) machines, and a process called lithography, which is the act of printing on chips. There are other foundries like Samsung and Intel, but they are not as advanced as TSMC, which currently holds 70% of the foundry market.\nWho makes these EUV machines? These are only being manufactured by ASML, which has a monopoly foothold in the EUV machine industry. China is fast catching up, but is still roughly 5 years behind.\npicture of ASML laboratory\nFun fact: there are no major competitors to ASML today. It took them 30 years to reach the stage they\u0026rsquo;re in today. They\u0026rsquo;ve integrated thousands of suppliers together to build a generator that fires 50,000 droplets per second. They also own the major company Cymer that makes these EUV sources. These light sources are so short (as short as 13.5nm) that there is no natural source for it.\nOK, now we know what a transistor is and how it\u0026rsquo;s made. Now we can understand GPUs (graphical processing units). A single GPU contains billions of transistors, which are packed onto a die (a raw block of silicon) manufactured with a 3nm fabrication technology. Don\u0026rsquo;t worry, we\u0026rsquo;ll get into these different types of fabrication technologies in just a bit.\nDie Shrinks: Going from 10,000nm -\u0026gt; 2nm in 5 decades Terminology Definition Die Microchip cut out from a silicon wafer Die shrink Manufacturing technique that redraws the same circuit design smaller Microprocessor Type of chip that is used in the CPU (central processing unit) of a computer Let\u0026rsquo;s take a quick history detour of die shrinks. Die shrinks mean that we\u0026rsquo;re redrawing the same circuit design smaller and smaller, shrinking it! Why do we need it to be smaller? This allows manufacturers (like TSMC) to produce a significantly higher number of dies on a single piece of silicon, lowering the cost it takes. There\u0026rsquo;s also an additional side benefit, that chips will use less power and generate less heat, allowing for higher clock speeds and performance.\nWhat\u0026rsquo;s the history of these dies and die shrinks? Well, it\u0026rsquo;s complicated, but I\u0026rsquo;ll try to do a quick tour in a couple of paragraphs. The first microprocessor was launched in 1971, at a 10,000nm (or 10 micrometers) line width. This means that the gate length (distance between drain and source electrodes) was at 10,000 nm. This means that electrons have to travel across this 10,000nm whenever a transistor switches on. A shorter gate = faster speed.\nAt this point of time (1971), this chip, the Intel 4004, was designed for a Japanese calculator company, Busicom. However, once Intel realized that this was much more useful at mass market, Intel repurchased the marketing and technology rights from Busicom. A few years earlier, in 1965, Gordon Moore made the observation you\u0026rsquo;ve heard of as Moore\u0026rsquo;s Law: that the number of transistors on a chip would keep doubling at low cost — originally about every year, which he revised in 1975 to roughly every two years.\nOver the next few decades, we went from 600nm -\u0026gt; 250nm -\u0026gt; 180nm -\u0026gt; 130nm -\u0026gt; 45nm. In the early 2000s, manufacturers hit a wall. There was no way to take the next jump and shorten gate length. However, a breakthrough from a TSMC engineer named Burn-Jeng Lin thought of adding water between the lens and the wafer. This was a huge bet by ASML in 2003-04, which was at that time a smaller European challenger behind Nikon and Canon. They went all in on immersion and won.\nNikon and Canon stuck to their guns on 157nm dry lithography, but by then it was too late. ASML had a huge headstart. Canon essentially exited leading-edge lithography, and while Nikon did eventually build immersion tools, it never recovered the lead, and later sat out EUV entirely.\nToday, your iPhones run on 3nm chips. GPUs today all use TSMC\u0026rsquo;s 5/4/3 nm variants. We\u0026rsquo;re currently at 2nm, and there\u0026rsquo;s targets to hit 1.6nm (TSMC\u0026rsquo;s \u0026ldquo;A16\u0026rdquo;) around late 2026–2027, with 1.4nm later and true 1nm not expected until the back half of the decade.\nUnfortunately, one sad fact is that these numbers no longer mean gate lengths. They’re used more for marketing. If anything they refer to transistor density over the previous node, measuring MTr/mm^2 (millions of transistors on a mm^2).\nFive decades of shrinking dies.\nThe Shift from CPU to GPU Let\u0026rsquo;s talk a bit about how GPUs got famous in the first place. It all started with the CPU.\nCPUs have been in place since 1971, since the first microprocessor. However, when games like Quake were introduced in 1990s, computers were lagging pretty badly trying to render graphics. I remember my own computer grinding to a halt whenever a game was played.\nGPUs were designed exactly by NVIDIA to solve this. Instead of having a few sophisticated cores, you\u0026rsquo;d have thousands of dumb cores. Each individual core is super weak, but when combined together in a GPU, the throughput is gigantic. These were great for games, since rendering a 4K image meant computing colors of 8M pixels independently.\nIn 2006, Jensen Huang, CEO of NVIDIA made a huge bet. That Moore\u0026rsquo;s Law is slowing. Single-threaded CPU performance was not optimized for the long run. He wanted to build a programming platform for scientific computing on graphics cards. This bet was targeted at scientists who wanted access to supercomputers. At that point of time, these supercomputers were multi-million dollar machines only owned by government labs and a few corporations.\nThis bet was called CUDA (Compute Unified Device Architecture). This allowed the CPU to offload parallelized computing tasks from the CPU to the GPU. This ended up being their moat. The ecosystem (PyTorch, TensorFlow, which we will cover in later chapters) ended up being built CUDA-first. Silicon and networking was also specialized around this architecture.\nThe first hint of AI leveraging GPUs came about in 2012. Three University of Toronto researchers, Alex K., Ilya S. and Geoffrey H. submitted a neural network called AlexNet. This was trained on two NVIDIA GTX 580 gaming GPUs in Alex\u0026rsquo;s bedroom. This proved that GPUs were feasible to train deep neural networks for the first time. These neural networks, which were mostly based on matrix multiplication was able to be achieved in someone\u0026rsquo;s bedroom. If the same neural net were to be trained on a CPU, it would have taken centuries.\nHow a GPU is structured. You can see that there are many more cores in a GPU! But at the same time, there\u0026rsquo;s less \u0026ldquo;Control\u0026rdquo; and \u0026ldquo;Cache\u0026rdquo; layers, these are shared. Taken from the CUDA programming guide.\nStructure of an NVIDIA GPU Acronym Definition SM Streaming Multiprocessor - self-contained core GPU Graphical Processing Unit - explained in prior sections GPC Graphical Processing Cluster - group of streaming multiprocessors MIG Multi-Instance GPU - sliced up into isolated logical GPUs to let providers run multiple tenants HBM High Bandwidth Memory - stacked DRAM chips sitting on the same package as the GPU DRAM Dynamic random-access memory - one transistor + one capacitor per bit. cheap but needs constant refreshing SRAM Static random-access memory - faster memory with six transistors, fast but expensive. CUDA NVIDIA\u0026rsquo;s proprietary programming model on their GPU architecture TSV Through-silicon via (tiny vertical wires) that connect HBM DRAM stacks. Now let\u0026rsquo;s take a look at a GPU. Fair warning: it starts to get very technical from here on out. I will try my best to break them down.\nThe first is an overall view of the Blackwell GPUs (launched Q4 2024). This is not the latest chip architecture: Rubin R100 was recently announced and plans to ship in a few months. I could not seem to find a good infographic on R100, so let me know if you do!\nNevertheless, this does give us a good sense of how it works end-to-end. Cleaned up version of the image from Blackwell\u0026rsquo;s overview.\nHere in Blackwell, we have 2 dies that have been welded together with a custom interconnect called NV-HBI. An interconnect is basically a physical wire linking two things together. In this case, NV-HBI is an ultra-low latency, proprietary die-to-die interconnect that powers 10 TB/s.\nNow, let\u0026rsquo;s think of each die as a city. Each city contains:\n🟩 4 GPCs (Graphics Processing Clusters), each holding 20 SMs A GPC is a district within the city. Think of it as a district housing a group of factories that share some local infrastructure. Within a GPC, there are 20 SMs, which means that there are 80 SMs per die, and 160 SMs total across both dies. That\u0026rsquo;s a lot of compute power!\n📋 GigaThread Engine + MIG Control The city-level dispatcher. Its job is simple: to receive work from the CPU (through the PCIe Gen 6) and farm it out to the GPCs. It also has an accompanying control portion, where the \u0026ldquo;MIG\u0026rdquo; part stands for Multi-Instance GPU. This lets the chip be sliced into up to 7 logical GPUs. Each of these GPUs looks isolated to different tenants, which is important for cloud providers running multiple customers on one physical chip.\n🟦 L2 Cache We have 50MB of shared SRAM (static random-access memory) in each L2 cache. Remember, SRAM is expensive, high-access RAM. Given this dual-die architecture, any SM can read what the other SM wrote. L0 and L1 caches are within the SM, which we\u0026rsquo;ll cover in the next section.\n🟦 8 HBM3E stacks (288 GB total, up to 8 TB/s) Surrounding the dies, we have these High-Bandwidth Memory (HBM) stacks. In order to build them, we stack DRAM (dynamic random-access memory), with 4 stacks per die.\nThese DRAMs are connected by Through-Silicon Vias (TSVs), aka microscopic vertical wires drilled through the silicon, connected by an interposer which is a thin slab of silicon that sits beneath the GPU die and HBM stacks.\nThe \u0026lsquo;E\u0026rsquo; in HBM3E stands for extended – a refresh of the HBM3. Only three companies in the world make the HBM: SK Hynix (~55% share), Micron and Samsung (~20% each). HBM4 is already being sampled and ramping up for NVIDIA\u0026rsquo;s Rubin.\nOn SRAM vs DRAM: SRAM uses 6 transistors per bit and holds its value as long as it\u0026rsquo;s powered on. It\u0026rsquo;s expensive since it\u0026rsquo;s bulky, but much faster. DRAM uses 1 transistor + 1 capacitor, but the whole array refreshes thousands of times per second. As such, SRAM lives on the GPU die which is closer to the GPC, and faster but more expensive. On the other hand, DRAM lives off the die. Both are however volatile, and gets lost once power gets cut. We\u0026rsquo;ll dive into this deeper during the memory-compute wall.\nLast but not least, we have the I/O paths surrounding the image:\nNVLink v5 (1.8 TB/s) → connects to other GPUs via NVSwitch PCIe Gen 6 (256 GB/s) → connects to the host system NVLink-C2C (900 GB/s) → connects to a paired CPU coherently (the Grace+Blackwell \u0026ldquo;superchip\u0026rdquo; combo). These are interconnects which we\u0026rsquo;ll learn about shortly. Within a Streaming Multiprocessor Now within a Graphical Processing Cluster (GPC), there are 20 Streaming Multiprocessors (SM).\nWhat does each SM contain? Let\u0026rsquo;s take a look. Each SM is split into 4 partitions, which are all done in parallel. Remember we mentioned that Blackwell Ultra has 160 of these SMs, so this means 640 partitions in total.\n📋 L0 Instruction Cache This is a super fast cache that sits next to the Warp Scheduler. When the warp scheduler needs the next instruction (aka what to do, be it a matmul or a different operation), it pulls from this.\n📋 L1 Instruction Cache (top strip) This contains the recent instructions that are shared across all 4 partitions. If there\u0026rsquo;s an L0 miss, it hits the L1 Cache.\n📣 Warp Scheduler (32 threads/clock) This is the shift manager, where upon every clock cycle, it picks a warp (32 threads) and issues an instruction. Remember, this is a GPU here, so all 32 threads in the warp execute the same instruction simultaneously on different data.\n📤 Dispatch Unit The dispatch unit next to the Warp Scheduler helps to find the right execution unit. Now, there\u0026rsquo;s going to be different execution units for different purposes. A CUDA core is used for regular multiplication (e.g. 3 x 2), a Tensor Core is used for a matrix multiplication, and an SFU is used for transcendental functions (exponential, logarithmic, trigonometric, etc.).\n🟦 Register File (64KB) This is the fastest possible storage. Registers are sitting physically adjacent to the execution units. And working values are kept here. This is used especially to tune performance. Keep note of this, as we will return to this when we talk about the memory-compute bottleneck.\n🟪 CUDA Cores This is where work happens. Each partition contains an assortment of execution units here. FP32 refers to 32-bit floating point math, INT32 is for integer math, and FP64 is for scientific computing.\nA key point to note here: every time a clock cycle happens, a CUDA core does one multiply-add. A fun point, CUDA cores were introduced in 2006 in the GeForce 8800 GTX, as a shader.\n🟪 Tensor Cores Remember how AlexNet was introduced in 2012? Everyone wanted to train neural networks after that happened. However, CUDA cores were limited to one multiply-add per clock. This is because they were for scalar arithmetic. Given that neural networks are matrix multiplications, there needed to be a different type of cores.\nAt GTC 2017 (Nvidia\u0026rsquo;s conference), the Volta V100 was launched, and the Tensor Core was introduced. It specialized in 4x4 matrix multiplication (matmul). That\u0026rsquo;s 64 multiply-adds per clock, already a 64x improvement! With each V100 SM handling 8 Tensor Cores, that was an insane amount of matmul work it could perform, roughly 125 TFLOPs.\nBy the time Blackwell arrived, new floating points FP6 and FP4 were supported with adaptive precision selection. In layman terms, it was extremely powerful at 15 petaflops (10^15) per second. This means 15,000,000,000,000 floating operations per seconds.\nAgain, this is all NVIDIA. We\u0026rsquo;ll take a look at other architectures like Google down the road. CUDA cores and Tensor Cores don\u0026rsquo;t exist outside of NVIDIA.\nGleaning across the next two: 🟧 SFU - These Special Function Units help to take care of transcendental operations (sine, cosine, exponential, logarithm, softmax, math functions that cannot be represented by basic algebra). 🟥 LD/ST Units: Move data between registers and larger memory tiers.\n🟦 Tensor Memory (TMEM) — 256 KB New in Blackwell. A dedicated SRAM pool that is reserved exclusively for tensor (multi-dimensional scalar/vectors) operations. Used to stash work in progress instead of forcing data back to L1 Cache.\n🟦 L1 Data Cache / Shared Memory — 256 KB (configurable) Shared workshop SRAM that all 4 partitions can access.\nBlackwell by the numbers When you add it all up, a single Blackwell Ultra SM contains:\n4 partitions 128 CUDA cores (32 × 4 partitions) 64 INT32 units, 64 FP64 units 4 fifth-generation Tensor Cores (1 per partition) 256 KB total register file (64 KB × 4 partitions) 256 KB Tensor Memory (new) 256 KB L1 / shared memory 4 Texture units 4 Warp Schedulers, 4 Dispatch Units L0 instruction caches per partition, shared L1 instruction cache for the whole SM LD/ST units, SFUs across all partitions Multiply by 160 SMs on the full chip and you get the scale:\n~20,480 FP32 CUDA cores total ~10,240 INT32, ~10,240 FP64 640 fifth-generation Tensor Cores ~40 MB of TMEM across the chip (a new tier of on-die SRAM) ~40 MB of L1/shared memory ~100 MB of L2 cache (shared between the two dies, fully coherent) Now that\u0026rsquo;s a lot! No wonder it costs around $30-40k for each GPU.\nNVIDIA vs Google vs Others Now let\u0026rsquo;s take a break from numbers. Let\u0026rsquo;s go back to history and understand who else is in this space.\nThere are three other major players. AMD, Google and Amazon.\nAMD sells chips called Instinct. Google rents TPUs, or Tensor Processing Units. And Amazon also rents their chips called Trainium/Inferentia.\nThere\u0026rsquo;s also Groq and Cerebras which are newer companies, formed in 2016. You might have heard of Cerebras as the \u0026ldquo;first AI-era IPO\u0026rdquo;, and Groq as a company started by a former TPU engineer that\u0026rsquo;s been bought by NVIDIA in 2025.\nI won\u0026rsquo;t cover them here in detail, but it\u0026rsquo;s worth checking them out. Groq bets on inference using a new LPU (Language Processing Unit), and Cerebras is making one huge chip to reduce interconnect tax.\nLet\u0026rsquo;s put these four (NVIDIA, AMD, Google and Amazon) into a table.\nNVIDIA GPUs AMD Instinct GPU Google TPUs AWS Trainium / Inferentia What it is Industry standard merchant GPU Lower cost GPU but software is slowly catching up Custom ASIC (application-specific integrated circuit) Specialized chips for training / inference. Sold / Rented Sold Sold Rented Rented Market Share 80% 5-10% 6-8% 2-3% Software CUDA - 15 year moat ROCm - open-source software JAX + TensorFlow (we\u0026rsquo;ll cover this in a future post) Neuron SDK Customers Almost everyone Microsoft and Meta, recently OpenAI Google Internal, Anthropic Amazon, Anthropic What\u0026rsquo;s interesting is that internal adoption - in the past Google and AWS used to run on NVIDIA GPUs within their clouds. But they\u0026rsquo;re slowly starting to displace NVIDIA inside their own clouds. For instance, TPU is 75% of Google\u0026rsquo;s Gemini and Trainium is 50% of AWS\u0026rsquo;s Bedrock. Anthropic is another interesting customer, because it\u0026rsquo;s serving Claude across both Google and AWS clouds. And of course, as of three weeks ago, Anthropic announced a deal with SpaceX to use all the compute capacity at Colossus 1 in Memphis (300MW worth). This is spread across at least three silicon types now - AWS, Google Cloud and NVIDIA (the Colossus GPUs, from what was xAI).\nA Quick History: NVIDIA vs Google vs AWS Most of the history revolves around NVIDIA and Google. NVIDIA starts at around 2006 with CUDA, and then Google starts their TPU journey around 2013, before bringing them public around 2018.\nIf there\u0026rsquo;s just one takeaway from each of these two companies, it\u0026rsquo;s this. Nvidia asks: How do I make threads more productive? Google asks: How do I keep the grid fed? How can I adjust my schedules?\nNVIDIA We\u0026rsquo;ll talk about launches within NVIDIA, since each of these are important and have interesting bets.\nThe codenames refer to GPU architecture codenames.\n2006: Launches Tesla with the CUDA architecture. This is the year where Jensen bet that parallel, individually weaker compute is going to be more important than large CPUs. 2010: Fermi. This introduced a full IEEE-754 floating point, teaching computers how to do decimal/math operations. ECC memory was also a milestone here where this memory could detect single corrupt bits that could ruin the entire job. 2012: Kepler. CUDA core counts jumped to 1500. This was the year AlexNet was trained on two GTX 580s as well. 2017: Volta (V100). Each of these chips now have the first letter of their chip\u0026rsquo;s name, as well as 100, which means the most performant chips. 2020: Ampere (A100). 8-GPU boxes were added, with new number formats BF16 and TF32. Basically, the less number of bits, the faster the compute, but less precision. For neural networks, these were helpful. 2022: Hopper (H100): Named after Grace Hopper, this was when the streaming multiprocessor turned into a dispatcher. 2024: Blackwell (B200). These are dual dies that we talked about, and chips are now optimized for FP4 (4-bit floating point numerals). 2026: Rubin (R100) \u0026lt;\u0026ndash; We are now here! A new 3rd gen transformer engine was introduced recently. Google TPUs 2013: Internal TPUs get developed. Silicons are spent mostly on a systolic matrix engine. Let\u0026rsquo;s stop here. The architecture of a TPU is important and I want to cover it. You\u0026rsquo;ll hear these a lot. The basis of the TPU\u0026rsquo;s matmul engine is a giant systolic matrix engine called MXU. The reason it\u0026rsquo;s called systolic is because it pumps data through the chip in rhythm (similar to the systole of a human heartbeat).\nA visual is probably the best way to look at the difference between a GPU and TPU. CPU: A few cores that handle one hard task each and stores results to memory. GPU: Thousands of small cores that store results to memory. TPU: Accumulates partial sums to pass to the next grid. Never has to return to memory.\nEach of these have different execution models you\u0026rsquo;ll hear about as well: CPU → SIMD (single instruction, multiple data), GPU → SIMT (single instruction, multiple threads), and TPU → Systolic .\nOf course, this might seem great. However, remember that we are only looking at one thing that TPUs excel at, matmul. They\u0026rsquo;ve been heavily optimized for matmul.\nBut these systolic arrays are weaker at branching heavy logic, sparse workloads, and irregular control flows when it comes to GPUs. For instance, softmax, a very important function in neural networks can run slower in a TPU than a GPU. As such, Google added other separate specialized hardware blocks called VPU (Vector Processing Unit), and Compilers to help with these other operations.\nAlright, let\u0026rsquo;s finish the history of TPU. 2017: TPU v1 paper goes public, and proved that GPU was not the only way for ML going forward. 2018: v3. MXU is scaled up with liquid cooling. Starts opening to Google Cloud customers. 2021: v4. SparseCore (specialized engine for data dependent embeddings) and Optical Circuit Switching also launches. 2023: v5e and v5p both get launched. v5p is especially interesting since trillion-parameter training is finally launched with a 3D torus shape. 2024: Trillium (v6e). 4x the multiply-accumulate count with a smaller HBM on each chip. The goal here was to optimize for inference throughput and cost efficiency with less dependencies on chip memory. 2025: Ironwood (v7). 64x more chips in one domain. 2026: TPU 8t and 8i. This time, it follows AWS\u0026rsquo;s lead by separating training and inference chips, and leans on Broadcom and MediaTek to help implement these chips with TSMC.\nAWS Trainium / Inferentia In 2019, AWS realized that they were too dependent on NVIDIA, and for a hyperscaler their size, this wasn\u0026rsquo;t the best position to be in.\nGiven that Google had proved that hyperscalers could build ASICs, they started to build their own chips, starting with Inferentia (for inference).\n2019: Inferentia enters the market, marketed as \u0026ldquo;good enough for a cheaper price\u0026rdquo;. 2021: Trainium adds a new training chip. 2024: Trainium2 launches with 30-40% better price performance, with Anthropic as the main customer. 2025-2026: Trainium3 ships and processes over 50% of Bedrock\u0026rsquo;s token throughput. Provides up to 2.52 PetaFLOPs of FP8 compute.\nBrief Segue into Floating Points \u0026amp; Numeral Formats Traditional computing standardized around FP32 (32-bit floating point). This was great for scientific computing and simulations. However, over time, research showed that neural networks could tolerate lower precision. Throughput also scaled like crazy with the reduction of bits. For instance a H100 chip can deliver 67 TFLOPs at FP32 and 3958 TFLOPs at FP8.\nThis kicked off a round of increasingly specialized numerical formats:\nFP32 → traditional scientific computing FP16 → early deep learning acceleration BF16 → wider exponent range for more stable training TF32 → NVIDIA’s Tensor Core optimized training format FP8 / FP4 → ultra-low precision formats optimized for modern large-scale inference. A brief visualization of floating point numeral formats\nWe\u0026rsquo;re currently at FP8/FP4 with NVIDIA\u0026rsquo;s Rubin and Google\u0026rsquo;s TPU, continuously learning how low number of bits an AI workload can tolerate. The art of performing this lossy compression is called quantization, where we are trying to reduce the memory footprint in order to increase inference speeds.\nInterconnects Before we jump into the final section, which I think is the most interesting problem of this piece, we need to talk about interconnects.\nWhy does it even matter? As GPU chips compute gradients during matmul operations, they need to share their gradients with each other. TPUs might have solved some of these problems with systolic operations, but that happens only within a chip.\nTo share these gradients, they use interconnects.\nNVIDIA has a full interconnect stack strategy. We talked about most of these in the prior section.\nWithin the chip: NV-HBI (NVIDIA\u0026rsquo;s High Bandwidth Interface) runs at 10 TB/s between the two dies. Chip-to-chip: NVLink 6 runs at 3.6TB/s for the Rubin series. Throughput roughly doubles every generation. Rack-to-rack: NVLink Switches which were introduced in Blackwell. Server-to-server: InfiniBand owns this layer. This is built by Mellanox, which was acquired by NVIDIA in 2020. Whereas Google on the other hand takes a bit of a different approach (take note of their torus topology and OCS tech):\nWithin the chip: No need for an interconnect Chip-to-chip: Inter-chip Interconnect (ICI): Runs at 9.6 Tbps. Uses a 3D torus topology which is a 3D lattice that loops back to itself. This is a lot cheaper at scale and given its twisted shape, every chip is connected to many others. Optical Circuit Switching (OCS) is also a key innovation here that we mentioned earlier. At provision time, the interconnect topology can be rewired to match the workload. It\u0026rsquo;s also important to look at what else is happening across the landscape. Other interesting companies are:\nCerebras: No interconnect. Don\u0026rsquo;t even cut the chip wafer into multiple chips. Build one gigantic chip. Groq: No dynamic networking. All network traffic is statically scheduled by the cluster. As well as new technologies that are being developed:\nNVLink Fusion: Allowing third-party chips to plug into NVLink. Developed by NVIDIA. Ultra Ethernet: Open standard built by AMD, Microsoft, Meta. This competes with InfiniBand and NVLink. Optical I/O: Photonic chiplets and interposers that route signals as light instead of copper traces. Today signals leave chips electrically and need to be converted to light and then back. If we remove that conversion, the energy savings is massive. Ayar Labs and Lightmatter are two companies that are tackling this space. Now we\u0026rsquo;ll talk about the biggest problem that GPUs have faced for awhile: the memory-bandwidth.\nThe Biggest Issue: Memory Bandwidth While GPU floating-point ops (aka compute) have been scaling exponentially every few years, memory bandwidth has not kept up. Memory Bandwidth = the speed at which data travels between the GPU chip and memory banks. This is now too slow.\nThis leads to underutilized compute on the GPU (meaning that you\u0026rsquo;re not getting the most of what you\u0026rsquo;re paying for).\nRooflines and Arithmetic Intensity Before we dive in further, we need to know what rooflines are. Roofline models help us understand whether an algorithm is memory bound or compute bound.\nLet\u0026rsquo;s take a fictitious example. Say your chip is a kitchen. Your chef is your compute (e.g. GPU cluster). It can perform 1000 TFLOP/s of knife work. The runner is your memory bandwidth, where it has 2TB/s of legs sprinting to the pantry (HBM) and back.\nIs your kitchen limited by chef speed or runner speed? To do that, we need to calculate the peak compute: 1000 TFLOP/s ÷ 2 TB/s = 500 operations per byte fetched. This means that recipes that do more than 500 things with each byte (e.g. ingredient) are bottlenecked by the chef. Less than that, and it\u0026rsquo;s bottlenecked by the runner.\nHere\u0026rsquo;s a real example with a graph. A chip that can do 1000 TFLop/s of compute and pull 2TB/s from memory (bandwidth), the arithmetic intensity is calculated as 1,000 ÷ 2 = 500 FLOPs per byte.\nThis means that if your algorithm\u0026rsquo;s Arithmetic Intensity is below 500, you\u0026rsquo;re bandwidth bound (limited by how fast you can move tensors). If it\u0026rsquo;s Arithmetic Intensity is above 500, you\u0026rsquo;re compute bound. I\u0026rsquo;ll refrain from turning Arithmetic Intensity into an acronym, well\u0026hellip; because it\u0026rsquo;s also AI. Now we have two algorithms, Algo 1 and Algo 2.\nAlgo 1 — matrix × one vector\nMath done: 2 × 10,000 × 10,000 = 200 million FLOPs Data read: 10,000 × 10,000 × 2 bytes = 200 MB Arithmetic Intensity = 200M ÷ 200M = 1 FLOP/byte 1 is way below the ridge of 500, so it\u0026rsquo;s bandwidth-bound. Each number you load gets used exactly once, so there\u0026rsquo;s no way to climb. Actual speed ≈ 1 × 2 TB/s = 2 TFLOP/s — just 0.2% of the chip\u0026rsquo;s peak. This means that the chip is mostly sitting idle waiting for memory. What a waste!\nAlgo 2 — matrix × matrix\nMath done: 2 × 10,000 × 10,000 × 10,000 = 2 trillion FLOPs Data read: 2 × 200 MB = 400 MB Arithmetic Intensity = 2T ÷ 400M = 5,000 FLOPs/byte Now, 5,000 is well past the ridge of 500, so it\u0026rsquo;s compute-bound and runs at the peak compute full 1,000 TFLOP/s. It\u0026rsquo;s the same matrix, but now every number you load gets reused thousands of times, so memory keeps up easily and the chip stays busy.\nSolving Memory Bandwidth All the above architecture that we studied help to address this memory bandwidth issue: large amounts of SRAM, L1-2 caches, quantization (remember the floating point stuff?), interconnects (NVLink, etc.), and moving memory as close to compute as possible.\nIn particular, the most interesting solution here is optimizing speed through HBMs, or High-Bandwidth Memory.\nSK Hynix, a Korean company, conceptualized stacking DRAM dies vertically and connecting them together using TSVs (through-silicon vias). This allows enormous amounts of data to move in parallel.\nHowever, HBM3e is getting increasingly expensive, given that manufacturing and packaging difficulties increases with dies per HBM stack.\nOne paper, by Xiaoyu M. and David P. that is most interesting shares more about four new research areas to address these memory and interconnect issues. I\u0026rsquo;d actually advise checking out the paper if you have time.\nThey propose four different hardware shifts, and I\u0026rsquo;ll try my best to dissect them, though the paper does a much better job:\nHigh-bandwidth Flash: stack NAND flash instead of DRAM. With HBM, you can only stack so much DRAM before heat and cost gtes high. Flash is cheaper, denser and its capacity keeps doubling. However Flash has its own write-endurance limits, but is useful for data that\u0026rsquo;s written rarely and read enormously (like model weights!) Processing-near memory (PNM): For datacenter LLM inference, it seems that shards with PNM can be 1000x larger, which would allow these partitions to have a low communication overhead. With processing-in-memory (PIM), arithmetic units inside DRAM means (1) weak compute since these units are fabricated for DRAM and (2) creation of many shards which creates a lot of communication overhead. 3D memory-logic stacking: In 2D solutions, the HBM sits besides the processor. This means that the memory bandwidth is capped by how much edge the logic die has. As the area grows quadratically, the shoreline only grows linearly. With 3D stacking, it runs the entire 2D area of the die, which allows bandwidth to scale with area quadratically instead of with the parameter. Right: 3D stacking, taken from the paper.\n4. Low-latency interconnect. Build topologies inspired by tori with high connectivity. Reduce work outside of the chip network and process within the network. Optimize chip design that land small packets directly into SRAM. Improve reliability through local standby spares and accepting good enough results. Again, more details in the paper that I can\u0026rsquo;t cover in a short paragraph.\nFive things to remember It all comes back to one problem: Memory Bandwidth. Compute got fast much faster than memory could feed it. Every acronym in this piece (HBM, SRAM, NVLink, quantization, 3D stacking) is an attempt to shrink the distance between the two. Three companies, three chokepoints. Designers like NVIDIA hand a blueprint to TSMC (~70% of all chip fabrication), which can\u0026rsquo;t print it without ASML\u0026rsquo;s EUV machines (a literal monopoly). Each of these three companies form a strong dependency chain. \u0026ldquo;nm\u0026rdquo; is mostly marketing now. We went from 10,000nm to 2nm, and the transistors got faster and denser the whole way. But shrinking the chip was never going to fix the memory gap. If anything it made all that compute harder to feed. NVIDIA and Google answer the same question differently. NVIDIA asks how to make each thread more productive (CUDA, Tensor Cores). Google asks how to keep the whole grid fed (systolic arrays that never go back to memory). Both are interesting strategies that we\u0026rsquo;ll see how it turns out over the next few years. The roofline tells you which side of the wall you\u0026rsquo;re on. In order to know if your algorithm is starved for memory, work out its arithmetic intensity: how many operations you do per byte loaded. You\u0026rsquo;re either compute-bound or memory-bound. That\u0026rsquo;s it! You made it to the end of the first part of this series. Stay tuned for Part Two: Data \u0026amp; Model Architecture. Learn about how models are made of. We\u0026rsquo;ll cover the paper that started it all (\u0026ldquo;Attention is All You Need\u0026rdquo;), plus talk about transformers and diffusion models. And of course, we\u0026rsquo;ll cover how training data is prepared for these models (what sources? how is the data decontaminated and filtered?)\nAs I work on this series, I\u0026rsquo;d appreciate feedback and comments. What did you like? What would you like to see?\nReferences (TO CLEAN UP) https://horace.io/brrr_intro.html https://x.com/MainzOnX/article/2044462083010662771 https://jax-ml.github.io/scaling-book/ https://people.eecs.berkeley.edu/~kubitron/cs252/handouts/papers/RooflineVyNoYellow.pdf\n","permalink":"https://sidwyn.com/posts/unpacking-ai-hardware/","summary":"\u003cblockquote\u003e\n\u003cp\u003e\u003cem\u003eThis was initially published on \u003ca href=\"https://www.pathtostaff.com/p/unpacking-ai-the-hardware-behind\"\u003ePath to Staff\u003c/a\u003e, but I\u0026rsquo;m bringing it here as I start to write more about learning AI.\u003c/em\u003e\u003c/p\u003e\u003c/blockquote\u003e\n\u003cp\u003eWelcome back to Path to Staff. I recently left Meta for personal reasons (not laid off!), and have found much more time to write. This means learning as much as I can and distilling what I learn into these articles.\u003c/p\u003e\n\u003cp\u003eNow back to the topic: As an engineer who never really interfaced much with AI, I realized I didn\u0026rsquo;t know much about it at all. Since I’ve more time now, I\u0026rsquo;ve spent the past few weeks diving really deep to understand AI from the bottom up. And I want to share those learnings with you today.\u003c/p\u003e","title":"Unpacking AI: The Hardware Behind AI"}]