A ground-up explanation of reinforcement learning fundamentals, policy gradients, and how proximal policy optimization (PPO) is used to align large language models with human feedback.
Post archive
Blog posts
Writings gathered into a readable stack.
-
-
A first-hand account of hand-dyeing a yard of linen using Procion MX dye, including how to calculate dye quantities and make soda ash from baking soda.
-
A practical walkthrough of C# for Python developers, covering task-based async patterns, file input classes, and building a .NET console application.