ByteByteGo Newsletter
Subscribe
Sign in
Home
Sponsoring ByteByteGo
Become an AI Engineer Cohort
Archive
About
Latest
Top
Discussions
Learn Claude Code, evals, AI systems, and more: ByteByteGo Live is here
Most online courses never get finished (~4% completion). Live cohorts get ~40%, roughly 10x higher. Live courses are the only courses people actually…
14 hrs ago
•
ByteByteGo
173
4
A Guide to Application Networking Basics
In this article, we will look at the various aspects of networking in detail.
Sep 10
•
ByteByteGo
87
4
How Smart Model Routing Can Cut LLM Costs by 10X
Cost reduction isn’t a given. It also depends on the types of requests the application receives, the price difference between models, and how well the…
Sep 9
•
ByteByteGo
252
5
11
Built for Reliability: How American Express Processes Payments at Scale
In this article, we will try to understand how the transaction runs through such a cell-based architecture and how the payments are processed even when…
Sep 8
•
ByteByteGo
235
3
5
How to Deal With Errors and Failures in LLM-Powered Applications
Apart from normal processing, the application also sends data to a large language model (LLM). It then uses the model’s response to carry out a task.
Sep 7
•
ByteByteGo
280
5
9
EP224: MCP vs RAG vs AI Agents
An AI agent is kind of an AI system where the agent performs the task autonomously and takes the decisions.
Sep 5
•
ByteByteGo
309
7
11
How Databases Keep Their Sanity with Concurrency Control
So how do we handle such bugs? This is what we are going to try to answer in this article.
Sep 3
•
ByteByteGo
94
1
Why Your RAG System Is Only as Good as Its Translator Model
In this article, we’re going to look at how this embedding model works in an RAG setup and what makes it such a critical part of the system.
Sep 2
•
ByteByteGo
270
13
How to Shrink a Language Model Without Making it Too Dumb
Models have grown roughly 100-fold in a few years, while consumer graphics memory has roughly doubled. It’s not just a matter of tightening things up to…
Sep 1
•
ByteByteGo
258
2
5
August 2026
What Happens Inside an AI Chatbot Between Enter and the First Word?
In this article, we are going to look at this entire journey in detail.
Aug 31
•
ByteByteGo
256
5
10
Background Work: From Cron Jobs to Distributed Systems
In this article, we will look at various such strategies to perform background work in detail.
Aug 27
•
ByteByteGo
109
9
How to Make LLMs 3X Faster
In this article, we will look at how speculative decoding works.
Aug 26
•
ByteByteGo
278
5
13
This site requires JavaScript to run correctly. Please
turn on JavaScript
or unblock scripts