Bryan Catanzaro, VP of Applied Deep Learning Research at NVIDIA, walked us through how his team builds the company’s open models, the reasoning behind their architecture, and why NVIDIA open-sources so much of it.
ByteByteGo
https://blog.bytebytego.com · 52 posts · history since 2026 · active
Yesterday
23 Jul
Why does something as simple as reading the time become a hard problem for distributed systems?
22 Jul
In this article, we try to explore the collective thinking into a smaller set of practices and explain the reasoning behind each one, rather than asking anyone to memorize a numbered list.
21 Jul
We sat down with Anupam Singh, senior vice president of engineering at Roblox, to hear from him about the world model that Roblox is using to make its multiplayer games look photorealistic, the key insights that have come from taking that approach, and the next big thing that the Roblox team is focusing on.
18 Jul
Agents are capable on their own. Combined with tools and other agents, their capabilities compound.
16 Jul
In this article, we will understand multi-tenant architecture from the basics, along with its various benefits and challenges.
15 Jul
In this article, we will look more closely at the different solutions by following the support pipeline from first principles, show why a tail of cases resists automation regardless of model quality, and use these three approaches to understand how these can be handled.
14 Jul
In this article, we will look at how that learning actually happens, starting with why instruction-following alone falls short, then walking through the two main methods for teaching preferences (RLHF and DPO).
13 Jul
To understand what it actually takes to ship agents at that scale, we spoke with Marco Casalaina, VP of Products for Microsoft Core AI.
11 Jul
A Docker container starts with a single command, but that command has to be turned into a running Linux process. Here is what actually happens.
10 Jul
Our 7th cohort of Becoming an AI Engineer starts tomorrow, Saturday, July 11. This is a live, cohort-based course created in collaboration with best-selling author Ali Aminian and published by ByteByteGo.
9 Jul
When is the data complete enough to be moved to the compute stage?
8 Jul
In this article, we will walk through that progression. We will also look at how an agent is structured, what choices the model makes on every turn, what scaffolding holds it together, and when an agent is actually the right pattern to reach for.
7 Jul
In this article, we will look at the various architectural forks the teams building these models encountered and the decisions they took.
6 Jul
Our 7th cohort of Becoming an AI Engineer starts in less than a week. This is a live, cohort-based course created in collaboration with best-selling author Ali Aminian and published by ByteByteGo.
4 Jul
For this article we spoke with the team behind World, including Tiago Sada and Lily Gordon at Tools for Humanity, on how they try to solve this problem.
2 Jul
When an application grows geographically, it is logical to start serving it from a second region to improve latency and availability.
1 Jul
In this article, we will look at the entire journey in detail and challenges the OpenAI engineering team faced.
30 Jun
In this article, we will look at what the research preview covers and the concept of an interaction model proposed by Thinking Machines.
29 Jun
In this article, we will try to understand how that architecture gets built, from the constraint that forces it to exist all the way to the tradeoffs that follow.
27 Jun
RAG connects LLMs to your data and there are three different ways to do it.
25 Jun
In this article, we will look at some of the most important anti-patterns in service architecture, how they happen, and how they can be avoided.
24 Jun
In this article, we will explore those constraints through three layers of model design, look at the tradeoffs that come with each approach, and investigate the production systems that combine both small and large models.
23 Jun
If you’re on the same journey of making your work with agents more productive and enjoyable, I hope this gives you a head start and shortcuts some of your own exploration.
22 Jun
Individual gains do not become organizational gains on their own. This is the playbook for making that leap. Let’s dive in.
20 Jun
Twelve models worth knowing in 2026, each with one standout strength.
18 Jun
In this article, we will look at the basics of observability in detail with concepts like logs, metrics, and traces explained in detail.
17 Jun
We’re launching Cohort 2 of our 2-day intensive, cohort-based course, Build with Claude Code, taught by John Kim, who has trained hundreds of engineers at Meta to use Claude Code in real production workflows.
16 Jun
In this article, we will look at how open-weight models have transformed the AI landscape.
15 Jun
In this article, we will walk through how inference works and why the field’s optimization techniques exist.
13 Jun
Over to you: Which layer of the stack do you think is the hardest to get right in production?
11 Jun
In this article, we’ll go through the main deployment strategies used in production today, looking at how each one works, what it costs, and when it makes sense to use.
10 Jun
We’re looking for multiple part-time instructors to teach AI and engineering cohort-based live courses.
9 Jun
We sat down with John Kucera, Salesforce’s CPO of Agentforce, to learn what separates agents that deliver real business value from those that stall after a good demo.
8 Jun
To understand how teams keep this under control in production, we sat down with Scott Breitenother and Sid Sijbrandij, co-founders of Kilo, an open-source coding agent that runs through a lot of these loops every day.
6 Jun
Latency, throughput, and bandwidth often get used interchangeably, but each one tells a different story about performance.
4 Jun
In this article, we follow the journey of a web request one hop at a time.
3 Jun
The hardest part of data analysis isn’t writing SQL. It’s finding the right tables to use in the first place and understanding semantically how to use data.
2 Jun
This piece is a working guide for engineers who want to land on the productive side of that split.
30 May
In this article, we will learn how they built this flywheel and the key takeaways.
28 May
In this article, we will look at the most significant failure mode patterns in distributed systems and the standard approaches to deal with each of them.
27 May
In this article, we will look at how Airtable’s data infrastructure team built its architecture, the challenges they faced, the tradeoffs they accepted, and why the choices they made only make sense once their data is properly understood.
26 May
In this article, we examine the constraints Vercel faced, the choices they made in response, and the optimizations that produced the speedup.
25 May
In this article, we will look at how the CockroachDB engineering team built this index and the challenges they faced.
23 May
Ask an LLM about your company's data and it will guess. The two patterns that fix this are RAG and agents, and they solve different problems.
22 May
The first cohort starts in about a week: May 28-29, 2026.
21 May
In this article, we will look at each of these patterns in detail, along with their advantages.
20 May
In this article, we will understand how Netflix built this system and the challenges it faced.
19 May
For Snap, machine learning is closer to the product itself than a feature on top of it.
18 May
Grab’s data engineering team had a problem that looks familiar to anyone who’s maintained shared infrastructure.
16 May
An AI agent can be thought of as a simple While-loop.
15 May
Our 6th cohort of Becoming an AI Engineer starts tomorrow, Saturday, May 16. This is a live, cohort-based course created in collaboration with best-selling author Ali Aminian and published by ByteByteGo.