Full Width [alt+shift+f] Shortcuts [alt+shift+k]
Sign Up [alt+shift+s] Log In [alt+shift+l]
1

100 Lessons From Building a 100-Person Org

from Miguel Carranza [alt+shift+b] in programming

RevenueCat now has more than 150 team members. Around 100 of them are in my org. This is still a little crazy to me. My own org is now bigger than any company I had worked at before starting RevenueCat. As a technical founder, I always preferred computers over people. Computers are deterministic. People are emotional. But if you find real PMF, team building becomes unavoidable. You have to build more. You have to move faster. You have to do more things in parallel. And at some point, the company stops being a product that a small group of people can brute force into existence and becomes an organization that needs to keep improving without breaking itself every week. I have made a lot of mistakes over the last almost 9 years. Some were obvious in hindsight. Some were expensive. A few were probably unavoidable. These are 100 of the lessons I have learned getting to a 100-person org. Recruiting is a long game. Just because someone says no today does not mean it is a no forever. Build the relationship, leave a good impression, and keep the door open. Hiring is always a priority, even when you are ahead of the headcount plan. If the company is growing, it is much easier to fall behind than to stay ahead. External recruiting agencies rarely work. Their incentive is not to find the best fit; it is to close a candidate as fast as possible. Excellent recruiters are very hard to find. When you do find one, everyone’s life gets a lot easier. The brand compounds like every other company metric. Candidates who were unreachable a year ago may eventually reply. Great people attract more great people. Hire enough of them and recruiting starts to get its own kind of virality. Most people oversell when recruiting. Don’t. The best candidates see through the BS and value honesty about both the good and the bad. Candidates appreciate knowing the salary and process upfront. It saves time for both sides. Have a compensation philosophy you can explain and defend. People might take a pay...
3rd Jun 2026

Stay updated

Get a weekly newsletter with the top 5 articles worth reading every week.

More from Miguel Carranza

The Dopamine Problem Every New Manager Face

I recently wrote an internal Management 101 guide for new RevenueCat managers. It is not meant to be a universal management guide. Most universal management advice is either too vague to be useful or too generic to survive contact with a real company. My goal was simpler: explain what we expect from managers, why the role matters, what tends to feel weird at first, and where to start when things get messy. One of the sections I think matters most is the dopamine problem, because it is one of the most immediate changes. The weirdest part of becoming a manager is that your old scoreboard disappears. As an IC, you can point to things you shipped. Code merged. Bugs fixed. Systems improved. Customers helped. Even hard weeks usually have visible artifacts. Management work is slower and less satisfying in the short term. You can have a great week and nothing obvious happened. A project got slightly clearer. A candidate stayed warm. A conflict did not become a fire. A direct report heard feedback early enough to fix something. A team avoided a bad architectural decision. A future manager took one more step toward readiness. That is real work. It just does not always feel like it. This is one of the hardest parts of the transition from IC to manager, especially for Engineering Managers. You are used to being useful in a very concrete way. You know how to debug the issue, review the design, write the code, fix the incident, and unblock the team. Then suddenly a lot of your work becomes indirect. You need to stop trying to solve the problem in front of you. You need to make the team better at solving that type of problem in the future. It is also getting harder in the age of AI. Agentic coding is even more addictive. You can describe what you want, watch the agent work, review the diff, and get a working prototype or bug fix in minutes. The loop is fast and incredibly satisfying. This is why many new managers cling to coding, technical details, or direct execution. It gives them the old dopamine. It also feels useful when there is urgency: a project is at risk, a technical issue is blocking the team, or execution is moving slower than you think it should. Sometimes stepping in directly is the right call. If production is broken, do not sit there having a philosophical debate about leverage. Help. But most of the time, direct execution is a patch around the real problem. If the project only moves when you personally drive it, something else is broken. Maybe ownership is unclear. Maybe the team does not have enough context. Maybe the technical plan is too complex. Maybe the bar is too low. Maybe someone needs feedback. Maybe you have the wrong people on the problem. Maybe you are still acting like the tech lead because that is more comfortable than acting like the manager. Your job is to make the team capable of executing without you, either by hiring the right people, training people up, clarifying ownership, or raising the bar. If you keep solving the work yourself, you prevent yourself from learning the new job and prevent the team from growing into it. You need to rewire your satisfaction from “I shipped the thing” to “the team is stronger and shipping better because of the work I did.” That takes time. Some practical advice: Set weekly goals for yourself. Not vague goals like “be a good manager.” Concrete goals: write the 30/60/90 plan, give feedback to X, unblock Y, clarify ownership for Z, source X candidates, review hiring funnel metrics, talk to another manager about a problem. Create accountability. I use weekly emails. Use whatever works, but do not let weeks disappear into meetings. A full calendar is not evidence of impact. Write things down. Writing forces clarity and creates artifacts when the work itself is abstract. If you cannot write the decision clearly, you probably have not made the decision. If you cannot write the feedback clearly, you probably do not understand the behavior well enough yet. Look for compounding work. Better onboarding, better interview calibration, clearer team ownership, stronger tech leads, training people up, and improved decision-making pay off repeatedly. They may not feel as satisfying today, but they make the company more capable. Management is not a promotion from building. It is a different way of building. You are building the system around the work: the team, the standards, the ownership model, the feedback loops, the hiring bar, the culture, the decision-making. The hard part is that the system does not always give you a green checkmark at the end of the day. Do it anyway. Originally published as an X article on July 10, 2026 and archived here.

10th Jul 2026 • 2 votes
The Long Run

They say only one out of 1000–2000 startups becomes a unicorn, and that only 1% of humans will ever finish a marathon. If you had asked my mother, though, she would have told you it was a thousand times more likely I’d build a billion-dollar tech company than finish a marathon. And I would’ve agreed. I was that kid in school with straight As who still almost failed PE. Two laps around the soccer field and I’d be hyperventilating. On New Year’s Eve last year, my co-founder Jacob slipped on the snow and broke his ankle. A gnarly injury. Fast-forward to March 1st, and he announces to the entire company and our investors that he is going to run the California International Marathon on December 7th. And also, he invited people to join him in his crazy adventure. I considered it for a split second and literally laughed out loud. This guy will do it, I thought. He’s delusional and stubborn enough to crawl the last 10 miles if necessary. But me? No chance. The longest run of my life was a single 10K, once, years ago. I ran it, and then immediately retired from the sport. But the idea stayed in the back of my mind. If Jacob crippled, needing surgery, with months of PT, could commit to running a marathon… what excuse did I realistically have? I was healthy. I had a head start. And honestly, after the twins were born, my exercise routine was basically nonexistent. Running was something I could technically do anywhere. My wife, fully aware of how much I hated running and how hard this would be for me, became a supporter. So I downloaded Runna. And I ran. I hated it. Deeply. Running felt like the most boring, repetitive, form of suffering ever invented. But something weird happened: I started seeing progress. Three miles stopped feeling like a death sentence. Then I ran 10. Then one random morning before giving a conference talk, I ran a half marathon by myself. And it didn’t even feel that hard. I started waking up earlier and earlier to get my mileage in before the morning meetings. Discipline compounds quietly, until suddenly you’re a different person. This year I traveled a lot. Yet running turned out to be the easiest habit to keep. I ran more than 900 km in 18 different cities. Running became a new, exciting way to explore a city. A nice way to start the day with a small win before all the startup chaos. This weekend, we did it. We ended the race in way more pain than we probably should have, but we crossed the finish line. And the biggest lesson I learned, bigger than training plans, pronation, and fueling… is this: You need to be a little crazy and do hard things. And that’s kind of the whole point. Life isn’t about being the best in the world. It’s about being the best version of yourself. Expanding what you thought your limits were. Rewiring your internal narrative from “I could never do that” to “actually, maybe I can.” I’m insanely proud of Jacob for this comeback story, and grateful (even if involuntarily) that he dragged me into another ridiculous adventure. It was worth it. And maybe this story nudges you toward something you’ve always dismissed as impossible. After running these 26.2 miles, building a decacorn feels easy. Okay, not easy, but you get the point. Will I do it again? I don’t know. If I had put this much time into surfing, I would suck a lot less at it and probably enjoy it more 😅. But I’ll keep running. It keeps me fit, clears my mind, and makes me a better version of myself. Which, in turn, helps me be a better founder, partner, and dad. Special thanks To Runna, genuinely a fantastic training app (and RC user!). To Marina, for the endless support and the childcare coverage while I disappeared for very long runs. To Andrew, Harlan, Bobby, Susannah, Sarah, Nacho, Chris, Mica, Matt, and Olga, for the advice, motivation, and for putting up with my running updates.

7th Dec 2025 • 1 votes
My role as a founder CTO: Year Seven

2024 has come and gone, and it’s time for my annual post. What a year for startups—like squeezing five regular years into one. Do you remember the Apple Vision Pro, the DMA regulation, founder mode, or the o1 launch? All of that happened in just the last twelve months. It’s also been wild at RevenueCat. My journal is full of stories that could fill a whole book or even a few HBO Silicon Valley seasons. Some are inspiring, some are hilarious, and others are honestly gnarly. Due to limited space, the need for context, and respecting everyone’s privacy, I’ll cover only the most interesting topics at a high level. Looking back, it was a good year for RevenueCat. Actually, a great one. Perhaps our best since 2020. We have plenty to celebrate: We accelerated again, and we hit our C10 revenue plan. We made our first acquisition and welcomed an amazing founder to the team. We signed our first multi-million-dollar contracts. We went to more than 20 events around the world… …including hosting our own conference, featuring its own award ceremony. We became the #1 payments SDK on iOS. Our swag went to 11. We launched over 80 user-facing features. We showed up in Times Square and along Highway 101. We raised a mini Series C and welcomed two new board members. We landed in Japan for the first time. Our API grew beyond 2B requests per day. We are processing nearly twice the TAM we had when we started the company. OpenAI is a friend of the cat. We continued building a winning team. We hired people who were on my to work with bucket list even before we started the company. By all the metrics, 2024 was our year of shipping and selling. We absolutely helped developers make more money. My role If you compare our progress to the goals I wrote about in last year’s blog post, it’s clear we succeeded. But the journey itself was a lot rockier than I had imagined. Many things didn’t go as planned, somes hires did not work out, and some strategy changes were really hard to push through. At the start of the year, besides my usual responsibilities, I also set four personal goals to help me scale with the company. I wanted to: Ship code more consistently. Talk to at least one customer every single day. Get more involved in areas outside of engineering. Stay very close to my co-founder Jacob, giving him my full support. I completely missed goal #1. Honestly, that hurts because I love building software. But I’m not too upset about it—our customers care that the team ships, not that I personally do. And in a way, I did help the team ship. As for the other three goals, I hit them, at least based on the company’s results. Still, my self-perception wasn’t always great. I found myself acting much more like a co-founder/executive than a CTO. Some weeks were brutal, with constant context switching across tasks and teams I didn’t always enjoy. At one point, there were over 60 people in my org, spread all around the world, which is pretty intense for an introverted computer kid. A lot of bullshit escalates to the top, making me question my entire existence some days. And then life threw serious personal emergencies at some of our team members too. Like I said last year, life is what happens when you’re busy building your startup. As we got closer to 100 people, it was statistically unavoidable that we’d face a few life-changing traumas—sometimes several all at once. As a founder, you need to be supportive and empathetic, but also protect your mental health. As a human, it’s tough. Add in a couple of two-year-olds who were often sick and not sleeping, and it felt like a ticking bomb. Did I burn out for the first time in my life? I don’t think so, but it got close. I’m confident being a founder (and having a co-founder) kept me going. I care too much and must stay resilient. If I’d been just an employee, I might have tapped out. But enough about the tough parts. Aside from being a professional BS handler, here’s how I spent most of my time: Lots of travel This was the year I traveled the most in my life—and looking back, I probably should have traveled even more. I visited customers, helped with sales, conducted executive interviews, and spoke at a few conferences. It’s hard being away from two little kids, but each trip turned out to be worth it. Support One of my biggest worries this year was our support function. Everything was fine, but our Support Engineering Manager was going on parental leave, and the bus factor was scary. I’ve seen support crises before—they’re not fun. It wouldn’t have killed the company, but at our size, it might have forced us to pull senior engineers into support and slow down our product velocity. Time was short, so we tried a few things that worked: We hired two new Developer Support Engineers who already knew our product—former customers! Their onboarding was smooth, and they hit the ground running. We split the support team into two pods with their own leads. Each pod handles certain tickets, and they collaborate with each other instead of relying too heavily on one person. We finally set up our first on-call rotation for emergencies. Sales and post-sales Reviewing sales and implementation calls, giving technical input, collecting enterprise customer feedback for product and engineering, joining calls, and even doing some in-person visits. Writing more As our team continues to grow across the globe, not everyone has the same direct interaction with me or Jacob as before. A lot of our culture and collaboration style was once passed through observation, but now needs to be written down to reach everyone faster. These days, my code editor is basically replaced by Google Docs. I’ve been publishing more internal and external documents—like our Engineering Strategy and an updated How to Work with Miguel. Product delivery We still have a few details to refine, but our founder Shipping and Timeline reviews have been valuable. It gives Jacob and me a high-level view of everything in progress, lets us offer feedback, and helps us dig deeper where needed. It’s a great way to see each team’s capacity, find bottlenecks, keep a sense of urgency, and deflate anything that’s not truly important for the customers. Re-orgs This year brought a couple of big reorgs in product and engineering. Overall, they went well, but the puzzle gets more complex with each new piece. We created sub-teams to narrow their focus and keep things running smoothly. For the first time, we had a couple of management layers between me and our ICs. One major change was shutting down our Enterprise/Reactive team. The idea was solid at first, and they delivered plenty of value, but eventually they became the random tasks team. Our new plan is to reinforce the rest of the teams: quick enterprise requests go to the right team to handle them reactively, while bigger projects move to the main roadmap (which keeps us disciplined). The engineers from that old team will help bootstrap new teams as we hire in 2025. Hiring Our engineering hiring goals weren’t super ambitious, but we still reached them. The pace was a bit uneven, and some roles took longer to fill than desired. However, when Hiring Managers took ownership of the process (with recruiting as a support) it made a huge difference in the quality of our candidates. We also brought back the founder interview stage: Jacob or I spoke with every single candidate before making an offer, and we plan to keep doing this for the foreseeable future. Random founder stuff Not my main focus, but I still spent a fair amount of time handling people-related topics, operations, investor, customers and partners relations, fundraising, company policies, planning, etc. Learnings: deepened insights On culture Shipping is king. Yes, deadlines are stressful, but failing to ship and getting stuck in endless debates is far worse. It’s depressing. If someone isn’t a good culture fit, it will be more than a single incident. Over time, it becomes pretty clear to everyone. Top performers tend to measure themselves against the very best in the company. Reassure them they are doing a great job. On the other hand, those who underperform will look to other underperformers to gauge their own progress. Even top performers will struggle if they don’t fully align with the vision. Nothing beats talking to customers. Encourage everyone to do it. Better in person. Most managers don’t have a strong incentive to be strict, so finding the right balance takes a lot of calibration. Only once everyone is aligned, you can truly delegate. On hiring Managers hate hiring because it’s binary: a lot of “no”, but eventually one “yes” can change everything. Staying consistent really helps. Simple things, such as weekly updates keep everyone accountable. People management sucks. If someone’s main motivation is to be a manager for the title, or so they can “lead,” that’s a red flag for me. Best managers end up being those who never planned on becoming one in the first place. They actually roll up their sleeves and can do the work. Big ideas are great, but they need to be executed. This matters even more when it comes to executives. A truly great exec can change your life and it will feel like a brand new company. Given their influence, anything less than great will eventually turn into a big mess. Spend time with executive candidates in person. Watch how they work and make sure they’re real builders. Previous founders and true engineers are usually a little bit de-risked, but they’re still not guaranteed. I also like to write a very detailed onboarding doc, clarifying context, expectations, and what success or failure looks like. Having them write a 30/60/90-day plan helps us all align. I’ve learned not to rely on their past pedigree. The real key is their glass eating endurance. On company building Re-orgs are inevitable at a growing startup, and they will feel scary or emotional for people who aren’t used to them. I find it helpful to be super clear about why we’re doing it, and to share the fallback plan if things don’t work out. But I’ve learned that by the time you think you need a re-org, you’re already behind. We now plan to reassess our team structure twice a year for optimal shipping. A numeric goal (like an SLA or revenue target) is just a proxy. Missing it isn’t the end of the world if you learn in the process. But if you surpass it without recognizing what’s broken underneath, you’ll be masking critical issues. Perfectionists are great, but can have a tough time at startups. Some do fine if they’re focused on one specific thing or working as an IC. But as soon as they have to juggle multiple tasks, the chaos can feel overwhelming. Help them embrace it. Things will break. It’s about continuously reprioritizing, and stopping small cracks from becoming big fires. No agenda == no meeting. Synchronous time is expensive. The only exception is the occasional unscheduled call. After you reach around 50 people (especially in a remote setup), documenting every process change becomes crucial. I’ve learned the hard way that simply talking about a new process isn’t enough. Founders can’t talk to everyone all the time anymore, and misunderstandings or gossip can spread fast. Not everyone will read everything, but at least there’s a single source of truth to reference. Best people can stretch quite a bit—they’ll often rise to the challenge. But it’s wise to keep an eye on their limits before they burn out or become a bottleneck. On scaling as a founder A great EA is life-changing. Family will be supportive, but it’s not fair to offload all the stress on them. They will end up feeling helpless. Having a solid co-founder, a network of peers, or a good executive coach makes a world of difference. Founders are the ultimate guardians of the culture. It’s constant work, and it will feel relentless—especially when things are going well and it’s easy to get entitled. The best team members will help uphold the standard, but you cannot expect them to do all the policing. At this stage, it’s stupid not to level up your lifestyle in ways that reduce stress or save time. This might include childcare support, having a second car, investing in a better mattress, or hiring help with housekeeping. The startup is bigger than its founders, and the goal is to continuously remove yourself from the critical path. Still, it’s easy to forget that, as a founder, you literally brought everything into existence from thin air. Imposter syndrome often creeps in when you step outside your core expertise, but if your gut feeling is strong, it’s worth paying attention to. Stay open-minded, yet remember that no one knows the company quite like you do. The future Next year is going to be another big one. We’ll keep shipping and selling, while finally tackling our design and UX debt. We’ll keep investing heavily in our infrastructure. Not just for reliability, but also for real-time data. We want RevenueCat to feel fast, accurate, and easy to use. We’re also upgrading our self-serve and enterprise support, aiming for a truly world-class experience. In many ways, we’re finally seeing the original vision Jacob and I had back in 2017 come to life. We’ll be launching new product lines too, and if we execute well, we will be just a couple of years away from hitting $100M in revenue. We’ll keep building a winning team. We’ll hire about 45 people, 30 in Engineering, Product, and Design. It’s a challenge, but totally doable. As for me, my personal goals haven’t changed much, but my perspective has. Jacob and I used to joke that being happy and winning can’t happen at the same time. But why not? We’re in a privileged position to shape our own path and change anything we don’t like. After a lot of reflection and coaching, I realized I was simply too hard on myself. I was feeling depressed by the constant BS even though we were winning. A hack that helped was working with my EA to set weekly goals and then sending a public update to the full team. It let me see the real progress behind all the drama and back-to-back meetings and stay transparent with everyone. Next year, I’ll avoid meetings before 9 AM, keep an eye on calendar creep, and hold myself accountable to exercise and doing what helps me decompress. I know, it’s obvious. I also plan to travel more. Especially to the Bay Area, which is clearly back again. Meeting up with other founders and customers is always worthwhile. Another thing that made last year tough was having half my direct reports on parental leave for about half of the year. Now they’re back, and I can feel the momentum returning. Jacob has also taken over product again, now that he’s stepped away from directly owning operations and people. Increased shipping velocity has been noticeable. We have all the pieces in place. All that’s left is to keep pushing forward: shipping, selling and enjoying the ride. If there’s one thing I learned in Silicon Valley, it’s that no goal is too crazy if you refuse to give up. I really hope you enjoyed reading this post. As always, my intention was to share it with complete honesty and transparency, avoiding the hype that often surrounds startups. If you are facing similar challenges and want to connect and share experiences, please do not hesitate to reach out on Twitter or shoot me an email! Special thanks to my co-founder Jacob, my EA Susannah, the whole RevenueCat team, and our valuable customers. I also need to express my eternal gratitude to all the CTOs and leaders who have been kind enough to share their experiences over these years. Shoutout to Dani Lopez, Peter Silberman, Alex Plugaru, Kwindla Hultman Kramer, João Batalha, Karri Saarinen, Miguel Martinez Triviño, Javi Santana, Matias Woloski, Tobias Balling, Jason Warner, and Will Larson. Our investors and early believers Jason Lemkin, Anu Hariharan, Mark Fiorentino, Mark Goldberg, Andrew Maguire, Gustaf Alströmer, Sofia Dolfie, and Nico Wittenborn. I want to convey my deep gratitude to my amazing wife, Marina, who has been my unwavering source of inspiration and support from the very beginning, and for blessing us with our two precious daughters. I cannot close this post without thanking my mom, who made countless sacrifices to mold me into the person I am today. I promise you will look down on us with pride the day we ring the bell in New York. I love you dearly.

6th Jan 2025 • 88 votes
Full Circle

I’m back in Spain for my brother’s wedding. I rarely visit during the summer. The heat in my hometown is brutal, around 40 degrees Celsius (over 100 Fahrenheit for my imperial friends). Most people escape to the coast, just like my family did when I was a kid. I haven’t been here in years. As I drive along the coast, I find myself reflecting on a tweet about money and happiness, a vivid memory pulls me back in time. It’s August 22nd, 2007. The iPhone, the first real smartphone, has just been announced. It’s so cool, but of course I cannot afford it. It’s not even going to be released in Spain. I’ve just gotten my driver’s license, and I’m about to dive into my third year of Computer Science. I set up my clunky TomTom navigator knockoff, and hit the road. I’m on my way to meet Marina for our first real date. She’s cool, pretty, and kind. She likes the same music as I do, and even has distant relatives in California, the place we jokingly plan to visit someday (if I ever get enough cash). I’m listening to a pirated Blink-182’s self-titled CD. Pop-punk is pretty niche in the south of Spain, and it’s dying. Blink-182 has split up, and I missed my window to see my favorite band live. The Atlantic Ocean is as flat as a lake. This corner of Huelva’s coast is sheltered from any real waves, a stark contrast to the world-class surf breaks I drool over in magazines. And suddenly, reality hits: my childhood dreams of building a tech company in Silicon Valley, while vacationing in Southern California feel impossibly far. Even getting through my degree feels like a pipe dream. School isn’t fun anymore. It’s grueling, especially the parts I thought I’d enjoy, like algorithms and Data Structures. I might never become a Software Engineer. I feel stuck, trapped by my lack of direction. I am seriously considering quitting. But no degree means no job in the US, and good tech gigs are rare here in Spain. The only cool company is Tuenti, a new startup that is cloning Facebook. I’m nowhere near smart enough to land a job there though. Flash forward to today, 17 years later. It’s almost laughable to think about how hopeless things once seemed. Even now, it doesn’t feel like I’ve “made it.” The path and the results look nothing like what teenage me envisioned, but somehow, I’m realizing I’ve kind of checked off every box. I married Marina, and we live in Southern California with our two beautiful identical kids. We’ve become American citizens, and I’ve lived in the Golden State for nearly a third of my life. I’ve worked as a Software Engineer at a Silicon Valley startup, learned from the best, and found the best co-founder I could ask for. We launched our own company. Smartphones? They’re in everyone’s pockets now. Our product is in a third of all new apps shipped in the US. We’ve helped developers reach millionaire status, and we’ve made more money than I ever thought possible. But I’ve learned that a lot of money is a relative term. Somehow, I managed to hire insanely talented engineers—a bunch of them, ironically, from Tuenti. Blink-182 is back together, and I’ve been fortunate enough to see them live five times. I’ve even bumped into Tom Delonge after surfing world-class waves a few times. I’m literally just realizing how surreal all of this is. I tend to get caught up in the chaos of what’s next—the next big fire, the next goal—but sometimes you’ve got to stop, be present, and reflect on how far you’ve come. It wasn’t easy. It wasn’t without loss, sacrifice, and a fair share of doubts. Am I truly happy? Maybe not in a perfect, all-the-time kind of way. There are external things humans cannot control. But when I look at my life, I realize there’s no real reason not to be. The journey has been was worth it so far: the ups, the downs, the unexpected turns. So, here’s to your journey, whatever it looks like. Keep going, keep dreaming. It might not turn out the way you envisioned it, but it’s only impossible if you quit.

28th Aug 2024 • 94 votes

More in programming

Clip of me singing Despard in Ruddigore in 2013

A clip of me singing a funny song from Gilbert and Sullivan’s Ruddigore back in 2013

6 hours ago • 1 votes
How and Why fork() Uses Copy-on-Write

In this video, we look at why fork() needs copy-on-write, how it works inside the kernel, and a memory usage problem that Instagram encountered with Python.

12 hours ago • 1 votes
What we lost when we lost comments

Comments require commitment, but they’re worth it.

18 hours ago • 1 votes
Lighthouse map

Lovely global map with animated lights sweeping the waters

20 hours ago • 1 votes
Warming up the Puma master before it forks

Basecamp 5 runs on Puma in cluster mode: one master process with preload_app! and 63 single-threaded workers per host, deployed as a Docker container with Kamal. We serve Basecamp from several sites. Each site has its own web hosts and a read replica of the database, and writes go to a single primary database in one of them. On our busiest hosts, each deploy left up to 2,000 requests waiting while the new workers warmed up. We reduced those queues by running signed-in requests through the app in the Puma master, before it forked the workers. Why 63 single-threaded workers? Basecamp has always served web requests from processes rather than threads. It ran on Unicorn, which only does processes, until we moved to Puma in January 2025, and we kept the same setup: workers (Concurrent.physical_processor_count * 1.3).ceil threads 1, 1 preload_app! On a 48-core host that’s 63 workers, each handling one request at a time. We chose 1.3 after benchmarking HEY in 2023, when we moved our apps out of the cloud and onto our own hardware. We tested several combinations of workers and threads with a mix of GET and POST requests on a 32-vCPU VM. Every multithreaded configuration we tested was slower and handled fewer requests than single-threaded workers. Adding workers beyond about 1.2 to 1.3 per vCPU brought little benefit. The threaded workers spent a lot of their time waiting for Ruby’s global VM lock. That made single-threaded workers a good fit for this workload, and we use the same setup for Basecamp. An app that spends more time waiting on its database or other services may benefit from more threads, so benchmark your own app. The other reason is the app itself. Basecamp has class-level state in places and has never needed to be thread-safe. With one request per process, it still doesn’t. Processes do use more memory than threads, and preload_app! reduces the difference. The master loads the app once and the workers share its memory through copy-on-write until they write to it. Shopify’s comparison of Ruby execution models explains the trade-off well. In the HEY benchmark the best setup came to about 260 MB of PSS per core, where PSS counts each shared page once, split between the processes using it, and the gap to a threaded setup was smaller than we’d expected. What Puma does on each host when a container starts: one master, then 63 forked workers that share its memory until they write to it. Two things about this setup matter for the rest of the post. A worker that’s compiling or loading something is fully blocked — there’s no other thread to pick up the next request. And whatever the master has in memory before it forks, all 63 workers share. Whatever they build after the fork, they build 63 times. What happens when we deploy Kamal starts the new container alongside the old one, and kamal-proxy moves the host’s traffic across as soon as the health check passes. At that moment, the new workers have handled health checks but no customer requests. preload_app! means the master loads the app once and the workers inherit it through fork. That covers the code. It doesn’t cover anything Ruby and Rails set up on first use: YJIT compiled code. YJIT compiles a method once it’s been called a certain number of times. The master calls very little during boot, so every worker compiles the same methods again on its own first requests. Compiled templates. Action View turns each ERB template into a Ruby method the first time it’s rendered. The schema cache. Active Record reads each model’s columns from the database the first time that model is used. Inline caches and memoized values throughout Ruby, Rails and the app. All 63 workers did all of this at once, while serving the traffic the old container had been handling a second earlier. In the test environment with YJIT on, the first request to a project page on a cold process took 652 ms, 151 ms of it YJIT compiling. The same request to a warm process took 28 ms. In production, CPU time per request peaked at around 200 ms while kamal-proxy moved traffic to the new container, against about 30 ms once the workers had warmed up. A host with spare CPU absorbs this. Every one of our web hosts has 48 cores and 63 workers, but each Amsterdam host serves around 250 requests per second, against 25 to 60 at our other sites. In Amsterdam the slow first requests turned into a queue. At a peak-hour deploy, the Puma backlog on an Amsterdam host reached anywhere from 250 to 2,238 requests, and kamal-proxy’s p99 response time hit about 10 seconds. Eron, our Director of Operations, had been tracking this since June. Another server in Amsterdam would help, but it would take weeks to arrive, so we also wanted to make deploys cheaper on the hardware we already had. What didn’t work We tried a few things first. In June, Donal tested the first two on a single Amsterdam host, comparing it with its neighbors, and they ruled out two likely causes. Warming each worker’s database connections. Puma’s before_fork hook clears the master’s connections, and each worker opened its own on its first request. Opening them in before_worker_boot instead made no difference. Queries on a freshly booted production host were already under a millisecond, so connections weren’t the problem. A synthetic request in each worker. Next, each worker made a few requests in before_worker_boot to an internal controller that touched every model. That ran the middleware, routing and Active Record paths, but it ran them in 63 workers at once — exactly the CPU spike we were trying to avoid. And a request with no real data renders no real views, so most of the app stayed cold. Spreading YJIT compilation out. Delaying YJIT in each worker by a random interval spread the compiling out over a few minutes, but every worker still ran interpreted until its delay ended. The queue didn’t change. Reforking from a warm worker. This is what Shopify’s Pitchfork does: let one worker serve traffic until it’s warm, then fork the others from it. Puma has an experimental version called fork_worker, and on beta it worked — the reforked workers were warm after three to five requests, where fresh ones took up to 30 seconds. But with fork_worker the template is worker 0, and it keeps serving requests. If it exits, the workers waiting to be forked never start (puma/puma#3596). If it gets no traffic, the refork never happens, which is what we saw on beta. Instacart have a mold_worker patch that promotes a warm worker to a template that stops serving, but it isn’t in a Puma release. We have a branch of it, and we may come back to it. That last experiment did show us where the fix was, though. Everything a warm worker has that a cold one lacks is in its memory, and fork copies memory. The master already has the app loaded. It just never runs it. Run the requests in the master So now, before the master binds its socket and forks, it makes the app’s own requests, in-process, the way a signed-in user would. Rack has a hook for exactly this. Rack::Builder#warmup takes a block that’s called once with the built app, before the server starts. rails server builds the app from config.ru, so the change to boot is one line: require_relative "config/environment" warmup { WarmUp.configured.run } if ENV["WARM_UP"] run Rails.application With preload_app! this runs in the master, and the workers inherit whatever it did. Puma binds its socket after the app is built, so until the warm-up finishes the health check’s connection is refused and kamal-proxy keeps retrying. No request reaches a worker that hasn’t been warmed. The warm-up has three steps. After precompiling the views, it gives the page requests and schema loading a shared 20-second budget, checked before each page or model. 1. Precompile the views actionview_precompiler reads every template for its render calls and compiles each one with the locals it’s passed. For us that’s 1,394 templates in about two seconds. A first request to a project page then compiles 2 templates instead of 44. 2. Request the pages, signed in A small browser class makes the requests through Rack::MockRequest, with the two cookies a real sign-in sets, then goes back for each page’s lazy Turbo frames: class WarmUp::Browser def initialize(signed_in_as:) @client = Rack::MockRequest.new(Rails.application) @headers = { "HTTP_USER_AGENT" => "Basecamp warm-up", "HTTP_COOKIE" => cookie_for(signed_in_as), "bc3.warm_up" => true } end def visit(path) page = get(path) frames_in(page).each { |id, src| get(src, "HTTP_TURBO_FRAME" => id) } end private def get(path, headers = {}) @client.get("https://#{host}#{path}", @headers.merge(headers)) end def frames_in(page) Nokogiri::HTML5(page.body).css("turbo-frame[src]").map { |frame| [ frame["id"], frame["src"] ] } end end The requests are signed in. The user is a monitoring account we already use for automated checks, and the pages are its own project, Campfire, to-dos, documents and messages. Public pages weren’t enough: after warming up with signed-out pages only, the first signed-in request to the projects page still took 131 ms, because authentication, the signed-in controllers and their views had never run. With signed-in pages it took 40 ms. cookie_for writes the same signed cookie the sign-in controller does, using the app’s own cookie jar, so there’s no API token and no secret to store. The frames are followed. The busiest HTML requests in production aren’t pages at all but Turbo frames — the sidebar badge, the inbox, the navigation menus. The browser parses each page and requests its <turbo-frame src> URLs with the Turbo-Frame header, so those controllers and views get warmed too. Our first four pages turned into 60 requests. The requests are excluded from rate limiting. They are internal, so they do not count against the rate limits that apply to real visitors. 3. Load the rest of the schema The page requests load the schema for the models they touch. The last step loads the rest, from the read replica: ApplicationRecord.reading do models.lazy.take_while { time_left? }.each { |model| model.load_schema if model.table_exists? } end The step checks 261 models and loads any schema information still missing. Those database round trips add up when the primary is far away: outside a request, Active Record uses the writing role, and from a host a long way from the primary each round trip is tens of milliseconds. Reading from the local replica brings the step down from about 20 seconds to 3.5. The pages go first because they load most of the schema anyway. If the time budget runs out, the step stops, logs how many models it got through, and the workers load the rest on first use like they always did. Rails can also load the schema from a dumped cache file at boot (bin/rails db:schema:cache:dump), which would make this step unnecessary. We don’t ship one in our image yet, because the dump needs a database to read from at build time, and we have several databases to cover. It’s on the list. What to close before the fork Running requests in the master opens things the master never opened before, and every worker inherits them. Two processes writing to the same socket will corrupt each other’s traffic, so you need to know what’s open before you fork. The way to find out is to list the master’s open file descriptors — ls -l /proc/<pid>/fd — before and after a warm-up, in an environment set up like production. Development wasn’t enough for us: it stores files on disk, so our S3 connections only showed up in production. Then, for each thing that’s open, check how its library handles a fork. We found three kinds: Already handled. Plenty of libraries detect a fork on their own, either by recording the PID they connected from and reconnecting in the child, by opening per-process files, or by resetting their thread pools. Redis clients, metrics libraries and concurrency libraries tend to be in this group. Check, but you probably don’t need to do anything. Already closed. Database connections are the classic one, and most Puma configs already clear them in before_fork. Anything else that’s opened per process — we have a SQLite cache the workers open on boot — needs closing when the warm-up finishes. Needs a new step. HTTP clients with keep-alive connections are the ones to look for: cloud SDKs with connection pools, tracing exporters, error reporters. They usually have no fork handling at all. We empty the aws-sdk connection pools in before_fork, and we run the warm-up untraced so the OpenTelemetry exporter never opens its connection to Tempo in the first place. Once that’s done, before_fork finishes with Process.warmup, which Ruby 3.3 added for this purpose: a major GC, a heap compaction, and every surviving object promoted to the old generation, so the memory pages the workers share change as little as possible afterwards. Choosing the pages The first list was the four pages that ran the busiest requests on beta. Once the warm-up was live, production showed us which endpoints were still cold. For one deploy, we compared each endpoint’s mean duration in the six minutes after kamal-proxy moved traffic to the new container with the same endpoint an hour later, then multiplied the difference by the number of requests in those six minutes. That gives the extra time each endpoint cost us because it was cold: Endpoint Cold Warm Requests in 6 min Extra seconds Campfire 246 ms 70 ms 6,490 1,140 Projects (JSON API) 84 ms 50 ms 22,077 771 Docs & Files 262 ms 177 ms 4,996 421 To-dos tool 205 ms 113 ms 4,018 371 To-dos (JSON API) 33 ms 16 ms 18,738 320 The pages already in the warm-up showed what to expect: the project page kept a 36 ms gap after a deploy, and the to-do page 10 ms. We’ve proposed adding these five requests, and expect them to add about five to seven seconds to the page step. The two JSON endpoints were a surprise. The warm-up’s page list had no API requests in it, so nothing on the API path had run before the first real request: not the API controllers, and not the Jbuilder templates rendering real records. Precompiling the views covers JSON templates too, but it isn’t a substitute for running the request. Results The warm-up is on for all 68 web hosts. With the first four pages it took 12 to 16 seconds per host: about 2 seconds to precompile the views, 7 to 9 for the 60 requests, and 3.5 for the schema. Deploys take that much longer per host, and we raised the deploy timeout from 30 to 60 seconds to cover it. In Amsterdam, at a peak-hour deploy: During deploy Before After Peak Puma backlog per host 250–2,238 requests 19–223 requests Peak kamal-proxy p99 about 10 s 2.4–4.8 s Peak CPU time per request 201–214 ms 88–132 ms Peak database time per request 56–69 ms 39–47 ms The same eight hosts at three deploys on 1 October, an hour apart, as the warm-up went from one host to four to all eight. The deploy in the middle, with four hosts warmed and four not, shows why every host needed the warm-up. Each warmed host recovered faster on its own: mean request duration peaked at 130 to 173 ms, against 203 to 311 ms on the hosts that weren’t warmed. But the backlog on all eight was about the same, because they were all waiting on the same database. Mean request duration on each host at the 07:21 UTC deploy. Blue hosts warmed up in the master before forking, orange hosts did not. Memory came down too. The workers now share compiled templates, YJIT code and the schema with the master instead of each building their own copy. On beta, the view precompiler alone took a busy worker’s private memory from 174–202 MB to 119–135 MB. Thirty minutes after the deploy, the web containers used about 39 GB less memory than the previous day’s containers at the same age and traffic. Amsterdam served most of our traffic at the times we tested. In Amsterdam, each new container used about 2 GB less just after traffic moved to it, which lowers the peak while the old and new containers overlap. Working with Claude Claude Code helped throughout. It combed through the per-worker backlogs and per-endpoint timings in Prometheus and Loki after each deploy, worked out the cold-versus-warm cost of each endpoint, and prepared the changes and the pull request descriptions with the benchmarks in them. We decided what to try, deployed it and read the results. If you do this Warm the master before it forks. Compile common code and templates and load their schema in the master, so workers inherit that work. With preload_app!, Rack::Builder#warmup runs before the workers start accepting traffic. Use the app’s real requests. Public pages, internal endpoints and synthetic queries warm the paths they run and nothing else. Signed-in requests to real records, frames included, run what production runs. Measure the cold penalty per endpoint. The difference between an endpoint’s cold and warm duration, times its request count after a deploy, ranks the pages worth adding. Ours weren’t the ones we’d have guessed, and two of them were JSON. Check what the warm-up leaves open. List the master’s file descriptors after a warm-up and account for every one before the fork. Two of ours needed changes. Set a time budget. A warm-up that runs long on one slow host fails the deploy on that host. Ours gives the page requests and schema loading a shared 20-second budget, checked before each page or model, puts the most valuable pages first, and logs what it skipped. Reforking from a warm worker, as Pitchfork does, solves the same problem continuously rather than once at boot, and it would warm paths no fixed list of pages covers. We may still get there: our branch brings Instacart’s mold_worker up to date with Puma’s main branch and fixes the bugs we found in it. But warming the master works with the Puma we already run, took a few days to implement, and substantially reduced the queues after deployment.

yesterday • 1 votes
📚 BoredReading

You seem to be enjoying this.

Join free to unlock everything.

Create free account

Already have an account? Sign in