Full Width [alt+shift+f] Shortcuts [alt+shift+k]
Sign Up [alt+shift+s] Log In [alt+shift+l]
1

Top 20 Helpdesk Interview Questions (with answers)

from Founder's blog [alt+shift+b] in programming

Help desk interviews usually test two things at once: whether you understand the technical basics, and whether you can explain those basics to a person who may already be annoyed, confused, or late for a meeting. That combination matters. A support technician who can diagnose a network issue but makes the user feel stupid will struggle. So will a friendly person who guesses randomly at technical problems. The best candidates show both: methodical troubleshooting and calm, useful communication. Below are common help desk and desktop support interview questions, with sample answers you can adapt to your own experience. Technical Help Desk Interview Questions 1. Can you tell me about yourself? Keep the answer focused on the job. Mention your IT training, support experience, certifications, customer service background, and the kind of technical problems you have handled. Avoid turning it into a life story. A good answer gives the interviewer several useful follow-up paths. For example: "I have been building my support skills through Windows troubleshooting, networking fundamentals, and customer-facing work. I enjoy breaking problems down, documenting what I find, and helping users get back to work without making the process more stressful for them." 2. A user says their monitor is not working. What do you check first? Start with the simple physical checks before assuming a complicated failure. Confirm that the monitor has power, the brightness is not turned all the way down, the video cable is connected securely, and the computer itself is powered on. If the monitor has multiple inputs, make sure the correct input is selected. If those checks do not solve it, continue with another cable, another monitor, or another port. From there, you can investigate graphics drivers, docking stations, sleep state problems, or hardware failure. 3. What is Safe Mode used for? Safe Mode starts Windows with a limited set of drivers and services. It is useful when a normal startup fails,...
8th Mar 2026

Stay updated

Get a weekly newsletter with the top 5 articles worth reading every week.

More from Founder's blog

Best AI helpdesk solutions in 2026

Every helpdesk vendor now sells "AI". What they actually ship differs a lot. Some offer a chatbot that answers customers before a ticket exists. Some offer a copilot that drafts replies for agents. Others hand you a platform and a consultant's invoice. The pricing models differ too: per seat, per conversation, per "resolution", or a pool of credits that nobody on the buying committee fully understands. We looked at six options, from a free do-it-yourself build to a full enterprise ITSM suite. Each one gets the same treatment: what it does well, where it falls short, and what it costs as of September 2026. Disclosure: this review is published on Jitbit's blog, and Jitbit is one of the products below. We tried to hold it to the same standard as the others, and we list its weaknesses too. Competitor prices come from public pricing pages and recent third-party breakdowns. Vendors change prices often, so check before you sign. How we evaluated Grounding. Does the AI answer from your documentation and ticket history, or from general model knowledge? Coverage. Does it deflect tickets (customer-facing), help agents (agent-facing), or both? Cost predictability. Can you forecast next quarter's bill without a spreadsheet model? Time to value. Days, weeks, or a multi-month implementation project? 1. Homegrown & Free: a Copilot Studio agent grounded on SharePoint Best for: internal IT support in organizations that already run on Microsoft 365. This option gets overlooked because nobody sells it. If your company already lives in Microsoft 365, you can build a decent self-service IT assistant in an afternoon without buying a helpdesk. The recipe is simple: put your end-user documentation on one SharePoint site, point a Copilot Studio agent at it, and publish the agent to Teams, where your employees already spend the day. Here is how to do it well. How to set it up Create a dedicated SharePoint communication site, for example "IT Help". Keep the knowledge the agent will use on this one site, not scattered across team sites. A single, clean source gives better answers than a large, messy one. Write one page per problem. "Reset your VPN token", "Printer on floor 3 shows offline", "Request a new laptop". Use the error message or symptom a user would type as the page title, and put the fix in numbered steps. Retrieval works on chunks of text, so short focused pages beat a 60-page PDF manual. Avoid scanned PDFs and screenshots-only pages, because the agent cannot read text inside images reliably. Check permissions. The agent answers using the identity of the person asking, so it only sees what that person can see in SharePoint. Give all employees read access to the site. Keep admin-only runbooks on a separate site so they never leak into employee answers. Create the agent in Copilot Studio (copilotstudio.microsoft.com). Add the SharePoint site URL as a knowledge source, and keep authentication set to "Authenticate with Microsoft", which SharePoint knowledge requires. Turn off general knowledge. In the agent's generative AI settings, disable the option that lets the model answer from its own general knowledge. You want "I don't know" rather than a confident, invented fix for your internal VPN. Write short instructions. For example: "You are the IT help assistant for Contoso employees. Answer only from the IT Help site. Give numbered steps. If the answer is not in the documentation, say so and link to the ticket form." Add an escalation path. Change the fallback topic so unanswered questions end with a link to your ticket form or IT mailbox. More advanced teams add an agent flow that opens a ticket directly, but each tool call uses extra credits. Publish to Teams and Microsoft 365 Copilot, then pin the agent in Teams for all employees through the Teams admin center. If people have to look for the agent, adoption stays low. Review analytics every week. Copilot Studio shows which questions went unanswered. Each one is a missing SharePoint page. Write it, and the agent improves by the next morning. This loop matters more than anything else on this list. Pros. There is no new vendor, no procurement cycle, and no new login. Employees ask questions in Teams, where they already are. Identity, permissions and compliance come from your existing Microsoft tenant, so security review is short. The knowledge base is ordinary SharePoint pages that anyone in IT can edit, with no proprietary format to migrate later. For the large share of internal tickets that are really "how do I..." questions (password resets, VPN, printers, Wi-Fi, software requests), a well-written site plus a grounded agent can deflect a meaningful share of volume in the first month. Cons. This is a self-service answer engine, not a helpdesk. There is no ticket queue, no SLA tracking, no assignment, no reporting on agent workload, and no audit trail of who fixed what. When the bot can't help, the issue goes to email or a form, and you still need something to manage it. Answer quality depends entirely on the documentation. A thin or outdated SharePoint site produces thin or outdated answers. Copilot Studio's credit-based licensing is also hard to forecast, and agents can be disabled automatically once a tenant goes 25% over its prepaid capacity. It works for employees only. Do not try to turn it into a customer-facing support bot. Price. Effectively free if your users already have Microsoft 365 Copilot licenses. Microsoft does not charge credits when licensed users talk to employee-facing agents, within fair-use limits. Without those licenses, Copilot Studio bills in credits: $200 per month for a pack of 25,000, or $0.01 per credit pay-as-you-go through Azure. A generative answer grounded on a knowledge source costs 2 credits. With the optional tenant-graph grounding it costs 12. In practice, that is roughly two to twelve cents per answer, and a few hundred dollars a month covers a mid-sized company's internal IT questions. The real cost is the staff time to write and maintain the SharePoint pages. You would need that documentation with any tool on this list anyway. 2. Jitbit Helpdesk Best for: mid-sized support and IT teams that want a full ticketing system with AI included, not billed separately. Pros. AI is built into the ticketing workflow, not added as a separate product. An AI assistant panel in every ticket drafts replies grounded on your knowledge base, indexed external documentation, canned responses, and similar closed tickets. It also summarizes long threads, cleans up agent drafts, and turns solved tickets into new KB articles, which helps teams that are just starting a knowledge base. Automation rules can run AI triage on every incoming ticket (sentiment, category and priority in one pass) and can post first-line AI replies either directly to the customer or as internal notes for review. The live chat widget has an AI auto-responder that hands off to a human when the visitor asks for one. For technical teams, the AI can call your own HTTP endpoints or MCP servers, so it can look up an order status or a subscription without an agent. Model choice is unusually open: GPT and Gemini are included, or you can bring your own key for Claude, Azure OpenAI or AWS Bedrock. That matters for data-residency and HIPAA reviews, and Jitbit signs a BAA on its Enterprise plan. It is also one of the few products here that still offers a self-hosted, on-premise edition with the same AI features. Cons. Jitbit is a smaller vendor, and its integration marketplace is much smaller than Freshdesk's or ServiceNow's. Its customer-facing bot is a helpful extra on top of a helpdesk, not a dedicated conversational platform. Teams with very high chat volume in B2C, which is Fin's specialty, will find fewer controls for multi-step procedures and channel orchestration. Like every grounded AI, it only performs as well as the knowledge base behind it. On-premise customers who want semantic search and document indexing must run an extra Docker-based add-on, and they supply their own model API keys. The entry-level Freelancer plan does not include AI credits. Price. Flat plans with no per-resolution or per-session fees: Freelancer $29/month (1 agent), Startup $69/month (4 agents), Company $129/month (7 agents), Enterprise $249/month (9 agents, $29 per extra agent). Annual billing saves about two months. Unlimited AI credits are included from the Startup plan up. For a seven-agent team, that works out to about $18 per agent per month with AI included. On most of the other paid products here, AI alone costs more than that per agent. 3. Fin (by Intercom) Best for: high-volume customer-facing support where deflection rate is the main metric. Pros. Fin is arguably the most capable customer-facing AI agent on the market, and it is the product the rest of the category measures itself against. It handles multi-turn conversations well, follows written "procedures" for multi-step tasks such as refunds or account changes, and works across chat, email, and voice. Importantly, Fin does not require Intercom's helpdesk. It can run on top of Zendesk, Salesforce, Freshdesk, Front, and others, so you can add it to an existing stack without migrating. Its reporting is built around resolution rate, which makes the ROI case easy to present to a CFO. Cons. Outcome-based pricing is fair in theory and hard to budget in practice. Fin counts an "assumed resolution" when a customer simply stops replying after its last answer. That includes customers who gave up. Your bill grows with volume, not with headcount, so a seasonal spike or a product incident that floods support also increases your AI spend. Fin is built for customer-facing deflection. Agent-side features are stronger inside Intercom's own inbox, so standalone deployments on other helpdesks get less of the product. For internal IT support, it is the wrong tool. Price. $0.99 per billable outcome (a resolution, a procedure handoff or a disqualification), and $9.99 per qualified sales lead. Running standalone on another helpdesk has a minimum of 50 outcomes a month. Inside Intercom, you also pay for seats: Essential $29, Advanced $85, or Expert $132 per seat per month on annual billing. A ten-seat Advanced team resolving 2,000 conversations a month pays roughly $2,800, and about 70% of that is Fin. 4. Tidio (with Lyro AI) Best for: small e-commerce stores that mainly need a website chatbot. Pros. Tidio is one of the fastest ways to get an AI chatbot onto a website. Lyro, its AI agent, learns from your FAQ and site content in minutes, and setup takes almost no technical skill. It integrates well with Shopify, WordPress, and the common e-commerce stack, so it can answer order and shipping questions that make up most of a small shop's inbox. The visual Flows builder handles cart-abandonment nudges and lead capture alongside support. For a one- or two-person store, the free tier is enough to try it before paying. Cons. Tidio is chat-first, and the ticketing side is basic. There is little in the way of SLAs, complex routing, asset management, or the reporting a growing support or IT team eventually needs. Lyro is billed by conversation, and the free and Starter plans include only 50 Lyro conversations in total, not per month. After that, it stops answering. The plan structure has a big gap: the jump from Growth to Plus is more than tenfold, with no middle tier. Teams often report that real costs run two to three times the advertised price once AI and Flows add-ons are included. It is not designed for B2B, internal IT, or regulated industries. Price. Starter $29/month, Growth from $59/month (scaling to about $349 depending on volume), Plus $749/month, and Premium from $2,999/month. Lyro AI is a separate add-on starting at $39/month, and Flows costs another $29/month. A typical Growth setup with Lyro and Flows costs around $127/month before overages. Annual billing saves about 17%. 5. Freshdesk (with Freddy AI) Best for: mid-market customer support teams that want a mainstream, feature-complete suite and can live with add-on pricing. Pros. Freshdesk is a mature, full-featured helpdesk with omnichannel support, a large app marketplace, solid automation, and a familiar interface that new agents learn quickly. Freddy AI covers both sides of the job. The Freddy AI Agent deflects questions on chat, email and the web widget. Freddy Copilot helps human agents with reply drafts, summaries, tone adjustment and suggested solutions. Freshworks also sells Freshservice for internal IT, so organizations can standardize on one vendor for both customer and employee support. Cons. AI costs are split across several meters and plan requirements. Copilot is available only on Pro and Enterprise, so Growth customers cannot buy it at any price. The customer-facing AI Agent is billed by session, and unused sessions expire at the end of each billing cycle instead of rolling over. Freshworks raised plan prices in 2026 for the first time in five years and cut the free plan to a six-month program for one or two agents. Model choice and on-premise deployment are not available. Price. Growth $19, Pro $55, Enterprise $89 per agent per month (annual billing). Freddy Copilot adds $29 per agent per month ($35 on monthly billing). The AI Agent includes 500 sessions, then costs $49 per 100 sessions. The realistic entry point for agent-side AI is $84 per agent per month (Pro plus Copilot). A five-agent team on Pro with Copilot and moderate bot usage spends roughly $650-700 a month. 6. ServiceNow (with Now Assist) Best for: large enterprises running IT, HR and facilities service management on one platform. Pros. ServiceNow is the default choice for enterprise ITSM, and its AI benefits from being close to all of that data. Now Assist summarizes incidents, drafts resolution notes, generates KB articles from resolved cases, and powers a Virtual Agent that can complete requests end to end instead of just answering questions. Because it sits on the same platform as the CMDB, change management, and HR and facilities workflows, the AI can act on real records: reset an account, provision software, or open a change request. No other product here offers that level of process depth, governance, and audit trail. In 2026 ServiceNow moved to AI-native tiers, so AI is now included by default rather than sold as a separate SKU. Cons. Cost and complexity. Implementations usually take months and a certified partner, and the platform typically needs dedicated administrators after go-live. Pricing is quote-only and heavily negotiated, so two companies of the same size can pay very different amounts. AI usage now draws from consumption-based "Assist" pools, which adds another capacity figure to forecast and renegotiate. For organizations under about a thousand employees, ServiceNow is usually more platform than they need. Price. There is no public price list. In April 2026, ServiceNow replaced its Standard/Pro/Pro Plus/Enterprise tiers with Foundation, Advanced and Prime, and older SKUs reached end of sale in mid-2026. As a reference point, legacy ITSM Pro licenses were widely reported at around $100-135 per fulfiller per month. Adding Now Assist (Pro Plus) raised that by an estimated 50-60%, to roughly $200-215 per fulfiller per month. Minimum seat commitments, platform fees and implementation services often make the first-year total a six-figure amount. Side-by-side Solution Best fit Pricing model AI included? Copilot Studio + SharePoint Internal IT, Microsoft shops Credits, or free with M365 Copilot licenses Yes (it's only AI, no ticketing) Jitbit SMB and mid-market support and IT Flat monthly plans Yes, unlimited from Startup plan Fin High-volume customer support $0.99 per outcome, plus seats in Intercom It is the product Tidio Small e-commerce Plan plus Lyro conversation add-on Add-on from $39/month Freshdesk Mid-market customer support Per agent, plus Copilot seats and bot sessions Add-on, $29/agent plus sessions ServiceNow Large-enterprise ITSM Quote-only, per fulfiller, plus Assist pools Bundled in 2026 tiers The bottom line Start with the question you're really trying to answer. If you want employees to stop emailing IT about VPN passwords and you already pay Microsoft, build the SharePoint and Copilot Studio agent first. It costs almost nothing, and the documentation you write for it will make any helpdesk you buy later work better. If you need an actual ticketing system with AI included and a predictable bill, Jitbit and Freshdesk are the main choices, and they differ mostly in how AI is charged. If deflecting a large volume of customer chats is the whole goal, Fin sets the standard, as long as you are comfortable with a bill that grows with volume. Tidio fits small online stores. ServiceNow fits organizations large enough to need a service-management platform, with a budget to match. Whatever you choose, AI quality comes from your knowledge base, not the model. Every product here gives better answers with better documentation, so start writing it now.

27th Jun 2026 • 1 votes
Jitbit Helpdesk Bot is now in the Microsoft Teams Store

Our new Microsoft Teams app just passed Microsoft's certification and is live in the Teams Store: Jitbit Helpdesk Bot. Search "Jitbit" in the Teams app store and it's right there - no custom app uploads, no admin-center gymnastics. Unlike our old webhook integration (one-way channel notifications, now deprecated), this is a proper two-way bot: End users create tickets in chat - type new and get a real form with category and priority, or turn any Teams message into a ticket from the "..." message menu. my tickets lists your open requests Technician replies come back to Teams - as cards with an inline answer box, so the whole conversation can happen without opening the helpdesk Channels get actionable ticket cards - @mention the bot, type subscribe, and every new ticket shows up with Take it / Reply / Close buttons. Cards update in place when the ticket changes - whether that happens in Teams or in the web app Setup takes about two minutes: install the app, click "Generate pairing command" on the helpdesk admin page, paste the command into a chat with the bot. Users are matched by their Microsoft 365 email automatically. Self-hosted customers run their own private bot, so ticket traffic never touches our servers. The bot is included with all plans. Details and setup steps: the integration page and the documentation.

8th May 2026 • 1 votes
AI features now run on-premise

Short version: every AI feature we ship on the hosted helpdesk now runs on the self-hosted edition too. It's a separately licensed Docker add-on, perpetual license, one-time payment. Customer data stays on your network. Requires Jitbit Helpdesk v11.22 or newer. Why this took a while Our AI stack isn't a thin wrapper around an LLM API. It's a Python service running a local vector database (Qdrant), embedding models, and a RAG pipeline tuned against support-desk content. That's what makes "similar KB articles" surface the right article, and what keeps reply drafts grounded in your own documentation instead of hallucinated. Running that stack on our own hosting is one thing. Packaging it so your IT team can stand it up on your own hardware without babysitting Python dependencies is another. We sat on it until we had something we'd be comfortable supporting. What you get Similar-article suggestions inside tickets - ranked by semantic similarity, not keyword overlap Semantic KB search for end-users - results ranked by meaning AI-generated reply drafts grounded in your KB, writing-style rules, and the ticket context External documentation indexing - crawl your own docs, wiki, or any internal site and use it as AI context alongside the KB Choice of embedding model - free local model that runs on CPU, or OpenAI embeddings with your own key Generative provider - OpenAI, Azure OpenAI, or AWS Bedrock, customer brings their own key Feature parity with SaaS. No caveats about "basic" ChatGPT integration. How it ships A Docker Compose stack you drop in next to your existing Helpdesk install. Runs on Windows or Linux, bare metal or VM, Intel/AMD or ARM. Upgrades are zero-downtime and preserve Docker volumes — indexed data and cached models carry over. Full setup and system requirements are in the manual. Privacy This is the reason we finally built this. With the bundled local embedding model, nothing about your tickets, KB articles, or indexed documentation leaves your network. No embeddings sent to OpenAI. No vectors stored in a third-party service. The vector DB is a container on a host you own. If you want a generative model for reply drafts, you bring your own API key. Azure OpenAI and AWS Bedrock keep everything inside your existing cloud tenancy, with a BAA if you need one. For regulated on-prem buyers — healthcare, defense, financial services, government — this was the #1 reason you told us you couldn't adopt our AI features. It's no longer a reason. Pricing Licensed separately from the core Helpdesk product. Perpetual license, one-time payment, 1 year of updates included - same model as the rest of the on-prem lineup. No subscription, no per-agent add-on, no per-request fees. See the pricing page for the current price, and the on-prem AI landing page for the full feature list and requirements. Setup instructions live at jitbit.com/docs/ai-on-premise.

10th Feb 2026 • 1 votes
Will AI kill SaaS helpdesks?

Do we even need helpdesk software? I mean, just give an AI agent a markdown skill file with your FAQ and canned responses and let it answer tickets - right? I (obviously) gave it a lot of thought recently and here's my take: If your entire support operation is one person answering the same three questions, congratulations, you've solved customer service. Ship it. For everyone else operating in reality - helpdesk apps do not automate support. They automate the messy human stuff around it. Workflow state & accountability Support isn't just answering questions - it's tracking who's answering them, who dropped the ball, and whose turn it is to care. Who owns this ticket? What's the SLA status? Did the second-line team even look at it, or did it rot in a queue for three days while everyone assumed someone else was on it? Teams need audit trails, escalation chains, and - let's be honest - blame-able history. An AI agent can triage, prioritize, categorize, and even draft a lovely response. It cannot enforce a process across a 70-person org where half the team is in a different timezone and the other half is "working from home" (AKA "at the beach"). Nobody's ripping out a tool for that just because ChatGPT can answer "how do I reset my password?" slightly faster. Customer data gravity You cannot put your entire docs website + a knowledge base into a markdown file. I mean, you can. You'd just need a web crawler to index your docs, dump them into a RAG database, and build an MCP server on top. Then maybe index the old tickets so the agent can search history, and... wait, you've just built a helpdesk app. With years of ticket history. Thousands of macros and canned responses. Customer sentiment patterns. Resolution time benchmarks. That one weird workaround for that one enterprise client that nobody remembers but the system does. That's institutional knowledge. That's training data. An AI agent starting from a markdown file has none of it. It's the new hire who didn't read the wiki - except the wiki doesn't exist yet either, because the wiki is the helpdesk history. Compliance & trust Enterprise buyers need GDPR compliance. Data residency. HIPAA. Audit logs that prove exactly who accessed what and when. "I built an AI agent over the weekend and pointed it at our support email" works great for a single-founder startup, but doesn't survive a procurement review. It barely survives a security questionnaire. Actually, it doesn't survive a security questionnaire - it is the security questionnaire's nightmare scenario. So what's the actual play I'm not here to dismiss AI - helpdesk apps need to absorb it. Embed AI deeply into the helpdesk itself: auto-drafted responses, smart routing, ticket summarization, triage, sentiment detection. Give customers the AI benefit inside the tool they already use, so they never feel the need to replace it. Better yet - become an MCP tool in your AI-powered org or even the orchestration layer. Let customers plug in their own AI agents, but manage them through the helpdesk. Routing rules, fallback-to-human thresholds, confidence scoring, handoff protocols. The helpdesk becomes the control plane, not the answer engine. Which, by the way, is exactly what we're building at Jitbit. The future isn't pure-AI support (customers will revolt) and it isn't pure-human support (too expensive). It's the helpdesk that best orchestrates humans and AI together. The markdown file crowd will figure that out eventually - right around the time their first enterprise prospect asks for an audit trail.

12th Jan 2026 • 1 votes

More in programming

Clip of me singing Despard in Ruddigore in 2013

A clip of me singing a funny song from Gilbert and Sullivan’s Ruddigore back in 2013

6 hours ago • 1 votes
How and Why fork() Uses Copy-on-Write

In this video, we look at why fork() needs copy-on-write, how it works inside the kernel, and a memory usage problem that Instagram encountered with Python.

12 hours ago • 1 votes
What we lost when we lost comments

Comments require commitment, but they’re worth it.

18 hours ago • 1 votes
Lighthouse map

Lovely global map with animated lights sweeping the waters

20 hours ago • 1 votes
Warming up the Puma master before it forks

Basecamp 5 runs on Puma in cluster mode: one master process with preload_app! and 63 single-threaded workers per host, deployed as a Docker container with Kamal. We serve Basecamp from several sites. Each site has its own web hosts and a read replica of the database, and writes go to a single primary database in one of them. On our busiest hosts, each deploy left up to 2,000 requests waiting while the new workers warmed up. We reduced those queues by running signed-in requests through the app in the Puma master, before it forked the workers. Why 63 single-threaded workers? Basecamp has always served web requests from processes rather than threads. It ran on Unicorn, which only does processes, until we moved to Puma in January 2025, and we kept the same setup: workers (Concurrent.physical_processor_count * 1.3).ceil threads 1, 1 preload_app! On a 48-core host that’s 63 workers, each handling one request at a time. We chose 1.3 after benchmarking HEY in 2023, when we moved our apps out of the cloud and onto our own hardware. We tested several combinations of workers and threads with a mix of GET and POST requests on a 32-vCPU VM. Every multithreaded configuration we tested was slower and handled fewer requests than single-threaded workers. Adding workers beyond about 1.2 to 1.3 per vCPU brought little benefit. The threaded workers spent a lot of their time waiting for Ruby’s global VM lock. That made single-threaded workers a good fit for this workload, and we use the same setup for Basecamp. An app that spends more time waiting on its database or other services may benefit from more threads, so benchmark your own app. The other reason is the app itself. Basecamp has class-level state in places and has never needed to be thread-safe. With one request per process, it still doesn’t. Processes do use more memory than threads, and preload_app! reduces the difference. The master loads the app once and the workers share its memory through copy-on-write until they write to it. Shopify’s comparison of Ruby execution models explains the trade-off well. In the HEY benchmark the best setup came to about 260 MB of PSS per core, where PSS counts each shared page once, split between the processes using it, and the gap to a threaded setup was smaller than we’d expected. What Puma does on each host when a container starts: one master, then 63 forked workers that share its memory until they write to it. Two things about this setup matter for the rest of the post. A worker that’s compiling or loading something is fully blocked — there’s no other thread to pick up the next request. And whatever the master has in memory before it forks, all 63 workers share. Whatever they build after the fork, they build 63 times. What happens when we deploy Kamal starts the new container alongside the old one, and kamal-proxy moves the host’s traffic across as soon as the health check passes. At that moment, the new workers have handled health checks but no customer requests. preload_app! means the master loads the app once and the workers inherit it through fork. That covers the code. It doesn’t cover anything Ruby and Rails set up on first use: YJIT compiled code. YJIT compiles a method once it’s been called a certain number of times. The master calls very little during boot, so every worker compiles the same methods again on its own first requests. Compiled templates. Action View turns each ERB template into a Ruby method the first time it’s rendered. The schema cache. Active Record reads each model’s columns from the database the first time that model is used. Inline caches and memoized values throughout Ruby, Rails and the app. All 63 workers did all of this at once, while serving the traffic the old container had been handling a second earlier. In the test environment with YJIT on, the first request to a project page on a cold process took 652 ms, 151 ms of it YJIT compiling. The same request to a warm process took 28 ms. In production, CPU time per request peaked at around 200 ms while kamal-proxy moved traffic to the new container, against about 30 ms once the workers had warmed up. A host with spare CPU absorbs this. Every one of our web hosts has 48 cores and 63 workers, but each Amsterdam host serves around 250 requests per second, against 25 to 60 at our other sites. In Amsterdam the slow first requests turned into a queue. At a peak-hour deploy, the Puma backlog on an Amsterdam host reached anywhere from 250 to 2,238 requests, and kamal-proxy’s p99 response time hit about 10 seconds. Eron, our Director of Operations, had been tracking this since June. Another server in Amsterdam would help, but it would take weeks to arrive, so we also wanted to make deploys cheaper on the hardware we already had. What didn’t work We tried a few things first. In June, Donal tested the first two on a single Amsterdam host, comparing it with its neighbors, and they ruled out two likely causes. Warming each worker’s database connections. Puma’s before_fork hook clears the master’s connections, and each worker opened its own on its first request. Opening them in before_worker_boot instead made no difference. Queries on a freshly booted production host were already under a millisecond, so connections weren’t the problem. A synthetic request in each worker. Next, each worker made a few requests in before_worker_boot to an internal controller that touched every model. That ran the middleware, routing and Active Record paths, but it ran them in 63 workers at once — exactly the CPU spike we were trying to avoid. And a request with no real data renders no real views, so most of the app stayed cold. Spreading YJIT compilation out. Delaying YJIT in each worker by a random interval spread the compiling out over a few minutes, but every worker still ran interpreted until its delay ended. The queue didn’t change. Reforking from a warm worker. This is what Shopify’s Pitchfork does: let one worker serve traffic until it’s warm, then fork the others from it. Puma has an experimental version called fork_worker, and on beta it worked — the reforked workers were warm after three to five requests, where fresh ones took up to 30 seconds. But with fork_worker the template is worker 0, and it keeps serving requests. If it exits, the workers waiting to be forked never start (puma/puma#3596). If it gets no traffic, the refork never happens, which is what we saw on beta. Instacart have a mold_worker patch that promotes a warm worker to a template that stops serving, but it isn’t in a Puma release. We have a branch of it, and we may come back to it. That last experiment did show us where the fix was, though. Everything a warm worker has that a cold one lacks is in its memory, and fork copies memory. The master already has the app loaded. It just never runs it. Run the requests in the master So now, before the master binds its socket and forks, it makes the app’s own requests, in-process, the way a signed-in user would. Rack has a hook for exactly this. Rack::Builder#warmup takes a block that’s called once with the built app, before the server starts. rails server builds the app from config.ru, so the change to boot is one line: require_relative "config/environment" warmup { WarmUp.configured.run } if ENV["WARM_UP"] run Rails.application With preload_app! this runs in the master, and the workers inherit whatever it did. Puma binds its socket after the app is built, so until the warm-up finishes the health check’s connection is refused and kamal-proxy keeps retrying. No request reaches a worker that hasn’t been warmed. The warm-up has three steps. After precompiling the views, it gives the page requests and schema loading a shared 20-second budget, checked before each page or model. 1. Precompile the views actionview_precompiler reads every template for its render calls and compiles each one with the locals it’s passed. For us that’s 1,394 templates in about two seconds. A first request to a project page then compiles 2 templates instead of 44. 2. Request the pages, signed in A small browser class makes the requests through Rack::MockRequest, with the two cookies a real sign-in sets, then goes back for each page’s lazy Turbo frames: class WarmUp::Browser def initialize(signed_in_as:) @client = Rack::MockRequest.new(Rails.application) @headers = { "HTTP_USER_AGENT" => "Basecamp warm-up", "HTTP_COOKIE" => cookie_for(signed_in_as), "bc3.warm_up" => true } end def visit(path) page = get(path) frames_in(page).each { |id, src| get(src, "HTTP_TURBO_FRAME" => id) } end private def get(path, headers = {}) @client.get("https://#{host}#{path}", @headers.merge(headers)) end def frames_in(page) Nokogiri::HTML5(page.body).css("turbo-frame[src]").map { |frame| [ frame["id"], frame["src"] ] } end end The requests are signed in. The user is a monitoring account we already use for automated checks, and the pages are its own project, Campfire, to-dos, documents and messages. Public pages weren’t enough: after warming up with signed-out pages only, the first signed-in request to the projects page still took 131 ms, because authentication, the signed-in controllers and their views had never run. With signed-in pages it took 40 ms. cookie_for writes the same signed cookie the sign-in controller does, using the app’s own cookie jar, so there’s no API token and no secret to store. The frames are followed. The busiest HTML requests in production aren’t pages at all but Turbo frames — the sidebar badge, the inbox, the navigation menus. The browser parses each page and requests its <turbo-frame src> URLs with the Turbo-Frame header, so those controllers and views get warmed too. Our first four pages turned into 60 requests. The requests are excluded from rate limiting. They are internal, so they do not count against the rate limits that apply to real visitors. 3. Load the rest of the schema The page requests load the schema for the models they touch. The last step loads the rest, from the read replica: ApplicationRecord.reading do models.lazy.take_while { time_left? }.each { |model| model.load_schema if model.table_exists? } end The step checks 261 models and loads any schema information still missing. Those database round trips add up when the primary is far away: outside a request, Active Record uses the writing role, and from a host a long way from the primary each round trip is tens of milliseconds. Reading from the local replica brings the step down from about 20 seconds to 3.5. The pages go first because they load most of the schema anyway. If the time budget runs out, the step stops, logs how many models it got through, and the workers load the rest on first use like they always did. Rails can also load the schema from a dumped cache file at boot (bin/rails db:schema:cache:dump), which would make this step unnecessary. We don’t ship one in our image yet, because the dump needs a database to read from at build time, and we have several databases to cover. It’s on the list. What to close before the fork Running requests in the master opens things the master never opened before, and every worker inherits them. Two processes writing to the same socket will corrupt each other’s traffic, so you need to know what’s open before you fork. The way to find out is to list the master’s open file descriptors — ls -l /proc/<pid>/fd — before and after a warm-up, in an environment set up like production. Development wasn’t enough for us: it stores files on disk, so our S3 connections only showed up in production. Then, for each thing that’s open, check how its library handles a fork. We found three kinds: Already handled. Plenty of libraries detect a fork on their own, either by recording the PID they connected from and reconnecting in the child, by opening per-process files, or by resetting their thread pools. Redis clients, metrics libraries and concurrency libraries tend to be in this group. Check, but you probably don’t need to do anything. Already closed. Database connections are the classic one, and most Puma configs already clear them in before_fork. Anything else that’s opened per process — we have a SQLite cache the workers open on boot — needs closing when the warm-up finishes. Needs a new step. HTTP clients with keep-alive connections are the ones to look for: cloud SDKs with connection pools, tracing exporters, error reporters. They usually have no fork handling at all. We empty the aws-sdk connection pools in before_fork, and we run the warm-up untraced so the OpenTelemetry exporter never opens its connection to Tempo in the first place. Once that’s done, before_fork finishes with Process.warmup, which Ruby 3.3 added for this purpose: a major GC, a heap compaction, and every surviving object promoted to the old generation, so the memory pages the workers share change as little as possible afterwards. Choosing the pages The first list was the four pages that ran the busiest requests on beta. Once the warm-up was live, production showed us which endpoints were still cold. For one deploy, we compared each endpoint’s mean duration in the six minutes after kamal-proxy moved traffic to the new container with the same endpoint an hour later, then multiplied the difference by the number of requests in those six minutes. That gives the extra time each endpoint cost us because it was cold: Endpoint Cold Warm Requests in 6 min Extra seconds Campfire 246 ms 70 ms 6,490 1,140 Projects (JSON API) 84 ms 50 ms 22,077 771 Docs & Files 262 ms 177 ms 4,996 421 To-dos tool 205 ms 113 ms 4,018 371 To-dos (JSON API) 33 ms 16 ms 18,738 320 The pages already in the warm-up showed what to expect: the project page kept a 36 ms gap after a deploy, and the to-do page 10 ms. We’ve proposed adding these five requests, and expect them to add about five to seven seconds to the page step. The two JSON endpoints were a surprise. The warm-up’s page list had no API requests in it, so nothing on the API path had run before the first real request: not the API controllers, and not the Jbuilder templates rendering real records. Precompiling the views covers JSON templates too, but it isn’t a substitute for running the request. Results The warm-up is on for all 68 web hosts. With the first four pages it took 12 to 16 seconds per host: about 2 seconds to precompile the views, 7 to 9 for the 60 requests, and 3.5 for the schema. Deploys take that much longer per host, and we raised the deploy timeout from 30 to 60 seconds to cover it. In Amsterdam, at a peak-hour deploy: During deploy Before After Peak Puma backlog per host 250–2,238 requests 19–223 requests Peak kamal-proxy p99 about 10 s 2.4–4.8 s Peak CPU time per request 201–214 ms 88–132 ms Peak database time per request 56–69 ms 39–47 ms The same eight hosts at three deploys on 1 October, an hour apart, as the warm-up went from one host to four to all eight. The deploy in the middle, with four hosts warmed and four not, shows why every host needed the warm-up. Each warmed host recovered faster on its own: mean request duration peaked at 130 to 173 ms, against 203 to 311 ms on the hosts that weren’t warmed. But the backlog on all eight was about the same, because they were all waiting on the same database. Mean request duration on each host at the 07:21 UTC deploy. Blue hosts warmed up in the master before forking, orange hosts did not. Memory came down too. The workers now share compiled templates, YJIT code and the schema with the master instead of each building their own copy. On beta, the view precompiler alone took a busy worker’s private memory from 174–202 MB to 119–135 MB. Thirty minutes after the deploy, the web containers used about 39 GB less memory than the previous day’s containers at the same age and traffic. Amsterdam served most of our traffic at the times we tested. In Amsterdam, each new container used about 2 GB less just after traffic moved to it, which lowers the peak while the old and new containers overlap. Working with Claude Claude Code helped throughout. It combed through the per-worker backlogs and per-endpoint timings in Prometheus and Loki after each deploy, worked out the cold-versus-warm cost of each endpoint, and prepared the changes and the pull request descriptions with the benchmarks in them. We decided what to try, deployed it and read the results. If you do this Warm the master before it forks. Compile common code and templates and load their schema in the master, so workers inherit that work. With preload_app!, Rack::Builder#warmup runs before the workers start accepting traffic. Use the app’s real requests. Public pages, internal endpoints and synthetic queries warm the paths they run and nothing else. Signed-in requests to real records, frames included, run what production runs. Measure the cold penalty per endpoint. The difference between an endpoint’s cold and warm duration, times its request count after a deploy, ranks the pages worth adding. Ours weren’t the ones we’d have guessed, and two of them were JSON. Check what the warm-up leaves open. List the master’s file descriptors after a warm-up and account for every one before the fork. Two of ours needed changes. Set a time budget. A warm-up that runs long on one slow host fails the deploy on that host. Ours gives the page requests and schema loading a shared 20-second budget, checked before each page or model, puts the most valuable pages first, and logs what it skipped. Reforking from a warm worker, as Pitchfork does, solves the same problem continuously rather than once at boot, and it would warm paths no fixed list of pages covers. We may still get there: our branch brings Instacart’s mold_worker up to date with Puma’s main branch and fixes the bugs we found in it. But warming the master works with the Puma we already run, took a few days to implement, and substantially reduced the queues after deployment.

yesterday • 1 votes
📚 BoredReading

You seem to be enjoying this.

Join free to unlock everything.

Create free account

Already have an account? Sign in