More from Evan Jones - Software Engineer | Computer Scientist
You can't safely use the C setenv() or unsetenv() functions in a program that uses threads. Those functions modify global state, and can cause other threads calling getenv() to crash. This also causes crashes in other languages that use those C standard library functions, such as Go's os.Setenv (Go issue) and Rust's std::env::set_var() (Rust issue). I ran into this in a Go program, because Go's built-in DNS resolver can call C's getaddrinfo(), which uses environment variables. This cost me 2 days to track down and file the Go bug. Sadly, this problem has been known for decades. For example, an article from January 2017 said: "None of this is new, but we do re-discover it roughly every five years. See you in 2022." This was only one year off! (She wrote an update in October 2023 after I emailed her about my Go bug.) This is a flaw in the POSIX standard, which extends the C Standard to allow modifying environment varibles. The most infuriating part is that many people who could influence the standard or maintain the C libraries don't see this as a problem. The argument is that the specification clearly documents that setenv() cannot be used with threads. Therefore, if someone does this, the crashes are their fault. We should apparently read every function's specification carefully, not use software written by others, and not use threads. These are unrealistic assumptions in modern software. I think we should instead strive to create APIs that are hard to screw up, and evolve as the ecosystem changes. The C language and standard library continue to play an important role at the base of most software. We either need to figure out how to improve it, or we need to figure out how to abandon it. Why is setenv() not thread-safe? The biggest problem is that getenv() returns a char*, with no need for applications to free it later. One thread could be using this pointer when another thread changes the same environment variable using setenv() or unsetenv(). The getenv() function is perfect if environment variables never change. For example, for accessing a process's initial table of environment variables (see the System V ABI: AMD64 Section 3.4.1). It turns out the C Standard only includes getenv(), so according to C, that is exactly how this should work. However, most implementations also follow the POSIX standard (e.g. POSIX.1-2017), which extends C to include functions that modify the environment. This means the current getenv() API is problematic. Even worse, putenv() adds a char* to the set of environment variables. It is explicitly required that if the application modifies the memory after putenv() returns, it modifies the environment variables. This means applications can modify the value passed to putenv() at any time, without any synchronization. FreeBSD used to implement putenv() by copying the value, but it changed it with FreeBSD 7 in 2008, which suggests some programs really do depend on modifying the environment in this fashion (see FreeBSD putenv man page). As a final problem, environ is a NULL-terminated array of pointers (char**) that an application can read and assign to (see definition in POSIX.1-2017). This is how applications can iterate over all environment variables. Accesses to this array are not thread-safe. However, in my experience many fewer applications use this than getenv() and setenv(). However, this does cause some libraries to not maintain the set of environment variables in a thread-safe way, since they directly update this table. Environment variable implementations Implementations need to choose what do do when an application overwrites an existing variable. I looked at glibc, musl, Solaris/Illumos, and FreeBSD/Apple's C standard libraries, and they make the following choices: Never free environment variables (glibc, Solaris/Illumos): Calling setenv() repeatedly is effectively a memory leak. However, once a value is returned from getenv(), it is immutable and can be used by threads safely. Free the environment variables (musl, FreeBSD/Apple): Using the pointer returned by getenv() after another thread calls setenv() can crash. A second problem is ensuring the set of environment variables is updated in a thread-safe fashion. This is what causes crashes in glibc. glibc uses an array to hold pointers to the "NAME=value" strings. It holds a lock in setenv() when changing this array, but not in getenv(). If a thread calling setenv() needs to resize the array of pointers, it copies the values to a new array and frees the previous one. This can cause other threads executing getenv() to crash, since they are now iterating deallocated memory. This is particularly annoying since glibc already leaks environment variables, and holds a lock in setenv(). All it needs to do is hold the lock inside getenv(), and it would no longer crash. This would make getenv() slightly slower. However, getenv() already uses a linear search of the array, so performance does not appear to be a concern. More sophisticated implementations are possible if this is a problem, such as Solaris/Illumos's lock-free implementation. Why do programs use environment variables? Environment variables useful for configuring shared libraries or language runtimes that are included in other programs. This allows users to change the configuration, without program authors needing to explicitly pass the configuration in. One alternative is command line flags, which requires programs to parse them and pass them in to the libraries. Another alternative are configuration files, which then need some other way to disable or configure, to be able to test new configurations. Environment variables are a simple solution. AS a result, many libraries call getenv() (see a partial list below). Since many libraries are configured through environment variables, a program may need to change these variables to configure the libraries it uses. This is common at application startup. This causes programs to need to call setenv(). Given this issue, it seems like libraries should also provide a way to explicitly configure any settings, and avoid using environment variables. We should fix this problem, and we can In my opinion, it is rediculous that this has been a known problem for so long. It has wasted thousands of hours of people's time, either debugging the problems, or debating what to do about it. We know how to fix the problem. First, we can make a thread-safe implementation, like Illumos/Solaris. This has some limitations: it leaks memory in setenv(), and is still unsafe if a program uses putenv() or the environ variable. However, this is an improvement over the current Linux and Apple implementations. The second solution is to add new APIs to get one and get all environment variables that are thread-safe by design, like Microsoft's getenv_s() (see below for the controversy around C11's "Annex K"). My preferred solution would be to do both. This would reduce the chances of hitting this problem for existing programs and libraries, and also provide a path to avoid the problems entirely for new code or languages like Go and Rust. My rough idea would be the following: Add a function to copy one single environment variable to a user-specified buffer, similar to getenv_s(). Add a thread-safe API to iterate over all environment variables, or to copy all variables out. Mark getenv() as deprecated, recommending the new thread-safe getenv() function instead. Mark putenv() as deprecated, recommending setenv() instead. Mark environ as deprecated, recommending environment variable functions instead. Update the implementation of environment varibles to be thread-safe. This requires leaking memory if getenv() is used on a variable, but we can detect if the old functions are used, and only leak memory in that case. This means programs written in other languages will avoid these problems as soon as their runtimes are updated. Update the C and POSIX standards to require the above changes. This would be progress. The getenv_s / C Standard Annex K controversy Microsoft provides getenv_s(), which copies the environment variable into a caller-provided buffer. This is easy to make thread-safe by holding a read lock while copying the variable. After the function returns, future changes to the environment have no effect. This is included in the C11 Standard as Annex K "Bounds Checking Interfaces". The C standard Annexes are optional features. This Annex includes new functions intended to make it harder to make mistakes with buffers that are the wrong size. The first draft of this extension was published in 2003. This is when Microsoft was focusing on "Trustworthy Computing" after a January 2002 memo from Bill Gates. Basically, Windows wasn't designed to be connected to the Internet, and now that it was, people were finding many security problems. Lots of them were caused by buffer handling mistakes. Microsoft developed new versions of a number of problematic functions, and added checks to the Visual C++ compiler to warn about using the old ones. They then attempted to standardize these functions. My understanding is the people responsible for the Unix POSIX standards did not like the design of these functions, so they refused to implement them. For more details, see Field Experience With Annex K published in September 2015, Stack Overflow: Why didn't glibc implement _s functions? updated March 2023, and Rich Felker of musl on both technical and social reasons for not implementing Annex K from February 2019. I haven't looked at the rest of the functions, but having spent way too long looking at getenv(), the general idea of getenv_s() seems like a good idea to me. Standardizing this would help avoid this problem. Incomplete list of common environment variables This is a list of some uses of environment variables from fairly widely used libraries and services. This shows that environment variables are pretty widely used. Cloud Provider Credentials and Services AWS's SDKs for credentials (e.g. AWS_ACCESS_KEY_ID) Google Cloud Application Default Credentials (e.g. GOOGLE_APPLICATION_CREDENTIALS) Microsoft Azure Default Azure Credential (e.g. AZURE_CLIENT_ID) AWS's Lambda serverless product: sets a large number of variables like AWS_REGION, AWS_LAMBDA_FUNCTION_NAME, and credentials like AWS_SECRET_ACCESS_KEY Google Cloud Run serverless product: configuration like PORT, K_SERVICE, K_REVISION Kubernetes service discovery: Defines variables SERVICE_NAME_HOST and SERVICE_NAME_PORT. Third-party C/C++ Libraries OpenTelemetry: Metrics and tracing. Many environment variables like OTEL_SERVICE_NAME and OTEL_RESOURCE_ATTRIBUTES. OpenSSL: many configurable variables like HTTPS_PROXY, OPENSSL_CONF, OPENSSL_ENGINES. BoringSSL: Google's fork of OpenSSL used in Chrome and others. It reads SSLKEYLOGFILE just like OpenSSL for logging TLS keys for debugging. Libcurl: proxies, SSL/TLS configuration and debugging like HTTPS_PROXY, CURL_SSL_BACKEND, CURL_DEBUG. Libpq Postgres client library: connection parameters including credentials like PGHOSTADDR, PGDATABASE, and PGPASSWORD. Rust Standard Library std::thread RUST_MIN_STACK: Calls std::env::var() on the first call to spawn() a new thread. It is cached in a static atomic variable and never read again. See implementation in thread::min_stack(). std::backtrace RUST_LIB_BACKTRACE: Calls std::env::var() on the first call to capture a backtrace. It is cached in a static atomic variable and never read again. See implementation in Backtrace::enabled().
The read() and write() system calls take a variable-length byte array as an argument. As a simplified model, the time for the system call should be some constant "per-call" time, plus time directly proportional to the number of bytes in the array. That is, the time for each call should be time = (per_call_minimum_time) + (array_len) × (per_byte_time). With this model, using a larger buffer should increase throughput, asymptotically approaching 1/per_byte_time. I was curious: do real system calls behave this way? What are the ideal buffer sizes for read() and write() if we want to maximize throughput? I decided to do some experiments with blocking I/O. These are not rigorous, and I suspect the results will vary significantly if the hardware and software are different than one the system I tested. The really short answer is that a buffer of 32 KiB is a good starting point on today's systems, and I would want to measure the performance to go beyond that. However, for large writes, performance can increase. On Linux, the simple model holds for small buffers (≤ 4 KiB), but once the program approaches the maximum throughput, the throughput becomes highly variable and in many cases decreases as the buffers get larger. For blocking I/O, approximately 32 KiB is large enough to hit the maximum throughput for read(), but write() throughput improves with buffers up to around 256 KiB - 1 MiB. The reason for the asymmetry is that the Linux kernel will only write less than the entire buffer (a "short write") if there is an error (e.g. a signal causing EINTR). Thus, larger write buffers means the operating system needs to switch to the process less often. On the other head, "short reads", where a read() returns less than the maximum length, become increasingly common as the buffer size increases, which diminishes the benefit. There is a SO_RCVLOWAT socket option to change this that I did not test. The experiments were run on two 16 CPU Google Cloud T2D instances, which use AMD EPYC Milan processors (3rd generation, released in 2021). Each core is a real physical core. I used Ubuntu 23.04 running kernel 6.2.0-1005-gcp. My benchmark program is written in Rust and is available on Github. On localhost, Unix sockets were able to transfer data at approximately 9000 MiB/s. Localhost TCP sockets were a bit slower, around 7000 MiB/s. When using two separate cloud VMs with a networking throughput limit of 32 Gbps = 3800 MiB/s, I needed to use 6 TCP sockets to reliably reach that maximum throughput. A single TCP socket gets around 1400 MiB/s with 256 KiB buffers, with peaks as high as 2200 MiB/s. Experiment 1: /dev/zero and /dev/urandom My first experiment is reading from the /dev/zero and /dev/urandom devices. These are software devices implemented by the kernel, so they should have low overhead and low variability, since other tasks are not involved. Reading from /dev/urandom should be much slower than /dev/zero since the kernel must generate random bytes, rather than just zeros. The chart below shows the throughput for reading from /dev/zero as the buffer size is increased. The results show that the basic linear time per system call model holds until the system reaches maximum throughput (256 kiB buffer = 39000 MiB/s for /dev/zero, or 16 kiB = 410 MiB/s for /dev/urandom). As the buffer size increases further, the throughput decreases as the buffers get too big. This suggests that some other cost for larger buffers starts to outweigh the reduction in number of system calls. Perhaps CPU caches become less effective? The AMD EPYC Milan (3rd gen) CPU I tested on has 32 KiB of L1 data cache and 512 KiB of L2 data cache per core. The performance decreases don't exactly line up with these numbers, but it still seems plausible. The numbers for /dev/urandom are substantially lower, but otherwise similar. I did a linear least-squares fit on the average time per system call, shown in the following chart. If I use all the data, the fit is not good, because the trend changes for larger buffers. However, if I use the data up to the maximum throughput at 256 KiB, the fit is very good, as shown on the chart below. The linear fit models the minimum time per system call as 167 ns, with 0.0235 ns/byte additional time. If we want to use smaller buffers, using a 64 KiB buffer for reading from /dev/zero gets within 95% of the maximum throughput. Experiment 2: Unix and localhost TCP sockets Exchanging data with other processes is the thing I am actually interested in, so I tested Unix and TCP sockets on a single machine. In this case, I varied both the write buffer size and the read buffer size. Unfortunately, these results vary a lot. A more robust comparison would require running each experiment many times, and using some sort of statistical comparison. However, this "quick and dirty" experiment satisfied my curiousity, so I didn't do that. As a result, my conclusions here are vague. The simple model that increasing buffer size should decrease overhead is true, but only until the buffers are about 4 KiB. Above that point, the results start to be highly variable, and it is much harder to draw general conclusion. However, appears that increasing the write buffer size generally is quite helpful up to at least 256 KiB, and often needed as much as 1 MiB to get the highest localhost throughput. I suspect this is because on Linux with blocking sockets, write() will not return until it has written all the data in the buffer, unless there is an error (e.g. EINTR). As a result, passing a large buffer means the kernel can do a lot of the work without needing to switch back to user space. Unfortunately, the same is not true for read(), which often returns "short reads" with any data that is available in the buffer. This starts with buffer sizes around 2 KiB, with the percentage of short reads increasing as the buffer size gets larger. This means the simple model does not hold, because we aren't actually increasing the bytes per read call. I suspect this is a factor which means this microbenchmark is likely not representative of real programs. A real program will do something with the buffer, which will provide time for more data to be buffered in the kernel, and would probably decrease the number of short reads. This likely means larger buffers are in practice more useful than this microbenchmark suggests. As a result of this, the highest throughput often was achievable with small read buffers. I'm somewhat arbitrarily selecting 16 KiB at the best read buffer, and 256 KiB as the best write buffer, although a 1 MiB write buffer seems to be To give a sense of how variable the results are, the plot below shows the local Unix socket throughput for each read and write buffer throughput size. I apologize for the ugly plot. I did not want to spend the time to make it more beautiful. This plot is interactive so you can slice the data to the area of interest. I recommend zooming in to the left hand size with read buffers up to about 300 KiB. The first thing to note is at least on Linux with blocking sockets, the writer will almost never have a "short write", where the write system call returns before writing all the data in the buffer. Unless there is a signal (EINTR) or some other "error" condition, write() will not return until all the bytes are written. The same is not true for reads. The read() system call will often return a "short" read, starting around buffer sizes of 2 KiB. The percentage of short reads generally increases as buffer sizes get bigger, which is logical. Another note is that sockets have in-kernel send and receive buffers. I did not tune these at all. It is possible that better performance is possible by tuning these settings, but that was not my goal. I wanted to know what happens "out of the box" for general-purpose programs without any special tuning. Experiment 3: TCP between two hosts In this experiment, I used two separate hosts connected with 32 Gbps networking in Google Cloud. I first tested the TCP throughput using iperf, to independently verify the network performance. A single TCP connection with iperf is not enough to fully utilize the network. I tried fiddling with some command line options and with Kernel settings like net.ipv4.tcp_rmem and wasn't able to get much better than about 12 Gb/s = 1400 MiB/s. The throughput is also highly varible. Here is some example output with iperf reporting at 2 second intervals, where you can see the throughput ranging from 10 to 19 Gb/s, with an average over the entire interval of 12 Gb/s. To hit the maximum network throughput, I need to use 6 or more parallel TCP connections (iperf -c IP_ADDRESS --time 60 --interval 2 -l 262144 -P 6). Using 3 connections gets around 26 Gb/s, and using 4 or 5 will occasionally hit the maximum, but will also occasionally drop down. Using at least 6 seems to reliably stay at the maximum. Due to this variability, it is hard to draw any conclusions about buffer size. In particular: a single TCP connection is not limited by CPU. The system uses about 40% of a single CPU core, basically all in the kernel. This is more about how the buffer sizes may impact scheduling choices. That said, it is clear that you cannot hit the maximum throughput with a small write buffer. The experiments with 4 KiB write buffers reached approximately 300 MiB/s, while an 8 KiB write buffer was much faster, around 1400 MiB/s. Larger still generally seems better, up to around 256 KiB, which occasionally reached 2200 MiB/s = 17.6 Gb/s. The plot below shows the TCP socket throughput for each read and write buffer size. Again, I apologize for the ugly plot.
This is a post for myself, because I wasted a lot of time understanding this bug, and I want to be able to remember it in the future. I expect close to zero others to be interested. The C standard library function isspace() returns a non-zero value (true) for the six "standard" ASCII white-space characters ('\t', '\n', '\v', '\f', '\r', ' '), and any locale-specific characters. By default, a program starts in the "C" locale, which will only return true for the six ASCII white-space characters. However, if the program changes locales, it can return true for other values. As a result, unless you really understand locales, you should use your own version of this function, or ICU4C's u_isspace() function. An implementation of isspace() for ASCII is one line: /* Returns true for the 6 ASCII white-space characters: \t \n \v \f \r ' '. */ int isspace_ascii(int c) { return c == '\t' || c == '\n' || c == '\v' || c == '\f' || c == '\r' || c == ' '; } I ran into this because On Mac OS X, Postgres switches to the system's default locale, which is something that uses UTF-8 (e.g. en_US.UTF-8, fr_CA.UTF-8, etc). In this case, isspace() returns true for Unicode white-space values, which includes 0x85 = NEL = Next Line, and 0xA0 = NBSP = No-Break Space. This caused a bug in parsing Postgres Hstore values that use Unicode. I have attempted to submit a patch to fix this (mailing list post, commitfest entry). For a program to demonstrate the behaviour on different systems, see isspace_locale on Github.
Nearly all programs are written to access virtual memory addresses, which the CPU must translate to physical addresses. These translations are usually fast because the mappings are cached in the CPU's Translation Lookaside Buffer (TLB). Unfortunately, virtual memory on x86 has used a 4 kiB page size since the 386 was released in 1985, when computers had a bit less memory than they do today. Also unfortunately, TLBs are pretty small because they need to be fast. For example, AMD's Zen 4 Microarchitecture, which first shipped in September 2022, has a first level data TLB with 72 entries, and a second level TLB with 3072 entries. This means when an application's working set is larger than approximately 4 kiB × 3072 = 12 MiB, some memory accesses will require page table lookups, multiplying the number of memory accesses required. This is a brand-new CPU, with one of the biggest TLBs on the market, so most systems will be worse. Using larger virtual memory page sizes (aka huge pages) can reduce page mapping overhead substantially. Since RAM is so much larger than it was in 1985, a larger page size seems like obviously a good idea to me. In 2021, Google published a paper about making their malloc implementation (TCMalloc) huge page aware (called Temeraire). They report this improved average requests-per-second throughput across their fleet by 7%, by increasing the amount of memory that is backed by huge pages. This made me curious about the "best case" performance benefits. I wrote a small program that allocates 4 GiB, then randomly reads uint64 values from it. On my Intel 11th generation Core i5-1135G7 (Tiger Lake) from 2020, using 2 MiB huge pages is 2.9× faster. I also tried 1 GiB pages, which is 3.1× faster than 4 kiB pages, but only 8% faster than 2 MiB pages. My conclusion: Using madvise() to get the kernel to use huge pages seems like a relatively easy performance win for applications that use a large amount of RAM. Unfortunately, using larger pages is not without its disadvantages. Notably, when the Linux kernel's transparent huge page implementation was first introduced, it was enabled by default, which caused many performance problems. See the section below for more details. Today's default to use huge pages only for applications that opt-in (aka madvise) should improve this. The kernel's policies for managing huge pages have also changed since then, and are hopefully better now. At the very least, the fact that Google uses transparent huge pages for all their applications is some evidence that this can work for a wide variety of workloads. The second problem with larger page sizes is software incompatibility, since so much software is only tested on x86 with 4 kiB pages. Linux on ARM64 used to default to 64 kiB pages. However, this caused many problems (e.g. dotnet, Go, Chrome, jemalloc, Asahi Linux list of broken software). It appears that around 2020 most distributions switched to 4 kiB pages to avoid these problems (e.g. RedHat RHEL9 change in 2021, Ubuntu note about the page size change). Page size historical details Other CPU architectures have made different page size choices. Notably, iOS and Mac OS X on ARM64 uses 16 kiB pages (ARM64 aka aarch64 supports 4, 16, and 64 kiB pages, although specific CPUs will only support some of them). Alpha and Sparc used 8 kiB pages. PowerPC on Linux uses 64 kiB pages, although Redhat/Fedora are considering switching to 4 kiB due to the same compatibility issues. See page sizes used by Windows on various processors. Latency and throughput problems with transparent huge pages The Linux kernel's implementation of transparent huge pages has been the source of performance problems. When introduced, it was initially enabled for all processes and memory regions by default. This caused a large number of problems, which eventually caused the kernel's default to change to madvise, where programs have to opt-in to use huge pages (see Nelson Elhage's summary (2017), and Ubuntu bug that changed the default (2017/released 2019). The performance problems are rare high latency (e.g. operations being substantially slower than normal), throughput issues due to excess CPU consumption of the kernel background tasks, or substantial increases in memory usage. Some examples are Hadoop (2012), TokuDB/MySQL (2014), Redis/jemalloc (2015), TiKV/TiDB (2020). The problems seem to fall into the following categories: Increasing memory usage by making fragmentation worse: using transparent huge pages rounds allocations up to 2 MiB. If an application allocates many separate memory regions, this can cause lots of memory to be wasted. Most of the problems have been where an application uses a large amount of memory, then frees a lot of it, leaving "holes" in the large pages. Sometimes the kernel's transparent page policy can decide to turn these back into huge pages, which causes the memory usage to increase. For example, see a Go bug (2015) and the corresponding kernel bug report (2015). The fix for Go was to only return memory on huge page granularity. This also happened to Redis with jemalloc (2015) malloc implementations that are not huge page aware may add more kernel CPU overhead: When returning memory to the operating system, if the memory allocator is not aware of huge pages, it may return part of a huge page. This causes the kernel to split the huge page back into separate 4 kib pages. This adds overhead, and also fragments memory, making fewer huge pages available, causing the kernel to do more work the next time it tries to allocate a huge page. This article about TokuDB from 2014 suggests that it ran into this problem with jemalloc. The good news is that it now seems like all major malloc implementations (jemalloc, tcmalloc, mimalloc, and glibc malloc) all have some huge page support, which should make this less bad. slow memory allocations due to fragmentation (latency): When trying to allocate a huge page, the kernel may spend time moving memory around to free up a page. See a detailed thread about impacts on the JVM (2017). The kernel's current default is to only do this for regions that have opted in with madvise. This should mean that other processes won't be penalized too much by this, but it does mean the process that called madvise could be stalled briefly when allocating new pages. One way to avoid this is to immediately touch every huge page in an allocation, to cause the cost to happen up front. This would work well for allocations that are made at program startup, such as caches. fork() e.g. Redis: Calling fork marks all of the process's pages as copy-on-write. Then when a single byte on a page is modified, the page must be copied. Redis uses fork to create a read-only "snapshot" of memory, when writing a checkpoint to disk. Since huge pages are 512X larger than "normal" pages, the time to copy a page increases by 512X. It also means the memory usage is higher, since modifying a single byte causes 2 MiB to be copied, instead of only 4 kiB. Using fork() in this way with huge pages seems like a bad idea. See details about a workload that causes this behavior (2014). References Huge Page Demo Evan Jones 2022-01-18: My huge page demonstration program. Larger Pages: Richard Sites 2022-05-06: Argues we should increase the minimum page size to 64 kiB, and maintain compatibility by using access flags on 4 kiB sub-pages. Stack Overflow: Why is the page size 4 KB? Answer by Hadi Brais 2018-04-26: a great look at the history of why 4 kiB pages were chosen. Using huge pages on Linux: Erik Rigtorp 2020-10-08: A hash table benchmark in C++ with results for both transparent and explicit huge pages. Reliably allocating huge pages in Linux: Francesco Mazzoli 2021-11-22: Includes C code describing how to verify if an address is a huge page. Intel Coffee Lake Microarchitecture (2017 aka Core 9th gen): L1 Data TLB: 64 entries for 4 kiB pages / 32 entries for 2 MiB pages / 4 entries for 1 GiB pages ; L2 unified TLB: 1536 for 4 kiB/2 MiB pages; 16 entries for 1 GiB pages.
More in programming
Today we are releasing the version 1.0 of Lexxy. Lexxy is a rich text editor for Rails built on Lexical. It already powers Basecamp, Fizzy and many others, and it will become the default editor in Rails. I recently presented it in Rails World (slides, video coming soon). This is the article version of my talk. Trix hit a wall Trix has been our editor since 2015, and every Rails app’s editor since Action Text shipped in Rails 6. It’s small and reliable, and it has served millions of people for a decade. But in the last few years our customers kept asking for features like tables or code highlighting, and we kept struggling to deliver them. The reason is the Trix document model. A Trix document is a flat list of blocks. A block is a line of text with some labels attached, like quote, bullet list, bullet. There is no tree, and a block can never contain another block. Nesting is an illusion: at render time, adjacent blocks with the same labels get wrapped together. A flat list of blocks with very limited extensibility options That design bought a lot of simplicity, but you can’t express something like a table with it. Two cells next to each other would carry exactly the same labels, so Trix would merge them into one. The model can say “this bullet is one level deeper”. It cannot say “this cell is different from the cell next to it”. A tables issue has been opened since 2015! The model just can’t do it. The second problem was maintenance. An editor built on contenteditable behaves differently in every browser and even changes from time with operating system releases. In 2024, three iOS releases in a row broke typing, dictation or the caret in Trix, and each one cost us real effort to work around. Check this one as an example. Why Lexical This was a conversation we had at 37signals for years: 2022. We started a project to add tables to Trix. We gave up after a week. The document model can’t represent two-dimensional things. 2023. I built a proof of concept with Tiptap inside HEY. We liked it, but we never started a serious project with it. 2024. We built House, our own Markdown editor, for Writebook. Not WYSIWYG, but WYSIWYM: what you see is what you mean. A wonderful editor for long-form writing like books or technical documentation. 2025. We tried House in another product, and it didn’t fit. For most apps, WYSIWYG was just the right answer. 2025. We had the discussion again, and this time we looked at the whole field. Four years of the same conversation David ruled out Tiptap, CKEditor and the other commercial editors: an open source core with features kept proprietary, and a sales team behind them. We didn’t want our editor to depend on somebody else’s licensing decisions. Then we found Lexical: MIT, from Meta, and very powerful. I spent two weeks evaluating it, and in May we made the call to go with it. A tiny core. Pick the rest. Lexical’s core has zero dependencies and weighs forty-two kilobytes. In a way, it validates the approach that Trix pioneered. The document is an immutable state you never mutate directly, you just get new snapshots when performing updates; contenteditable is an input device and a rendering surface, never the source of truth. It has a DOM reconciler to update the actual DOM very efficiently, and other primitives to deal with handling commands and node transformations. Everything else, from lists to tables to markdown, is a package in the orbit. Lexxy uses thirteen of them. Lexical solved the maintenance problem too. Meta’s products like Facebook or Instagram use Lexical and their user count is in the hundreds of millions. This means that even small issues with new keyboards and devices are fixed fast by the dedicated Meta team that maintains it. Furthermore, Meta’s business is the products, not the editor. We much rather liked this structure of incentives for the long-term investment an editor represents. The iceberg The plan was simple. Pick Lexical, wire it up to Action Text, add a toolbar and ship it. Well, it didn’t go exactly like that. What you envision, and what's under the water Lexxy today is thirteen thousand lines of vanilla JavaScript on top of Lexical. A great editing experience is very hard to get right. An editor is a machine where the user can change the state in a thousand different ways. For example. you have two images one after the other and want to put the cursor between them, but there is nothing there to put a cursor in. Or somebody pastes from Google Docs, and you have to turn a pile of inline styles and empty spans into clean markup. And then Safari, and Android keyboards, and the clipboard, and undo, and… From the first pull request to Basecamp 5 The first pull request landed in May 2025. Fizzy launched with Lexxy in December, and Basecamp 5 in May this year. Basecamp was the real test: twenty years of content written with Trix, and people who use the editor all day, every day. Zoltán Hosszú and Samuel Péchèr were the key people who made this happen. The took a very green version of Lexxy, added a ton of features (including Tables) and polish, and they fixed countless bugs. They also pulled off a remarkable milestone: seamlessly switching millions of Basecamp users from Trix to Lexxy. What’s included? We didn’t want a to build Trix with tables. We had Lexical and we had agents to help, so we wanted to be ambitious here. We went for the whole package. Features In terms of major features: A color highlighter, built in instead of this being a custom Basecamp extension, as it was with Trix. Tables, with an interface we worked hard to keep simple and accessible. Markdown. You type it, you get rich text. Code blocks with syntax highlighting as you type, in more than twenty languages. Image galleries you can navigate and reorder with the keyboard. Prompts. Type a character, get a menu: mentions, emoji, or whatever your app needs. Links by pasting a URL over selected text. Previews of attachments like videos and PDFs, rendered as your app renders them. Action Text Native Lexxy is also Action Text native. Action Text stores attachments in a canonical format that Trix doesn’t speak, so it translates on save and again on render. We taught Lexxy to emit exactly that markup. What you see in the editor is what gets saved, and what gets saved is what your app renders. Your existing content, attachments and views keep working. That opened another door. Action Text now talks to an editor adapter, with an implementation for Trix and one for Lexxy, so switching is one line: config.action_text.editor = :lexxy. Here the credit goes to Sean Doyle, who started that pull request before Lexxy existed and took it to the finish line with our input. It ships with Rails 8.2, and we hope other editors will use it too. Extensibility And you can extend Lexxy. Extensions are built on Lexical’s own mechanism, and this is not a second-class API: Lexxy itself is thirteen extensions, tables included, and Basecamp has nine more. class MyExtension extends Lexxy.Extension { get enabled() { … } get allowedElements() { … } get lexicalExtension() { return this.defineExtension({ name: "my-extension", nodes: [ … ], register(editor) { … } }) } initializeToolbar(toolbar) { … } dispose() { … } } Lexxy.configure({ global: { extensions: [ MyExtension ] } }) My favorite of how extensible is Lexxy are voice notes in Basecamp: you record, you see the waveform while you talk, and it becomes a player inside the document. About a thousand lines, without forking or patching anything. Performance Lexxy is fast, because Lexical is fast. In a ten thousand word document, Trix takes thirty-eight milliseconds to process a keystroke. Lexxy takes four. Above fifty milliseconds, the editor starts feeling sluggish. Compared to Trix, Lexxy brought a whole new performance regime. Milliseconds per keystroke by document size Accessibility Accessibility in Trix was not great. In general, building accessible experiences for rich text editors built on top of contenteditable is quite hard. We had a dream team to help with Lexxy accessibility. Bruno Prieto worked with Michael Berger, our accessibility champion at 37signals, to bring the bar to where we wanted it to be. Bruno is an outstanding programmer who happens to be blind, so he knows one thing or two about accessibility, and he delivered. As a result, in Lexxy everything is reachable with the keyboard. The editor announces itself properly to screen readers, and it gets a thousand details right so that the editing experience using a screen reader is fantastic. You can learn more about accessibility in our docs. Security The latest AI models have resulted in an unprecedented explosion of vulnerabilities found, and we took this thread quite seriously. Lexxy counted with programmers of the caliber of Jeremy Daer and Mike Dalessio helping to make it more secure. We have put a lot of attention to sanitizing the editor contents, validating attachment URLs and making sure that the types of attachments and nodes the editor support are allow-listed. Lexxy also comes with preliminary Trusted Types support, to offer CSP-level control over certain DOM manipulation APIs. The trusted types policy is there, but we are not enforcing it everywhere yet. Agents We started Lexxy using Claude since day one. A main lesson was that an agent needs to drive the editor like a user does. The best decision we made in this project was moving the system tests from Capybara to Playwright: three browsers instead of one, a suite that runs in seconds, and a much more faithful clipboard, keyboard and focus. This represented a tremendous improvement in how agents could close the loop by themselves. Write a test, see it fail, fix it, see it pass. We have more than six hundred browser tests today. Moving the suite to Playwright changed how fast we could write tests With a solid testing foundation in place, we could start fixing bugs in large batches. As mentioned, getting a text editor right implies a ton of work, and the kind of backlog we got at some point would have have buried us in pre-agent times. We also used agents to validate the fixes: an agent reproduces the bug in the public Lexxy sandbox, checks that it’s gone with the branch applied, and labels the pull request. Agents were essential to get Lexxy done with the people and the deadlines we had: we are a small company, and the same people were shipping two products in parallel. 275 cards closed The new Rails default We believe Lexxy is the best rich text editor out there right now, and we are going to make it the default editor in Rails next. If you’re starting a Rails application today, use Lexxy. If you’re using Action Text with a standard configuration, switch. It’s one line, and we’ve worked hard to make it seamless.
Reading my recent computing retrospective, I realised there was a big section missing: the people in my life that made an impact and helped shape my career. Outside my immediate family, one person made an outsized contribution, and I’m fairly certain that without his influence my life would have taken a very different path. The fact that I’m still here in 2026, still writing code and being fortunate enough to have a career in something I love is testament to him. So I’d like to take a few moments to talk about my old secondary school teacher, George Dryden. Denied Back in 1995, I had a problem. I knew I wanted to study computing at university and build a career out of my passion, but there was a snag. For those unfamiliar with the UK schooling system, when you’re 15-16 you take a set of GCSE exams in a broad range of subjects. After that, you pick around 3 subjects to really focus on over a period of 2 years. These are called A Levels, and they are a big step up and are meant to prepare you for a degree-level course at university. Admission to university is also governed by these results - if you want to study computing, you’re going to need a computing A-Level, and most universities will only accept you (or “make an offer”) if you achieve a certain grade. And whilst I had taken computing at a GCSE level, my school did not offer a computing A-Level course. I instead had to settle on “Design & Technology”, which just didn’t inspire me. Instead of working on my portfolio and projects, I spent most of my time daydreaming and writing code on the Acorn Archimedes computers that were the staple of every 90s UK school. No disrespect to the teachers - they were all awesome - but it just wasn’t for me. I was miserable, and by the end of my first year, I was well on my way to failing outright with my entire future plans seemingly going up in smoke. Someone noticed That’s when George stepped in. He’d taught me computing right the way through my GCSEs, and with no A-Level course on offer, that was officially where his involvement was supposed to have ended. It didn’t. He had noticed my constant presence in the computing labs - before and after school, during lunch breaks, free “study” periods - working on some little pet project or digging into RISC OS internals. I remember him as warm, with a wicked, dry sense of humour, and a refreshingly spiky attitude to authority - I always got the sense he’d worked out for himself which rules were worth taking seriously and which ones weren’t. And he always had time for me. I spent years pestering him with questions that had nothing to do with anything on the syllabus, and he’d always find a way to answer them that actually made sense. He was just as supportive of my odd little obsessions. At one point I’d got deep into the BBS scene, which I thought was the coolest thing ever, and decided what the school really needed was an internal BBS running on its own network. So I wrote one. It was deeply cringeworthy, obviously - but George helped me put posters up around the school advertising it, and even gave it a mention in assembly one morning. I think about five people in total ever checked it out. It didn’t matter: a teacher had stood up in front of the entire school and treated my weird little project like it was worth something, and that was a hugely validating moment for me. Off the books He recognised the passion, and eventually he made a suggestion: What if I quit the Design and Technology course, and instead attempt the A-level course myself? Personally, I also suspect he was enjoying himself. There was some internal school politics behind why computing wasn’t offered at A-Level in the first place - I never knew the details - and I think the prospect of one of his students simply going out and getting the qualification anyway appealed to him on two separate levels. It would get me where I wanted to go, and it would wind up exactly the right people. It wouldn’t be easy, he warned. The school would be against it, plus it was a two-year course which I’d have to cram into one year. I’d have to do it all myself - studying, lesson planning, coursework - he couldn’t help me in an official capacity, but he said he’d advocate for me and help where he could. If I had assignments, he’d send them off to be graded and would give feedback in his own time. He’d enter me in for the exams and also gave me a set of keys to the computer lab so I could use it whenever I needed. It was the first time anybody outside my own family had really shown faith in my abilities and encouraged me to take a stand. It was a pivotal moment for me - I realised if I wanted something badly enough I would have to fight for it, but I could still make it happen. I didn’t have to take “NO” for an answer - a lesson I think I picked up from watching him as much as from anything he ever actually said to me. After a few weeks of dithering, I took the jump. I remember a few awkward meetings with the school administration but thanks to his behind-the-scenes work, I was soon following my dream. An intense year And yes, it was bloody hard work. I had to condense an entire two year course into under a year, be disciplined enough to produce my own study plan, and be critical enough of my own shortcomings that I could focus my study where it was needed. I pretty much lived and breathed it for months straight and was more-or-less a permanent fixture in the labs or school library poring over my course books. I’d make lists of questions and chat to George over lunch, and he’d provide guidance and encouragement. It was a lonely way to learn with no classmates to compare notes with, no lessons to turn up to, and right up until the end I had no real idea whether any of it was good enough - but bit by bit, it started to feel like something I could actually pull off. And sure enough, in the late spring of 1996, I sat down in an exam hall with my fellow students, the only one with an A-Level computing question paper in front of me. The final exam went by in a blur - I can only remember a few of the questions now (and a peculiar obsession with the Pascal language) - but I do remember the euphoria as the invigilator called “time’s up, pens down, close your papers NOW”. I had done it. A few nerve-wracking months later, my Mum drove me into school to pick up my results. I ripped open the envelope and saw it - I’d passed with an A grade! I literally ran up the stairs to George’s office next to the computing labs to thank him personally. I’d taken my camera into school to take a few last photos for memory’s sake and snapped this photo of him before I walked out the school gates for the last time: A different path Because of him, I managed to get into my university of choice, studying computing with a focus on networks. Because of that, I landed my first job working as a “webmaster”, and my career since has been one of the highlights of my life. All these years later, it’s a real privilege to be able to get up each morning and actively look forward to working in an industry I love. Without George stepping up for me and encouraging me to believe in myself, none of that would have happened. I wouldn’t have had the career I have, and I wouldn’t be where I am now. I met my wife when we both worked at a software company - she sat at the desk behind me - so even my home and family life can be traced back to that spring of 1996. And the A-Level itself was only half of what I took away from that year. The qualification opened the door to university, but the lesson that came with it was every bit as important: that a “no” isn’t always the end of it, and that sometimes the answer can be argued with. I’m so proud of what I managed to achieve all those years ago, and even more thankful to have had someone like George in my life to put me on the right track. Mr. Dryden I did see him again after I left. He drank in one of my local pubs - a pub I’d been going to for a good while before I was technically old enough to be in it - and I’d say hello if I spotted him in there, mostly in the months before I moved away to university. After that it was only a handful of times. For years I’d find myself scanning the room whenever I was back home and in for a pint, half expecting him to be at the bar. At some point I stopped seeing him altogether and eventually moved across the country. The trouble was I never really knew how to talk to him outside of school. He was always Mr. Dryden, or just “Sir”, I don’t think I ever once called him George to his face! I was (and still am if I’m honest) fairly socially awkward, and I never worked out how to phrase the thing I actually wanted to say: that he had changed the entire direction of my life, and I wasn’t sure he knew it. So instead I’d say hello, and ask how he was, talk about nothing much, and go back to my friends. Epilogue Sadly, 3 years ago now, I opened the latest issue of my old school alumni newsletter to read that he’d passed away. The photo at the start of this article was taken from his obituary article and I read that he’d had a long illness and had suffered from dementia at the end. I did write to him years ago by email - I don’t know if he ever got it, or was in any capacity to understand what he’d done for me, but I hope so. There’s an old saying by one of my favourite authors (Terry Pratchett) that “no one is finally dead until the ripples they cause in the world die away”. In one of his books, a character keeps the memory of his son alive by passing his name along a series of telegraph towers. It’s in that spirit that I’m writing this post - I debated it for many years as it’s very personal to me and I also have no contact with any of George’s family so I have no idea what they would make of it all. But even though it’s 30+ years ago now, I will never forget him or what he did for me - and at least now, if someone searches his name it’ll be recorded here for as long as I’m alive and running this site. Thank you, Sir. George Dryden 1942-2023
Let’s step inside the kernel and understand how it implements copy-on-write and what are its implications for the performance of user-space systems
Yesterday, I received this email as a response to You Can't Vibe Code Love. It's such a remarkable and powerful statement that I asked permission to share it here, in its entirety, with personal information redacted: Hey Jeff, Hope you and your family are doing well.
A frustrated Reddit post about being a condom between an AI and production made the rounds in our team. Here is why I think the opposite is true and what it means for how we review code, plan work and think.