Full Width [alt+shift+f] Shortcuts [alt+shift+k]
Sign Up [alt+shift+s] Log In [alt+shift+l]
1

Embedding the Sky

from John Wiseman [alt+shift+b] in technology

Two projects at the intersection of geospatial data and AI. First, a semantic search engine for aerial imagery, fine-tuned on OpenStreetMap-tagged data. Second, WARP—a custom neural model that learns embeddings of aircraft flight trajectories, trained on roughly a million aircraft tracks. Given any flight path, WARP can instantly retrieve the most similar trajectories from history. It enables applications from OSINT and anomaly detection to safety analysis and flight planning. The same latent structure that made word2vec transformative for language now becomes available for aircraft behavior. Presented at Spec.LA on December 18, 2025. I'm going to quickly show two things tonight. The past few years I've been working mostly in geospatial, aviation, and AI, and that's what these are. One is geospatial semantic search and the other is a foundation model for aircraft. I believe they're both extremely valuable and useful. # Both projects are about embeddings. Embeddings are points inside a latent or conceptual space. Parts of that space that are conceptually similar are closer together. There are neural models that will generate an embedding of an image or text or other things. With the right model, an embedding of a photo of a golden retriever is close to the embedding of the text "golden retriever" and far from the embedding of an image of an aircraft boneyard. # I showed an early version of PIMINTO at Spec a long time ago. I've refined it a lot since then. It uses a model that's fine tuned on satellite imagery and text describing the imagery based on OpenStreetMap tags. I'm going to give a live demo of the new site, which is public. piminto.obliscence.com | Demo video # Next is WARP, which is a model I'm building to create embeddings of aircraft paths. Back to aircraft, big surprise. # I trained the model on about a million aircraft...
27th Feb 2026

Stay updated

Get a weekly newsletter with the top 5 articles worth reading every week.

More from John Wiseman

Comparison of Models Used for Aviation Briefings

I have a collection of prompts and tools that I give to LLMs to generate daily aviation briefings. The briefings are supposed to be a summary of everything interesting that’s happened in the air in the past several hours, either globally or in a region of interest. The model should contextualize aircraft activity in terms of behavioral analysis, historical patterns, geography, and current events: The fire west of Piñon Hills stood at roughly 2,700 acres and 0% containment in the latest reports, with cooler overnight conditions helping crews strengthen lines. This morning’s air picture is recon and rotary rather than Saturday’s tanker parade: N8PQ, the Aero Commander 690A, is back overhead at 8,000 ft, K-MAX N107MW is working the fire at 4,200 ft, Bell 205 N33HX is staging at Fox Field, and Tanker 104, an Erickson Aero Tanker MD-87, sits on the ramp at San Bernardino International. Writing a worthwhile briefing is a challenging task. It’s agentic, with models using 15-20 tools including querying databases, searching the web, and writing and executing scripts to process data. It also requires judgment from a model on what to include in the briefing and how to put together separate pieces of data to tell a story without overstating a conclusion. The agents filter through a large amount of data looking for interesting items, find corroborating sources, take screenshots, and edit it all into a short (less than 6000 characters) report. I’ve tried Claude Opus 4.8, Fable 5, OpenAI GPT-5.6-Sol, and Kimi K3 on this task. The tl;dr if you’re paying retail API token prices: Opus 4.8 is cheap, fast, and its output is good. Kimi K3 costs about the same as Opus 4.8 and its output is of similar quality (it’s saved by the fact-checker) but it takes 5-6x as long as Opus 4.8. Fable 5 costs 2-5x as much as Opus 4.8 or Kimi K3. I like the feel of its output best, but it needs to be reined in by the fact-checker almost as much as Kimi K3. GPT-5.6-Sol is cheap and fast and makes few mistakes, but its briefings are too short and boring. I think GPT-5.6-Sol is the smartest model available at the moment, but for this task Fable 5 does better, and Opus 4.8 does almost as well for far cheaper. Some numbers The initial prompt for the task is about 25 K - 30 K tokens, depending on whether the briefing is for the world or just Southern California. (All token calculations are done with the OpenAI tokenizer; others may be different.) Runbook Tokens briefing-runbook-common 16,516 briefing-runbook-world 13,635 briefing-runbook-socal 8,039 The custom tools to give the models access to aircraft and geospatial data, and the guides on how to best use them consume another 50 K tokens. Custom MCP Tokens MCP Description 433 Tool Descriptions (10 tools) 3533 Resources (13 documents) 45713 I’ve done six briefings with Fable 5 and two briefings with each of the other models. The following table lists time and cost for a single briefing, for each model. Time includes the total time for research, review, corrections and posting. Costs are API-equivalent estimates based on current pricing with token caching included. Model # Briefings Median time (minutes) Median cost Opus 4.8 2 24.8 $10.06 Fable 5 6 31.5 $30.33 Kimi K3 2 132.4 $10.43 GPT-5.6-Sol 2 40.0 $15.63 After a model writes the first draft of a briefing it sends it and the supporting material to an adversarial fact-checker powered by GPT-5.6-Sol. The model then rewrites the report based on the fact-checker’s findings. For each briefing I classified the fact-checker’s findings into “major”, more substantive corrections, and “minor”, softer issues. The table below shows the median number of writer tool calls and the mean number of corrections per report for each model. Model # Tool calls Major corrections Minor corrections Opus 4.8 66 0.0 4.0 Fable 5 93.5 2.5 7.0 Kimi K3 125 2.5 7.5 GPT-5.6-Sol 116 1.0 2.0 The GPT-5.6-Sol reports used separate GPT-5.6-Sol reviewer sessions, following the same adversarial process as the other models. The review step added a median $1.95 worth of GPT-5.6-Sol tokens to the cost of a report, and review costs are already included in the costs listed above. Report quality Opus 4.8 Opus 4.8 did a good job of creating rich reports with strong themes, and it showed good editorial judgment by removing weak leads not supported by further research. As an example, it might lead with a developing emergency: N969WR … was cruising at FL450 across the Oklahoma/Texas panhandle when it squawked 7700 at ~2204Z and began a continuous descent, averaging around 2,300 ft/min. It combined that lead with a large firefighting mobilisation, an EA-37B, an E-6B, allied tanker movements and forward-looking airspace notices. This was varied and interesting without feeling like a list of unrelated aircraft. The fact-checker didn’t refute any of Opus 4.8’s conclusions, and didn’t find any unsupported claims. Corrections were mostly wording changes to avoid implying unwarranted precision. Fable 5 Fable 5 produced the most detailed and ambitious reports. It was strongest when it could connect several observations into a larger story. For example, the 15 July world report described a shared GPS spoofing pattern near Smolensk: Three airliners (Turkish, Belavia, Air Serbia) each plotted flying an impossible 1.2-nm, 55-kt circle around the same fixed point southwest of Smolensk while MLAT put the real aircraft 400 km away. The 14 July SoCal report turned an anonymous military track into a strong lead: A silent visitor working W-291 — ae685e, a US military hex with no callsign, came down from Oregon overnight at FL270 and has flown a broad 21,000-ft circuit off the San Diego/northern Baja coast for hours. Fable 5 also followed stories across several days. It tracked tanker relays, airlift movements, range activity and unusual aircraft between reports. Fable 5’s reports were detailed, visual, and interesting. They also needed significant corrections, an average of 2.5 substantive and 7 minor problems per report. Its main weakness was overprecision. Codex regularly corrected: exact start and end times track boundaries aircraft counts claims about every member of a group statements that joined events across an anchor time mission interpretations that went beyond the observed movement An example of a significant error (that GPT-5.6-Sol let pass but that I found on review) was its report on a Russian Be-200 amphibious aircraft. Its published heading said: A Russian EMERCOM Be-200ChS scooping the Donbas. The report described the aircraft as working “a 300-km chain of fire sites”, but when I saw the map that seemed implausible. I asked Fable 5 about it, and eventually it posted an update: The evidence now says ‘repeated brief landings at unmapped sites, purpose unknown’. I still don’t know what that aircraft was doing, but it’s an example of Fable 5 going significantly off course. Kimi K3 Kimi K3 is really slow, and the fact-checker had to correct it more than any other model, but the quality of its (corrected) output once it was finally done was excellent–as good as Opus 4.8. KC-135 declares an emergency over the Irish Sea, home safe at Mildenhall. Fox Field staged and scattered — seven fire aircraft on the ramp on the fire squawk all morning; the 737 Fireliner and RJ-85 launched ~10am and ferried north without a drop run. There’s just no reason to use Kimi K3 since it costs as much as Opus 4.8, is about as good, but much slower. GPT-5.6-Sol GPT-5.6-Sol shows so much restraint that its briefings are boring: dry, no flavor, and not particularly interesting. LATTE18 / N650RX, a trustee-registered Global 6500, flew roughly 4 broad clockwise circuits at FL360 from 01:28 to 02:14Z, then departed southwest. The direct ADS-B track measured about 38 by 16 nautical miles. Public aviation reporting associates this tail with SNC’s Army ATHENA-S fleet; SNC confirms its 2 ATHENA aircraft support US Army airborne ISR, but SNC does not name the tails. The same offshore block appears repeatedly in the recent archive. I generally like GPT-5.6-Sol but its briefings were just too dull.

18th Jul 2026 • 1 votes

More in technology

Computational tools for society’s most complex challenges

Associate Professor Cathy Wu uses reinforcement learning to help map out improvements to transportation and other multifaceted systems.

19 hours ago • 1 votes
Two more 27.0 design grumbles

After living with Apple’s 27.0 OSs since launch, I have some more annoyances to get off my chest. This time, it’s all about how tabs and menus have gotten worse. I’ve already ranted about the Liquid Glass material in general, but these two design changes in particular have really been grinding my gears. I’ll reiterate that Apple’s latest OSs look substantially nicer to me than the previous set… but that only makes these setbacks more glaring. Also, many of these issues aren’t nearly as bad in light mode — but I use dark mode exclusively on all platforms. Apple offers this appearance setting, so I think it’s fair to criticize them when it’s not holding up. First up, let’s talk tab bars. I think these looked awful in the original Liquid Glass redesign, and in 27.0 they look even worse. Below is an example of three tab bars from Safari in macOS. All of them are in dark mode. The top example is from macOS 26 with the “clear” Liquid Glass setting, the middle is macOS 27 with the default (mid-slider) version of Liquid Glass, and the bottom is macOS 27 with Liquid Glass at its most tinted. In each, the middle tab is selected (though I think the word “tab” is being quite generous to these globs). In macOS 26, there was practically no difference between the clear and tinted versions of tabs. Similarly, the clearest and default/middle tabs in macOS 27 are effectively the same. Because of this, I’m leaving out the redundant examples. Even though I still think it’s ugly, I vastly prefer the macOS 26 version of these three options. It offers the most contrast, and it makes more sense in dark mode: the background is darker and the foreground of the active tab is lighter. The middle example is what tabs look like in the default (mid-slider) version of Liquid Glass in macOS 27. There’s now only a very faint outline around the active tab, and practically no difference in background colours. To me, this is unreasonably subtle. The effect is even worse when there are a lot of tabs open. Lastly, there’s macOS 27 with the fully tinted Liquid Glass setting. It’s better, but it still looks less correct to me than the tab design from macOS 26. It’s difficult to put into words how much I loathe the look of this new tab bar design. I don’t mind the more “bubbly” look of Liquid Glass throughout the 27 OSs for the most part. It gives UI elements more dimension than in the 26 OSs. But it doesn’t work for tabs. Because the bubbly look is inside a trough, the active tab’s glass effect ends up looking like a blur on the top and bottom. This reduces contrast further and makes the active tab harder for me to pick out. Even in dark mode with full tint, I find the active tab less visually clear than in 26’s clear mode. Now, I’m sure there are at least a few people reading who don’t see what the fuss is about. If that’s you, I assure you that the difference is more stark when you’re not comparing things side by side. It’s not impossible for me to pick out the active tab, but I think it’s trickier than it needs to be! But, if you still don’t believe me, here’s a little experiment. Which of these do you think is most legible? The text/background colours in the above image are based on the foreground/background colours used in the tab bar instances above. First is clear in macOS 26, then the default from macOS 27, then fully tinted in macOS 27. I think they’re all pretty awful, but I prefer the macOS 26 clear version. Again, this is because I’m using dark mode. In dark mode, light text appears on a darker background. Similarly, active UI elements have a lighter background than their surrounding elements. I’m sure there are counter-examples, but this is how just about everything else works in Apple’s own apps! It should be noted that Safari uses the system default tab bar design. I also see this design in Apple’s Terminal app, in Pixelmator Pro, and elsewhere. I don’t use Xcode daily anymore, but you’ll also find them there — though in true Xcode fashion, they’re ever so slightly nonstandard and also don’t respect your tint setting. Below is a screenshot of Xcode using my current settings of dark mode with fully tinted Liquid Glass. Up next: menus. Below is an image showing four versions of the same menu in macOS. Top left is macOS 26 clear, top right is macOS 26 tinted, bottom left is macOS 27 default (mid-slider), and bottom right is macOS 27 fully tinted. It’s a similar story here. In macOS 26’s dark mode, I had no problem with system menus even when Liquid Glass was set to clear. In macOS 27, even in the fully tinted mode, the menus have much lower contrast. They also now lose all of Liquid Glass’s refractive effects when at their most tinted. I think this is less of an issue than the tab design changes, but it’s still a downgrade. Again, I’m certain many people don’t care about this. Some might wonder why I’m not turning on accessibility settings to help with these things, if they bother me so much. I’ve flirted with this (especially the “Increase Contrast” setting), but those settings have many knock-on effects. 1 But honestly, I don’t think accessibility settings should be required to have a reasonable amount of contrast in a design system. Maybe Apple disagrees, but I really hope more dark mode tweaks are coming. The “Increase Contrast” setting is under System Settings -> Accessibility -> Display -> Increase Contrast. Interestingly, you can use this setting along with the clearest version of Liquid Glass to almost get back to how things looked in macOS 26’s version of tinted. However, it adds contrast-y lines around many UI elements that I find extremely distracting. It also alters colours on some elements to, strangely, make them less contrast-y. It feels unevenly applied and poorly implemented in several apps. ↩

22 hours ago • 1 votes
Does AI exacerbate English-language inequality?

It’s an empirical question!

yesterday • 1 votes
The 2030 Census is in Trouble

Trump wants to add a citizenship question and ban questions on race

yesterday • 1 votes
Make tmux the OS

I recently watched the talk by Scott Jenson titled "Are we really going to use the same Desktop UX forever?" https://www.youtube.com/watch?v=V7AfAcQwLW0&t=445s. He's a great presenter, really articulate and concise. The kind of speaker that you'd

yesterday • 1 votes
📚 BoredReading

You seem to be enjoying this.

Join free to unlock everything.

Create free account

Already have an account? Sign in