Full Width [alt+shift+f] Shortcuts [alt+shift+k]
Sign Up [alt+shift+s] Log In [alt+shift+l]
118
Introduction IMPORTANT: Use the Right GPS Antenna! The Problem: SyncServer Refuses to Lock to GPS The GPS Week Number Rollover Issue Making the Furuno GT-8031 Work Again How It Works Build Instructions Power Supply Recapping The Future: A Software-Only Solution The Result References Footnotes Introduction In my earlier blog post, I wrote about how to set up a SyncServer S200 as a regular NTP server, and how to install the backside BNC connectors to bring out the 10 MHz and 1PPS outputs. The ultimate goal is to use the SyncServer as a lab timing reference, but at the end of that blog post, it’s clear that using NTP alone is not good enough to get a precise 10 MHz clock: the output frequency was off by almost 100Hz! To get a more accurate output clock, you need to synchronize the SyncServer to the GPS system so that it becomes a GPS disciplined oscillator (GPSDO) and a stratum 1 time keeping device. The S200 has a GPS antenna input and a GPS receiver module inside, so in theory this should be a matter of connecting the right GPS antenna. But in practice it wasn’t simple at all because the GPS module in the SyncServer S200 is so old that it suffers from the so-called Week Number Roll-Over (WNRO) problem. In this blog post, I’ll discuss what the WNRO problem is all about and show my custom hardware solution that fixed the problem. IMPORTANT: Use the Right GPS Antenna! Let me once again point out the importance of using the right GPS antenna to avoid damaging it permanently due to over-voltage. GPS antennas have active elements that amplify the received signal right at the point of reception before sending it down the cable to the GPS unit. Most antennas need 3.3V or 5V that is supplied through the GPS antenna connector, but Symmetricom S200 supplies 12V! Make sure you are using a 12V GPS antenna! Check out my earlier blog post for more information. The Problem: SyncServer Refuses to Lock to GPS When you connect a GPS antenna to a SyncServer in its original...
18th Aug 2024

Stay updated

Get a weekly newsletter with the top 5 articles worth reading every week.

More from Electronics etc…

Refurbishing a Tektronix TDS7104 Oscilloscope

Introduction The TDS7104 Common Failures Make an Image of the Hard Drive CR2032 Backup Battery Replacement and Display Brightness Do NOT Remove the Front Panel Scope Disassembly A Failed Attempt at Switching over to an SSD Reinstalling from Scratch: Windows 2000 Pro or Windows XP? Not All TEAC CD-224E Drives are the Same Installing Windows 2000 Pro on an Old Machine through a Virtual Machine Burning the Windows 2000 Pro Installation Disk onto a USB Stick Booting from USB Stick with a Plop Boot Manager Installing Windows 2000 Pro Installing Special TDS7104 Drivers Installing Tektronix Firmware The Scope is Working! Re-enabling the Existing License Cleaning Up The End References Footnotes Introduction A little over a month ago, I ran into a Tektronix TDS7104 at the Silicon Valley Flea Market, where else? Other than some dirty buttons, a few smudges here and there, and the usual assortment of calibration and asset tracking tags, the unit was in excellent cosmetic shape, but the price tag of $700 was way out of line: as I write this a try-before-you-buy TDS7104 can be had on Craigslist for the same price. But Paul, the seller/liquidator, has a habit of saying “I’ll make you a deal” and he did before I even asked: $300. That’s still a lot by flea market standards, but a pretty good price for a TDS7104… if you can get it to work. At home, the scope powered up right away and it booted straight into the main scope application. Other than a screen that was way too dim, everything seemed fine. But when I tried it again a few hours later, it got stuck at the BIOS screen with a CMOS battery error. In this blog post, I go over the steps I took to get the scope back in top shape. The TDS7104 The TDS7104 is a 4-channel oscilloscope with 1 GHz bandwidth and a maximum sample rate of 10Gs/s, though that’s only possible when using 1 channel. The sample rates drop to 5 Gs/s for 2 channels and 2.5 Gs/s for 3 or 4. Even by today’s standards, the specs exceed those of hobbyist class oscilloscopes, think Rigol and Siglent, though there’s a price to pay in terms of weight, 39 pounds, and volume: they’re as wide and deep as the earlier TDS700 series, for example, and much taller. The TDS7054 is its little brother, figuratively speaking only. In the same chassis, it has a 500 MHz bandwidth and 5 Gs/s. Unlike more advanced TDS7xxx models, the 7104 and 7054 have BNC connectors instead of custom Tektronix ones that require probes or adapters with prices that exceed today’s price of the scope itself. Introduced mid 2000, these scopes initially ran Windows 98 but they must have upgraded soon after to Windows 2000 Pro Embedded, because that’s what mine has and it has components with a late 2000 timestamp. The PC motherboard has the little-used NLX form factor. Mine was a RadiSys SF810 with a Socket 370 and a 100 MHz front-side bus. Originally, these scopes shipped with a dog slow 550 MHz Celeron, I got lucky with a 850 MHz Celeron. The fastest compatible Celerons with 100 MHz FSB go up to 1.4 GHz, but they’re pricy. You should be able to find 1.1 GHz versions for around $20 on eBay. Unlike my Agilent 54831, the 640x480 LCD screen has resistive touch control which makes it possible to use the advanced scope features without the need to connect a mouse. In addition to the PC motherboard, there is a PowerPC-based controller board that runs VxWorks like many other Tektronix products of that time, and a large acquisition board. In addition to a few hardware options such as a 4M/channel sample memory, up from a 500k default, there are plenty of software options for advanced measurements: jitter testing, USB certification testing, etc. Both the software and hardware options can be enabled with a license key. To the suprise of no one, that protection scheme was hacked long time ago… According to the labels on the chassis, my scope came from the PSD lab at Cypress Semiconductor, where it was used for things like measuring high-bandwidth signals such as the battery current on the Apple TV Remote. :o) Common Failures As always, you’ll find a bunch of hobbyists trying to revive this kind of scope on the EEVblog forum, Youtube and some blogs. Here are the most common failures: PC motherboard CMOS backup battery dead PowerPC backup battery dead Hard drive dead Power supply capacitors leaking I was lucky and only had to deal with issues 1 and 3, sort of. Make an Image of the Hard Drive Whether the machine boots or not, your first step should always be to make an image of the hard drive, a 6 GB IBM Travelstar in my case. Like my Agilent 54831, I thought that I’d have to open the case to access the drive, but you can just push on the spring-loaded black cover in the back and pull the drive sled out1. Nice! Remove the drive from the sled, plug it into a USB-to-IDE adapter, and extract the data. On Windows, I use HDD Raw Copy Tool. I often use Linux for this kind of maintenance, but since this scope is a Windows 2000 machine, I ended up needing a bunch of Windows-only tools. The Travelstar HD was running on fumes, because HDD Raw Copy Tool ran into a number of corrupt sectors during the copying operation. I was lucky, the impacted files were related to the French Windows 2000 manual, but it shows the importance of making an image of the drive ASAP. CR2032 Backup Battery Replacement and Display Brightness You’ll need to open up the case to get to the PC motherboard CR2032 backup battery. See the next 2 sections for that. After installing the new CR2032, the scope booted back up again, but the TekScope window had some weird corruption and waveforms didn’t render right. This was because the Chips & Technologies 69000 graphics card settings had been changed to a 256 color palette mode. It needs to be set to True Color 24-bit mode.2 Notice the presence of 2 video cards: an Intel 810 integrated graphics card and the Chips & Technologies 69000. The latter is responsible for driving the LCD screen. It has special hardware to render oscilloscope waveforms in overlay mode: they are sent by the acquisition board to the video memory through DMA3 without CPU intervention. While we’re on the topic of the display: after installing the CR2032, the LCD display was still very dim, to the point that I was researching replacement CCFL backlight tubes. That turned out to be entirely unnecessary: the TDS7104 doesn’t have a way to control the intensity of the LCD backlight. The previous users must have used it in a dark lab and dialed down the brightness by adjusting the gamma settings in Windows: Do NOT Remove the Front Panel I’m putting this section before the Disassembly one to make sure those with a low attention span get the message: chances are high that you don’t need to remove the front panel. And that’s good because, unlike the TDSnnn series scopes, the front panel has some plastic tabs that are very easy to break. That said, even if you do break them (I did!), the result is not catastrophic and you should be able to put the panel back firmly where it belongs with no one noticing a thing. (Click to Enlarge) The TDS7000 Series Service Manual makes it sound easy enough: To remove the trim ring, slide the flat end of a soldering aid into the side slot on the trim ring. Press in, then lift up to hook it underneath, then pry up. And from the pictures, it’s as if you can remove the front panel without removing anything else. That just didn’t work… The front panel consists of multiple click tabs: 1 on the left side, 1 on the right and then a bunch at the top and the bottom. So far so good. However, the left and right side also have 2 slide tabs that go into the metal rails. If you lift the left and right tabs too much, these plastic slide tabs break off. So you need to be very careful to make sure that you don’t lift the plastic trim too much, and that you slide the panel out while it stays parallel with the display. Or… you don’t touch it: you can do all PC maintenance, including replacing the floppy drive, without removing the front panel. Scope Disassembly I will continue my tradition of documenting the disassembly of test equipment in too much detail because nobody else does it. Even though the service manual technically describes how to do it, a few pictures go a long way to make it easier. To access the inside of the scope, you need to remove more than 30 screws. On the plus side, they’re all Torx-15 screws and they’re all the same length, so you don’t need to worry about keeping track of which screw goes where. Still, it takes a while and it validated my recent purchase of this cordless screwdriver, recommended by Shrirar over at The SignalPath. Unbutton the accessory bag This took me longer to figure out than I want to admit: you can just unclick the bag from the chassis, but the buttons can be very tight and if you’re not careful the fabric can tear. Use a flat-head screwdriver right next to each button to lever it off. Put the scope upright on its back feet It’s an unusual arrangement, but the easiest way to dismantle the scope is by putting it on its back feet: you don’t need to remove any screw from the back! Let me once again sing the praises of a sturdy equipment cart: it’s so much easier to walk around the cart than to muscle around bulky, heavy test equipment on a table. Remove the top panel 4 screws through the accessory bag buttons (“snap studs”) fix the top panel to the chassis. After removing this panel, you could remove the side panels already, but I found it much easier to remove the bottom panel next. Remove the bottom panel and loosen the black front connector trim Next, remove the 5 screws of the bottom panel as well as 3 screws that keep the black trim of the front BNC connectors in place. The black trim doesn’t need to be completely removed, only loosened because otherwise it will soon be in the way of some other screws. The bottom panel shall now be removed though. Just slide it down a bit and take it off. In the picture above, you can see 2 screws that aren’t marked in red. That’s because they don’t keep the bottom panel in place. But if you feel like it, you might as well remove them now too. Remove the handle and side panels With the bottom panel gone, the side panels are a breeze to remove after unscrewing the handle. I lied: these 2 screws are different than the others. But they’re a different color and impossible to get wrong. Remove the 2 sheet metal parts With the outer covers removed, you’re now staring at the sheet metal RF protection enclosure. It consists of 2 parts, each part covers 2 sides. Remove all the screws, take off the bottom part and then the top. Note how some of the bottom screws are hidden underneath the BNC connector cover. That’s why you had to remove its 3 screws of the black trim. Congratulations! For those who didn’t keep track: you’ve removed 32 screws! After removing the panels, you now have access to the acquisition board at the bottom and the PC motherboard at the top: One side has nothing but cooling fans, but from the other side you can see the power supply and an RS-232 port that you will need to connect if the backup battery of the PowerPC controller board expires. If you need access to those items, you’ve only done the easy disassembly part. On my unit, both the PSU and the controller backup battery were fine so I was done. Note on the picture above that the front panel has been removed. You do NOT have to do this for pretty much all restoration cases! And you really shouldn’t. Reinstall the bottom sheet metal cover All of my work on the scope was on the PC motherboard and I had to put the scope back in its horizontal position. To make sure that I didn’t accidentally damage the acquistion board, I put the bottom sheet metal cover back in its place. A Failed Attempt at Switching over to an SSD I’ve been using CompactFlash cards in the past to replace ailing hard drives. They work, but unless you buy a more expensive “industrial” card, they don’t have built-in wear leveling support. That is not a problem on a Rohde AMIQ that runs DOS, but on an OS like Windows with swap space, it could be4. So this time, I chose a 64 GB mSATA SSD ($35) and an mSATA SSD to IDE 44 Pin 2.5” adapter ($15)5. The standard way to move away from a failing hard drive to an SSD is to once again use HDD Raw Copy Tool to write the image to the SSD and that is that. I tried that with the 64 GB SSD, and while the scope got past the first-stage boot process, it errored out during the second stage when it tries to bring up the Windows GUI with a STOP: c0000218 {Registry File Failure} error. Older systems often had issues with partitions larger than 32 GB, so I bought a 32 GB mSATA SSD instead, $3 cheaper for half the capacity, but I got the same error. Just copying the drive image to an SSD worked fine for others, but for me it was a dead-end that I spent many hours trying to get around. I eventually decided to reinstall all the software from scratch, which was a whole other adventure. Reinstalling from Scratch: Windows 2000 Pro or Windows XP? I had wanted to avoid reinstalling the OS from scratch because I expected to run into a bunch of driver issues, but in the end I had no choice. While a number of people have reported that Windows XP can work on some of the TDS7104 motherboards, I decided to stick with Windows 2000 Pro because I know that works and I didn’t have a pressing need for more functionality, whatever that might be. The scope has Windows 2000 Pro Embedded, but I wasn’t able to find an installation disk for that and the regular version works fine too. The ISO file can be downloaded from the Internet Archive. The license key that’s printed on the back to the scope does not work with the regular Windows 2000 Pro. The Internet Archive one has a key that works, and other valid keys are just a Google away, but I didn’t even need one: I was never asked for a license key during the Win2k installation on the scope. The standard way to install Win2k Pro is with a CDROM drive. Unfortunately, the drive didn’t work which meant I had to open the whole machine again to install a replacement drive. Not All TEAC CD-224E Drives are the Same The TEAC CD-224E laptop drive in my TDS7104 got detected just fine by the BIOS and in Windows, but when you inserted a disc in the drive, neither the BIOS nor Windows could read from it. Since the RadiSys motherboard doesn’t support booting from USB stick, I decided to replace the TEAC CD-224E laptop drive with a ‘new’ one that I got from eBay for $20. Unlike the hard drive, the CDROM drive can’t be removed without opening up the TDS7104, but once the case is open, the effort is minimal. I first removed the floppy drive to have a bit more maneuvering freedom with the cables, but it’s not really necessary. Unplug the CDROM IDE cable Remove 2 screws The CDROM drive sits in a metal enclosure with a small adapter PCB that converts the CD-224E 50-pin slimline IDE connector to a standard PATA/IDE connector. I tested the broken drive with the adapter PCB and my USB-to-IDE dongle on my laptop to make sure the issue was with the drive and not the CDROM disc, and that didn’t work, as expected. With the new CD-224E/dongle combo, my laptop could read the installation CD just fine, but when I installed the new drive in the TDS7104, the BIOS couldn’t even detect the drive! I tried every BIOS setting under the sun, but no luck. There are many versions of the CD-224E, all with the same dimensions and slimline IDE interface, but clearly they don’t all behave the same. The version of the broken one is version A93 (2000), the new one is CD0 (2005). You can find A93 drives on eBay, but $69 is way too high for something that I’d be using exactly once. Installing Windows 2000 Pro on an Old Machine through a Virtual Machine (Another dead-end) It is allegedly possible to install Windows 2000 Pro on an old machine without CDROM and USB port by using a virtual machine. The process is convoluted: mount the installation CDROM ISO and the SSD onto the virtual machine. go through the first phase of the installation process until asked to reboot. now move the SSD to the old the machine (the scope) and proceed with the installation there. I once again spent a few hours getting this to work, but the scope never managed to make it to the Windows installation GUI. Burning the Windows 2000 Pro Installation Disk onto a USB Stick Alright, so I’m running out of options and USB is about the only storage interface left. The scope can’t boot from a USB stick directly but there is a way around that. Let’s first create a bootable USB stick with the Win2k installation ISO on it. Most of the time, you can use a utility like Balena Etcher to burn a CDROM ISO onto a USB stick, but of course that doesn’t work for the Windows 2000 Pro installation CDROM. Instead, you need to use WinSetupFromUSB to prepare the USB stick: Download, install, launch Select the USB stick as target Select Auto format with FBinst and use the FAT32 file system Add to USB disk: Windows 2000/XP/2003 Setup Select the mounted Win2K Pro ISO drive as source Press “GO” to copy Win2K Pro onto the USB stick Booting from USB Stick with a Plop Boot Manager Plop Boot Manager makes it possible to boot from a USB stick on machines that don’t support it. It goes like this: copy the plpbt.img image from the plpbt-5.0.15.zip archive to a floppy disk with a tool like WinImage, Rawrite32, or RawWrite for Windows.6 boot the Plop Boot Manager from floppy disk. the boot manager has a USB mass storage device driver select USB as boot device I tried hard to avoid the floppy disk route because my experience with floppy drives on old test equipment has been abysmal: none of them worked. Having no choice, I tried to copy the boot manager image with my USB floppy drive and… that didn’t work either. All these years the USB floppy drive, freshly bought from Amazon, was the culprit! Since the scope still worked fine with the IBM HD, I used its own floppy drive to put the image onto the floppy disc and that worked. After setting the BIOS to allow booting from floppy, the scope booted into the Plop Boot Manager just fine and it was able to boot the USB stick with the Windows 2000 Installation ISO. Plop doesn’t support USB hubs. The RadiSys motherboard has only 1 USB port which will be occupied by the USB stick, so you’ll at least need a PS/2 keyboard to do anything. Installing Windows 2000 Pro With the empty 32GB SSD plugged into the scope, the installation of Windows 2000 Pro was uneventful. There are 2 phases: the first one uses text mode and primarily copies all the necessary drivers onto the SSD. The machine then reboots and continues the installation in Windows GUI mode from the SSD, though the USB stick is still needed in a later stage. The TDS7104 has a bunch of specialty hardware that needs dedicated drivers, but those are not needed to get the OS up and running. At long last, I was able to see this image: Installing Special TDS7104 Drivers There is a great GitHub repo with a bunch of TDS7000-series software, including this Drivers directory. The README.md says that the driver should work for Windows 98 and XP, but the Chips and Technologies video driver definitely did not work for Win2k!7 I used Driver Collector to extract drivers from the original hard drive and that worked fine. You can find these drivers here. The 4 specialty drivers are for these components: Front panel This is the USB Device that’s listed under “Other Devices” Texas Instruments PCI-1225 CardBus Controller You need to install this driver twice, once for each port. Windows installed a default PCI-1225 driver for this, but that one doesn’t work, hence the exclamation mark next to it. The name of the driver .inf file is unsup.inf, for unsupported? Confusing, but that’s the one to use. PCI2PCI bridge That’s the Other PCI Bridge Device. Chips and Technologies 69000 video driver The default Windows driver for the C&T 69000 is what makes the screen work when running Windows, but it’s not sufficient to render measured signals in the TekScope application. For that, you need to update to the Chips and Technologies (Asiliant) 69000 driver. Installing Tektronix Firmware The TDS7104 and TDS7054 firmware v2.5.5 can be freely downloaded from the Tektronix website. The installation was painless, just launch the executable. The TDS7104 has a convoluted architecture where the PowerPC on the controller board can access files on the hard drive of the regular PC that are located in the c:\vxboot directory. Since the controller backup battery on my scope was still in good condition, I didn’t have to do anything special: the vxboot directory was created automatically during the firmware installation. One thing that was missing, though, was the advanced jitter license option. The Scope is Working! And with that, I finally had a working TDS7104 with SSD! The time from pressing the power button to having a waveform on the screen was much lower too: from 2min50s down to 1min35s. Re-enabling the Existing License The same GitHub repo that I mentioned earlier also has an unlock options directory with scripts to enable and validate license key features. On the Eevblog forum, plenty of people have been able to use it, but it’s not as user-friendly as other license key schemes. Most of the time, license keys are additive, with one license key per feature that must be enabled. On the TDS7104, there is 1 license key that enables all features at once. The validate script shows how that works with the license key and serial number of my scope: ./validate.py BREHZ9885D3MNKXHHYQCQRGQRW7C E1 91 73 BF F7 7B E4 C5 52 3D C7 3A E1 9E 71 8F 76 01 44 2F 54 00 00 C0 1B 79 48 00 00 00 00 00 00 00 00 A8 16 30 00 10 00 00 00 00 00 This key is for UID 1BC00000542F (S/N 21551, model TDS/DSA/DPO7104): CRC: 4879 Key is valid, active options: 00 00 00 00 00 00 00 00 08 00 00 00 00 00 00 00 00 00 We can see how that long string of gibberish contains: the serial number 21551 the model number TDS/DSA/DPO7104 a UID that is really just a combination of the serial number and the model a CRC an 18-byte or 144-bit bitmask I can recreate the license key by feeding these parameters back in the generation tool: ./gen.py tds7104 B021551 000000000000000008000000000000000000 XBGDV-K8GDM-KH7X3-979Y9-ZZ593-9ZRZZ-4837X-9VV5Z-T9HB I had to join the 18 bytes into one 72-digit hex number. The license key that comes out doesn’t match the original one, but after entering it into my scope, it worked just the same: The scope is very forgiving about the license keys: upper case, lower case, dash or no dash, it all seems to work. You can even reduce the number of hex digits in the license enable mask to a certain extent, and the license key will still work: ./gen.py tds7104 B021551 000000000000000008000000000000 7GWUZ-RRRMK-59LYT-978Y8-GZD93-8ZQGZ-C836X-8CVD What remains is the question which bit maps to which feature? This post in the eevblog forum has you partially covered here: ######################################################################## 4 # options masks/names/descriptions, conversion functions 5 6 # 01 - 1M 7 # 02 - 2M 8 # 04 - 3M 9 # 06 - 2M 2A 10 # 08 - 4M 11 # 00 00 00 00 00 00 04 - USB 12 # 00 00 00 00 00 00 20 - JT3 13 # 00 00 00 00 00 00 00 80 - ET3 14 # 00 00 00 00 00 00 00 00 08 - JA3 15 16 # 00 00 05 00 00 00 00 00 00 10 - ASM DDRA DJA 17 # 00 44 00 00 00 00 02 08 - SM ST J1 J3E 18 # 04 40 00 00 00 00 06 C0 10 - 3M JT2 USB2 ST 19 # 04 44 FF 03 00 00 8D A3 EF FF 17 - 10XL, MTH, PTH1, ASM, LT, DDRA, SLE, EQ, TDSDDM2, TDSUSB2, YDSCPM2, RTE, IBA, PCI, TDSDVI, TDSET3, SAS, TDSHT3, TBD, JA3, TDSPTD, TDSVNM, DPOPWR, TDSHT3v1.3, 73, 74, DJE, DJA, 77, 78, 79, SVE, SVP, SVM, SLA Note how JA3, advanced jitter analysis, indeed has bit 14 set to 1. Some people just use a mask of FFFFFFF....FFFF. Some of these analysis tools can once again be found in the same GitHub repo, or on the Tektronix website. This is all theoretical, of course. I don’t think I’ll ever have a hobbyist need for any of this… Cleaning Up The final act is cleaning. This scope was in exceptional condition, except for the knobs on the control panel. The knobs have a thin anti-slip layer on them that is a finger grease magnet. Removing that layer with isopropyl alcohol makes the knobs look like new without a noticeable difference in control. Just be careful about using 99% isopropyl, I think it attacks the plastic. 90% was fine. If some knobs are missing or cracked, the ones of a TDS220 are identical. You can buy knobs new or on eBay, but they’re expensive. If you really need a few, you might be better off buying a donor TDS220 instead. The End And with that, the scope is ready to be deployed to a shelf in my garage. One day I’ll need something with this kind of firepower but for everything else, a small scope with lower specs is way more practical. I like the scope better than the Agilent 54831 so that will probably hit Craigslist at some point. All words in this blog posts were written by a human. References TDS7000 Series User Manual TDS7000 Series Service Manual Eevblog forum - Tek CSA7404/TDS7000 repair project This is the place to go. Chances are that all your questions will be answered here as long as you have the patience to read through 41+ pages of discussion. TDS7054 repair Github Repo with a lot of resources Footnotes I obviously only figured this out after removing the enclosure… ↩ I didn’t try the 16-bit not-so-true color mode. ↩ The Agilent 54831 uses a similar overlaying method. ↩ In reality, I will never use this scope enough to ever run into an issue like this. ↩ You can still find native 2.5” IDE 44 laptop SSDs, like this one, but you pay $30 more for the same capacity. ↩ RawWrite for Windows is the one to use on a 32-bit Windows system, like the Win2k OS on the IBM Travelstar of the scope. ↩ If you install the incorrect driver, the scope will still boot with a working LCD screen, but once the Windows GUI starts, it will move its business to the Intel integrated GPU. You need a VGA monitor to follow what’s happening. Even if you later select the right driver, Windows somehow thinks that the old driver is good enough and just doesn’t do it, without any feedback. I had to manually delete the bad driver files from the SSD to finally make it work. ↩

3 weeks ago 1 votes
A Galois Field Arithmetic Primer

Introduction A Galois Field Introduction by Example Base Galois Fields A Real World Example of a Base Galois Field GF(2) Extended Galois Fields Extended Galois Field Addition Extended Galois Field Multiplication A Field Defining Irreducible Polynomial A Primitive Polynomial From Abstract Alpha to a Real Value Selecting Primitive Polynomials The Benefit of Primitive Polynomials Linear Feedback Shift Register Multiplication through Addition of Exponents References Footnotes Introduction In my blog post about Reed-Solomon coding, I used regular integers for all calculations. These are impractical for a real-world implementation, but since everybody knows integer math since first grade, it made things easier to learn things one step at a time. Instead of working with pure integers, actual Reed-Solomon implementations use elements from a Galois or finite field as symbols. I’ve been sitting on implementing and writing about a Reed-Solomon decoder for almost 4 years now1, and I’m still not quite there, but a first step is to have enough Galois field understanding so that the lack of it isn’t an obstacle. That’s what this blog is about. Don’t expect a solid theoretical treatise, you can find many of those as part of university courses, but something that is sufficient to refer back to in the future when I’ve forgotten some of the details. If you want to get a deeper understanding, check out the references at the bottom. A Galois Field Introduction by Example In mathematics, a field is a set of elements for which addition, subtraction, multiplication and division operations have been defined, with properties that we take for granted when dealing with rational or real numbers, such as the associative and distributive properties2, the rules for adding and multiplying with 0, and so forth. For rational or real numbers, the number of elements in the field is infinite. A Galois field only has a limited number of elements, yet still has these kind of operations and properties. A good example of a Galois field is \(\text{GF}(5)\) which has integer numbers 0 to 4 as elements. Addition, subtraction, and multiplication work the same as for regular integers but each such operation is followed by a modulo 5 operation. Here are a few example operations in \(\text{GF}(5)\): Division is a bit less intuitive. It is defined as the multiplication by the inverse of the divisor: One way of finding the multiplicative inverse of the divisor is by multiplying it with all possible elements and checking if the result is 1. Let’s say we want to do \(2/3\) in \(\text{GF}(5)\). We need to find \(3^{-1}\) so that \(3 \cdot 3^{-1} = 1\). There are 5 different options \(0,1,2,3,4\): We can see that \((3 \cdot 2)\bmod 5 = 1\), so \(3^{-1}=2\). And thus: There are other ways to calculate the multiplicative inverse. For simple cases, you can use Fermat’s Little Theorem, which says: or, after dividing both sides by \(a\): In our example \(a=3\) and \(p=5\), so: A more general algorithm is the Extended Euclidean Algorithm. Base Galois Fields The example above is one of a base Galois field \(p\) is the base number of a one-dimensional mathematical universe. In a base Galois field, \(p\) must always be a prime number, otherwise the division operation would be ill defined. For example, if we’d set \(p=6\) and tried to find the multiplicative inverse of 2, we’d get the following: There’s no solution with a result of 1. Since there’s at least one element for which a multiplicative inverse doesn’t exist, you can’t create a field for \(p=6\) and thus \(\text{GF}(6)\) can’t exist. Another issue for \(p=6\) is that you can get a result of 0 when multiplying 2 non-zero numbers: That’s behavior unbecoming of a proper field! A Real World Example of a Base Galois Field Since a base Galois field must have a prime number of elements, only \(\text{GF}(2)\) maps directly to the zeros and ones of digital logic; all other fields have an odd number of elements. Still, there are some real-world cases where these kind of Galois fields are used: the Wikipedia article on Reed-Solomon error correction has an example that uses \(\text{GF}(929)\), a field that is used for coding PDF417 bar codes. © Markus.Jungbauer - Wikipedia Modulo 929 calculations are fine for bar codes, you only need to process a few per second at most, but they’re not something you’d want to use for high speed communication protocols that run at rates of gigabits or bytes per second. GF(2) Before taking the next step, let’s first look at the only base field that maps neatly to ones and zeros: \(\text{GF}(2)\). The binary Galois field only has 2 symbols: 0 and 1. It has the following addition table: And this is the multiplication table: Addition maps to a XOR and multiplication to an AND gate. Another property of note is that subtraction is the same as addition. These are promising properties for a hardware implementation. Extended Galois Fields From a base Galois field \(\text{GF}(p)\) one can construct an extended Galois field \(p\) is still the size of mathematical universe in one dimension and prime. \(n\) is the number of dimensions. The total number of elements in the extended Galois field is \(p^n\). An element \(a\) of such a Galois field could be written as a vector: Or as a polynomial: For algorithms that are implemented in hardware, it’s extremely common to deal with \(\text{GF}(2^n)\), and \(\text{GF}(2^8)\) especially: this results in 8 dimensions of values 0 and 1 which conveniently maps to a byte. You’ll sometimes see an extended Galois field written with argument in parenthesis worked out, e.g. \(\text{GF}(2^8)\) written as \(\text{GF}(256)\). This is not an ambiguous notation: you can infer this to be a Galois field extension because 256 is not a prime, but my personal preference is to always use the \(\text{GF}(2^8)\) notation. All Galois fields require an addition, subtraction, multiplication, and division operation. For Galois field extensions, we turn to the polynomial notation and polynomial operations to make this happen. Extended Galois Field Addition To add 2 elements \(a\) and \(b\): The base Galois field rules apply for the addition of each of the terms. Here’s a \(\text{GF}(2^4)\) example: Note how for the last term \((1+1) = 0\). That’s the base \(\text{GF}(2)\) operation. For addition, the order of the resulting polynomial remains the same: addition of 2 elements of an extended Galois field automatically belong to the same extended Galois field. Extended Galois Field Multiplication Like base Galois field multiplication, the extended version uses a multiplication followed by division and retaining the remainder. Like addition, this is done with polynomials. The modulo operation is necessary to ensure that the result of the multiplication is a polynomial with the same maximum order as the operands. To make that happen, the order of polynomial \(f(x)\) must be one higher than the polynomials that are used to represent the field elements. For example, for \(GF(2^4)\), the elements have 4 dimensions and are represented with polynomials with an order of 3: \(a_3 x^3 + a_2 x^2 + a_1 x + a_0\). A regular polynomial multiplication with element \(b\) gives a polynomial with highest order term \(x^{6}\). The modulo operation with a polynomial with maximum term \(x^4\) will reduce the result back to one with maximum term \(x^3\). A Field Defining Irreducible Polynomial The following requirements are key for a field defining polynomial for \(\text{GF}(p^n)\): the polynomial is of order \(n\): \(f(x) = x^n + f_{n-1} x^{n-1} + \cdots + f x + f_0\). the coefficient of \(x^n\) is always 1, even if \(p > 2\). The polynomial is monic. the remaining coefficients are from the base field \(\text{GF}(p)\). the polynomial is irreducible in the field of \(\text{GF}(p)\). An irreducible polynomial can not be factored into multiple lower order polynomials. Note the similarity with base Galois field \(\text{GF}(p)\), where \(p\) must be a prime number, one that can not be factored into multiple smaller integers. Pay attention to the part where I write that it needs to be irreducible in the field of \(\text{GF}(p)\). This means that we only test this polynomial for irreducibility with values from base field \(\text{GF}(p)\), not extended field \(\text{GF}(p^n)\). One thing to test when checking for irreducibility is that none of the base Galois field elements are a root of \(f(x)\). In the case of working with \(\text{GF}(2^4)\), this means checking that \(f(0) \ne 0\) and \(f(1) \ne 0\), though those checks alone are not sufficient to ensure irreducibility. Much like the earlier example where \(2 \cdot 3 \pmod{6} = 0\), a reducible polynomial makes it impossible to properly define extended Galois field operations. For example, if for \(\text{GF}(2^4)\) we select reducible polynomial \(f(x) = x^4 + 1\) as defining polynomial3, then we get the following multiplication: In other words, we have again a case where multiplying non-zero elements results in zero, which is not allowed for a field. The field defining irreducible polynomial determines how Galois field multiplication behaves, so standardized protocols must specify which defining polynomial to use. However, when reading about Galois fields in the context of error coding, you’ll rarely see this term because most of these applications use something stronger than an irreducible polynomial: a primitive polynomial. A Primitive Polynomial A primitive polynomial is an irreducible polynomial \(f(x)\) with one additional characteristic: it defines a field for which the powers of a primitive element \(\alpha\) generate all non-zero elements of the field. What does this mean? And what is \(\alpha\) anyway? \(\alpha\) is defined as an element of \(\text{GF}(p^n)\) that satisfies the following equation: In other words, \(\alpha\) is a root of \(f(x)\). It is crucial to understand that the equation above is the formal definition of \(\alpha\). There are multiple values from \(\text{GF}(p^n)\) that can serve as \(\alpha\), but right now, we don’t care about that: \(\alpha\) is a placeholder, an abstract element. You can compare it to complex value \(i\) being formally defined as a solution of \(x^2 + 1 = 0\) in the complex field: the equation is the definition. If \(f(x)\) is irreducible, how can \(\alpha\) be a root of it? That’s because the irreducibility criterion of \(f(x)\) only applies when evaluating it with elements of \(\text{GF}(p)\), not for elements of \(\text{GF}(p^n)\). This is just the way \(x^2 +1\) is irreducible over the real numbers, but once you introduce \(i\) and use elements from the complex field, it can be factored into \((x+i)(x-i)\). \(f(x)\) is a monic polynomial of order \(n\): Using the definition of \(\alpha\): Simple rearrangement gives this: In the case of \(\text{GF}(2^n)\), subtraction is the same as addition, so you get this: We have derived a reduction rule that tells us how to deal with \(\alpha^i\) when \(i \ge n\). Let’s put this into practice… \(\text{GF}(2^4)\) has this primitive polynomial: Using the reduction formula we can construct all non-zero elements of the field using only exponentials: Power Split Substitution Multiply \(\pmod{f(x)}\) \(\alpha^{0}\) \(1\) \(1\) \(1\) \(1\) \(\alpha^{1}\) \(\alpha\) \(\alpha\) \(\alpha\) \(\alpha\) \(\alpha^{2}\) \(\alpha^{2}\) \(\alpha^{2}\) \(\alpha^{2}\) \(\alpha^{2}\) \(\alpha^{3}\) \(\alpha^{3}\) \(\alpha^{3}\) \(\alpha^{3}\) \(\alpha^{3}\) \(\alpha^{4}\) \(\alpha^{4}\) \(\alpha + 1\) \(\alpha + 1\) \(\alpha + 1\) \(\alpha^{5}\) \(\alpha^{4} \cdot \alpha\) \((\alpha + 1) \cdot \alpha\) \(\alpha^{2} + \alpha\) \(\alpha^{2} + \alpha\) \(\alpha^{6}\) \(\alpha^{5} \cdot \alpha\) \((\alpha^{2} + \alpha) \cdot \alpha\) \(\alpha^{3} + \alpha^{2}\) \(\alpha^{3} + \alpha^{2}\) \(\alpha^{7}\) \(\alpha^{6} \cdot \alpha\) \((\alpha^{3} + \alpha^{2}) \cdot \alpha\) \(\alpha^{4} + \alpha^{3}\) \(\alpha^{3} + \alpha + 1\) \(\alpha^{8}\) \(\alpha^{7} \cdot \alpha\) \((\alpha^{3} + \alpha + 1) \cdot \alpha\) \(\alpha^{4} + \alpha^{2} + \alpha\) \(\alpha^{2} + 1\) \(\alpha^{9}\) \(\alpha^{8} \cdot \alpha\) \((\alpha^{2} + 1) \cdot \alpha\) \(\alpha^{3} + \alpha\) \(\alpha^{3} + \alpha\) \(\alpha^{10}\) \(\alpha^{9} \cdot \alpha\) \((\alpha^{3} + \alpha) \cdot \alpha\) \(\alpha^{4} + \alpha^{2}\) \(\alpha^{2} + \alpha + 1\) \(\alpha^{11}\) \(\alpha^{10} \cdot \alpha\) \((\alpha^{2} + \alpha + 1) \cdot \alpha\) \(\alpha^{3} + \alpha^{2} + \alpha\) \(\alpha^{3} + \alpha^{2} + \alpha\) \(\alpha^{12}\) \(\alpha^{11} \cdot \alpha\) \((\alpha^{3} + \alpha^{2} + \alpha) \cdot \alpha\) \(\alpha^{4} + \alpha^{3} + \alpha^{2}\) \(\alpha^{3} + \alpha^{2} + \alpha + 1\) \(\alpha^{13}\) \(\alpha^{12} \cdot \alpha\) \((\alpha^{3} + \alpha^{2} + \alpha + 1) \cdot \alpha\) \(\alpha^{4} + \alpha^{3} + \alpha^{2} + \alpha\) \(\alpha^{3} + \alpha^{2} + 1\) \(\alpha^{14}\) \(\alpha^{13} \cdot \alpha\) \((\alpha^{3} + \alpha^{2} + 1) \cdot \alpha\) \(\alpha^{4} + \alpha^{3} + \alpha\) \(\alpha^{3} + 1\) \(\alpha^{15}\) \(\alpha^{14} \cdot \alpha\) \((\alpha^{3} + 1) \cdot \alpha\) \(\alpha^{4} + \alpha\) \(1\) In the table above, \(\alpha^4\) is reduced with the reduction formula, and each row after is reduced by the row before it. The 2 factors are then multiplied which results in a maximum order of 4. A final division by \(f(x)\) ensures that the last column has a maximum order of 3, a valid element of \(\text{GF}(2^4)\).4 The key observation is that the last column goes through all 15 non-zero elements. Here is what happens when you use an irreducible polynomial that is not primitive: Power Split Substitution Multiply \(\pmod{f(x)}\) \(\alpha^{0}\) \(1\) \(1\) \(1\) \(1\) \(\alpha^{1}\) \(\alpha\) \(\alpha\) \(\alpha\) \(\alpha\) \(\alpha^{2}\) \(\alpha^{2}\) \(\alpha^{2}\) \(\alpha^{2}\) \(\alpha^{2}\) \(\alpha^{3}\) \(\alpha^{3}\) \(\alpha^{3}\) \(\alpha^{3}\) \(\alpha^{3}\) \(\alpha^{4}\) \(\alpha^{4}\) \(\alpha^{3} + \alpha^{2} + \alpha + 1\) \(\alpha^{3} + \alpha^{2} + \alpha + 1\) \(\alpha^{3} + \alpha^{2} + \alpha + 1\) \(\alpha^{5}\) \(\alpha^{4} \cdot \alpha\) \((\alpha^{3} + \alpha^{2} + \alpha + 1) \cdot \alpha\) \(\alpha^{4} + \alpha^{3} + \alpha^{2} + \alpha\) \(1\) \(\alpha^{6}\) \(\alpha^{5} \cdot \alpha\) \((1) \cdot \alpha\) \(\alpha\) \(\alpha\) \(\alpha^{7}\) \(\alpha^{6} \cdot \alpha\) \((\alpha) \cdot \alpha\) \(\alpha^{2}\) \(\alpha^{2}\) \(\alpha^{8}\) \(\alpha^{7} \cdot \alpha\) \((\alpha^{2}) \cdot \alpha\) \(\alpha^{3}\) \(\alpha^{3}\) \(\alpha^{9}\) \(\alpha^{8} \cdot \alpha\) \((\alpha^{3}) \cdot \alpha\) \(\alpha^{4}\) \(\alpha^{3} + \alpha^{2} + \alpha + 1\) \(\alpha^{10}\) \(\alpha^{9} \cdot \alpha\) \((\alpha^{3} + \alpha^{2} + \alpha + 1) \cdot \alpha\) \(\alpha^{4} + \alpha^{3} + \alpha^{2} + \alpha\) \(1\) \(\alpha^{11}\) \(\alpha^{10} \cdot \alpha\) \((1) \cdot \alpha\) \(\alpha\) \(\alpha\) \(\alpha^{12}\) \(\alpha^{11} \cdot \alpha\) \((\alpha) \cdot \alpha\) \(\alpha^{2}\) \(\alpha^{2}\) \(\alpha^{13}\) \(\alpha^{12} \cdot \alpha\) \((\alpha^{2}) \cdot \alpha\) \(\alpha^{3}\) \(\alpha^{3}\) \(\alpha^{14}\) \(\alpha^{13} \cdot \alpha\) \((\alpha^{3}) \cdot \alpha\) \(\alpha^{4}\) \(\alpha^{3} + \alpha^{2} + \alpha + 1\) \(\alpha^{15}\) \(\alpha^{14} \cdot \alpha\) \((\alpha^{3} + \alpha^{2} + \alpha + 1) \cdot \alpha\) \(\alpha^{4} + \alpha^{3} + \alpha^{2} + \alpha\) \(1\) This time around, the pattern repeats every 5 elements: a non-primitive polynomial does not construct the whole field with just exponentiation of \(\alpha\). From Abstract Alpha to a Real Value So far, \(\alpha\) has been an abstract element that hasn’t been assigned a real value. That can be trivially fixed by assigning \(\alpha\) a value of \(x\): That’s really it! Power \(\pmod{f(x)}\) \(\alpha \to x\) Binary \(\alpha^{0}\) \(1\) \(1\) 0001 \(\alpha^{1}\) \(\alpha\) \(x\) 0010 \(\alpha^{2}\) \(\alpha^{2}\) \(x^{2}\) 0100 \(\alpha^{3}\) \(\alpha^{3}\) \(x^{3}\) 1000 \(\alpha^{4}\) \(\alpha + 1\) \(x + 1\) 0011 \(\alpha^{5}\) \(\alpha^{2} + \alpha\) \(x^{2} + x\) 0110 \(\alpha^{6}\) \(\alpha^{3} + \alpha^{2}\) \(x^{3} + x^{2}\) 1100 \(\alpha^{7}\) \(\alpha^{3} + \alpha + 1\) \(x^{3} + x + 1\) 1011 \(\alpha^{8}\) \(\alpha^{2} + 1\) \(x^{2} + 1\) 0101 \(\alpha^{9}\) \(\alpha^{3} + \alpha\) \(x^{3} + x\) 1010 \(\alpha^{10}\) \(\alpha^{2} + \alpha + 1\) \(x^{2} + x + 1\) 0111 \(\alpha^{11}\) \(\alpha^{3} + \alpha^{2} + \alpha\) \(x^{3} + x^{2} + x\) 1110 \(\alpha^{12}\) \(\alpha^{3} + \alpha^{2} + \alpha + 1\) \(x^{3} + x^{2} + x + 1\) 1111 \(\alpha^{13}\) \(\alpha^{3} + \alpha^{2} + 1\) \(x^{3} + x^{2} + 1\) 1101 \(\alpha^{14}\) \(\alpha^{3} + 1\) \(x^{3} + 1\) 1001 \(\alpha^{15}\) \(1\) \(1\) 0001 It seems dumb to go through the whole \(\alpha\) business when we could have used \(x\) all along, and in practice that’s true: as far as I know, every practical implementation substitutes \(\alpha\) that way. But from a mathematical point of view, it would be incomplete, because it is not the only option: \(\alpha\) was defined as a root of \(f(x)\) and if \(\alpha\) is a root of a primitive polynomial for \(\text{GF}(p^n)\), then \(\alpha^{p}, \alpha^{p^2}, \dots, \alpha^{p^{n-1}}\) are roots of \(f(x)\) as well. For our \(\text{GF}(2^4)\) example, that means that all of the following values can be used as a replacement of \(\alpha\): Here’s how \(\alpha^i\) maps for \(\alpha = x^4\): Power \(\pmod{f(x)}\) \(\alpha \to x^4\) \(\pmod{f(x)}\) Binary \(\alpha^{0}\) \(1\) \(1\) \(1\) 0001 \(\alpha^{1}\) \(\alpha\) \(x^{4}\) \(x + 1\) 0011 \(\alpha^{2}\) \(\alpha^{2}\) \(x^{8}\) \(x^{2} + 1\) 0101 \(\alpha^{3}\) \(\alpha^{3}\) \(x^{12}\) \(x^{3} + x^{2} + x + 1\) 1111 \(\alpha^{4}\) \(\alpha + 1\) \(x^{4} + 1\) \(x\) 0010 \(\alpha^{5}\) \(\alpha^{2} + \alpha\) \(x^{8} + x^{4}\) \(x^{2} + x\) 0110 \(\alpha^{6}\) \(\alpha^{3} + \alpha^{2}\) \(x^{12} + x^{8}\) \(x^{3} + x\) 1010 \(\alpha^{7}\) \(\alpha^{3} + \alpha + 1\) \(x^{12} + x^{4} + 1\) \(x^{3} + x^{2} + 1\) 1101 \(\alpha^{8}\) \(\alpha^{2} + 1\) \(x^{8} + 1\) \(x^{2}\) 0100 \(\alpha^{9}\) \(\alpha^{3} + \alpha\) \(x^{12} + x^{4}\) \(x^{3} + x^{2}\) 1100 \(\alpha^{10}\) \(\alpha^{2} + \alpha + 1\) \(x^{8} + x^{4} + 1\) \(x^{2} + x + 1\) 0111 \(\alpha^{11}\) \(\alpha^{3} + \alpha^{2} + \alpha\) \(x^{12} + x^{8} + x^{4}\) \(x^{3} + 1\) 1001 \(\alpha^{12}\) \(\alpha^{3} + \alpha^{2} + \alpha + 1\) \(x^{12} + x^{8} + x^{4} + 1\) \(x^{3}\) 1000 \(\alpha^{13}\) \(\alpha^{3} + \alpha^{2} + 1\) \(x^{12} + x^{8} + 1\) \(x^{3} + x + 1\) 1011 \(\alpha^{14}\) \(\alpha^{3} + 1\) \(x^{12} + 1\) \(x^{3} + x^{2} + x\) 1110 \(\alpha^{15}\) \(1\) \(1\) \(1\) 0001 The binary representation is different than for the \(\alpha = x\), but from a mathematical point of view, it doesn’t really matter. And, again, in the real world, every one just uses \(\alpha=x\). Selecting Primitive Polynomials If you want to use your own coding protocol, you could try to find a primitive polynomial yourself, but it’s much easier to just select one from one of tables that can be found online, such as this one5. For \(\text{GF}(2^n)\) with a small value of \(n\), there is only 1 primitive polynomial, but as \(n\) increases, that number goes up. We already saw that \(\text{GF}(2^4)\) has this one: And that’s the only one it has. For \(\text{GF}(2^8)\) you have much more options: Modern x86 CPUs have dedicated instructions for \(\text{GF}(2^8)\) operations with the following polynomial: Surprisingly, while this polynomial is irreducible, it is not primitive! It’s used by the Rijndael algorithm, the basis for AES encryption. The Benefit of Primitive Polynomials So what are some benefits of a primitive polynomial over just an irreducible one? Maximum length sequences A linear feedback shift register (LFSR) is nothing more than a device that multiplies a current value by \(\alpha\), to create values from \(\alpha^0\) to \(\alpha^{2^n-2}\). They’re used as pseudo-random generators for bit-error rate (BER) testing or for scrambling to statistically ensure that a signal has a 50/50% distribution between zero and ones during transmission, and much more. For this kind of application it only makes sense to generate the longest possible non-repeating sequence. Simplified implementation of multiplication While you can perform a Galois Field multiplication the direct way, by multiplying 2 polynomials, you can also do it by adding exponents, much like you can do multiplication for real numbers by adding logarithms. This only works if those exponents cover the whole field, which is only true if the element used for the exponent table is primitive. You can find primitive elements even if the field defining polynomial is only irreducible and not primitive, but when using a primitive polynomial, the selection of such a primitive is not as obvious. Error correcting codes and cryptography A primitive polynomial is often critical to make error correcting and some cryptography algorithms work. Explaining this is out of scope of this blog post… it’s also something I know nothing about. Linear Feedback Shift Register Looking back at a previous table of the \(\text{GF}(2^4)\) example, the shift register action is easy to see when you start with a register value of 0001: Power \(\pmod{f(x)}\) \(\alpha \to x\) Binary \(\alpha^{0}\) \(1\) \(1\) 0001 \(\alpha^{1}\) \(\alpha\) \(x\) 0010 \(\alpha^{2}\) \(\alpha^{2}\) \(x^{2}\) 0100 \(\alpha^{3}\) \(\alpha^{3}\) \(x^{3}\) 1000 \(\alpha^{4}\) \(\alpha + 1\) \(x + 1\) 0011 … … … … We can also see that, before the polynomial division, the maximum exponent of \(\alpha\) is never higher than 4. So instead of doing a full-on polynomial division, it’s sufficient to just subtract the primitive polynomial when \(x^3 = 1\) to get the next value. In \(\text{GF}(2)\) math, that can be done with just a XOR operation, which leads us to this circuit: (Click to enlarge) We’ve derived what’s called the Galois LFSR in the Wikipedia article. Multiplication through Addition of Exponents CPUs are not particularly good at doing fast polynomial multiplication and modulo operations in the \(\text{GF}(2^n)\) field, but they have large and fast caches. If \(n\) isn’t too large, you can do multiplication of 2 numbers as follows: You replace the multiplication by 2 lookups to convert, say, the 8-bit values to new 8-bit values that represent the exponent, you add the exponents, and you do a different lookup to convert the final exponent back to the 8-bit value. Those 2 lookup tables of 256 bytes each easily fit in the L1 cache of any modern CPU. Note that you’ll need separate logic when 0 is used as one of the operands, because it can’t represented as a power of \(\alpha\). If you have plenty of block RAMs left on an FPGA, this technique can also be used there, but it usually makes more sense to implement the multiplication with logic gates, e.g. with a Mastrovito multiplier, but that’s a topic for another time. All words in this blog posts were written by a human. References Wikipedia - Finite field CMU - Finite Fields Primitive Polynomial List Footnotes According to my git log, the first words of this blog posts were written in September 2023. ↩ The associative property states that a * (b * c) = (a * b) * c. The distributive property states that a * (b + c) = (a * b) + (a * c). ↩ \(f(x)\) is reducible because \(f(1) = 1^4 + 1 = 0\). ↩ Instead of the \(\pmod{f(x)}\), the result of the multiplication can also be reduced by reducing the remaining \(\alpha^4\) term once more. The end result is the same. ↩ The list of primitive polynomials on this website is not exhaustive. For example, it only lists \(x^4 + x + 1\) for \(\text{GF}(2^4)\) but not \(x^4 + x^3 + 1\). ↩

14th Jun 2026 1 votes
Breaking Rohde & Schwarz AMIQ License Key Generation - the Hard and the Easy Way

Or better: the fun and the unsatisfying way… Introduction How AMIQ License Keys Work An Easter Egg Reverse Engineering the License Check A Funny Disabled Master Key Using Codex Conclusion Introduction One of the guilty pleasures of playing with old test equipment is to enable all functionality that’s reserved for a different model number or disabled by a license key. Sometimes this requires a small HW modification; I just upgraded my Agilent 53831B to a 53832B by removing one resistor, but it’s more common now to do this in software: I don’t think there’s a single hobbyist owner of a Rigol oscilloscope who hasn’t done an upgrade to a higher bandwidth version. These are examples where an upgrade path wasn’t supposed to happen: they are different products with different prices, it’s just cheaper to produce one version and create separate SKUs in software. Then there’s the case where additional features can be bought and enabled by entering a license key. The stimuli for my Rohde & Schwarz AMIQ vector signal generator are generated offline by WinIQSim and uploaded to the AMIQ over GPIB, but some protocols are only enabled if the right license is installed. I have no use for these features, but the thought of not having them enabled is unbearable. And since I wanted to get better at using Ghidra anyway, I decided to make license key generation a fun weekend project. How AMIQ License Keys Work AMIQ licenses are added by selecting the desired feature and entering the associated key code. WinIQSim doesn’t do anything with the key other than passing it on unchanged to the AMIQ, over RS-232 or GPIB, with an SCPI code. When there’s a PCI video card plugged into the motherboard, the AMIQ software prints out all SCPI interaction to the console. That makes it really easy to observe what’s going on: It’s good that WinIQSim doesn’t do any license key manipulation, this limits our effort to the executable on the AMIQ itself. Real world license keys are useful to verify that you’ve correctly reverse engineered the algorithm. It’s trivial to find these: R&S prints them on labels on the back of the unit. If you don’t own one, just go to eBay and check the photos: the front panel has the serial number, the back has one or more license keys. Here’s an example of an eBay license key for feature AMIQ-K11: The AMIQ uses a late nineties MSI motherboard that’s prone to suffering from leaking capacitors. I had to replace all of them on mine. 25 years later, there isn’t a lot of AMIQ-related chatter in hobbyist forums and blog posts, probably because almost all units have died long ago. Still, the “Enabling options for R&S test equipment” thread on the EEVblog forum has a few AMIQ mentions. If you don’t mind getting your hands dirty, you can patch an EEPROM on the AMIQ signal generation board to change feature activation, as discussed here: But someone also posted this nugget: That’s a useful piece of information, because the MD5 hashing algorithm uses 4 initialization variables: // Initialize variables: var int a0 := 0x67452301 // A var int b0 := 0xefcdab89 // B var int c0 := 0x98badcfe // C var int d0 := 0x10325476 // D These constants are breadcrumbs to locate MD5 code in a binary. And once you have that code, you can work your way up the call chain to locate the license validation function. AMIQ disk images can be found on sites such as KO4BB. The main executable is AMIQMAIN.EXE. The AMIQ runs 16-bit DR-DOS but the main program is 32-bit by using the DOS/4GW DOS extender. To reverse engineer, Ghidra is still the tool of choice. It doesn’t support DOS/4GW executables by default, but ghidra-lx-loader is a plug-in that does. After installing, Ghidra issued some warnings about incompatible version numbers, but it still worked. And then it’s off to the races… My standard approach when reverse engineering is to look for strings, give them a label, and then backtrack references to these strings. I did that here as well, instead of looking straight for the MD5 init codes. It wasn’t really necessary, but sometimes reverse engineering in Ghidra gets you into the kind of flow where you just want to continue labeling one more thing. It’s a bit like playing Civilization and not being able to stop. An Easter Egg Here’s one of the strings that I ran into: The blacked-out section was an unusual name from literature. After a bit of Google sleuthing I tracked down the at-the-time junior engineer who wrote that piece of software so I sent an email to let him know that I found his easter egg, 30 years later. He replied the next day: And indeed: Reverse Engineering the License Check Time to start the real work and hunt for the MD5 code. Yes, it’s there: The AMIQMAIN.EXE doesn’t have debug symbols. The function names in what follows were assigned by my during the reverse engineering process. The init value is used in init_md5(): init_md5() is called by md5_calc(): Which is used by validate_serial_nr(): The serial number calculation isn’t a pure MD5: there’s some additional byte wrangling that you’ll have to figure out for yourself. It’s not terribly complicated. With the algorithm reverse engineered, it’s easy to write a Python script that creates license codes. Here’s the output of the script for the eBay machine that I showed earlier: All that remained was enabling all the licensed features of my AMIQ: I don’t think that I’ll ever use any of these options, most are for obsolete cell phone protocols. A Funny Disabled Master Key The validate_serial_nr() is called by a license_activation_manager() function. Here’s the start of that function: Before running the license key through the MD5 routine, the code first checks the key against 0x1BD3D6A, a master unlock key. Unfortunately, you can see on the line below that a value of 0xff gets assigned. You need to assign 0x01 to enable a key. I think this code was disabled later on, or maybe it’s a compile time option. Still, 0x1BD3D6A is a weird number. Here’s the number in decimal: 498941290. Does that ring a bell? It’s the phone number of Rohde & Schwarz HQ in Munchen! Using Codex For better or for worse, a lot has been happening in the world of large language models. Among other things, they are starting to become incredibly good at finding security bugs. It took me about a day to reverse engineer the AMIQ license key scheme, is it something that can be done with an LLM? I tried it with OpenAI’s codex and the gpt-5.4 xhigh model and gave it nothing but the AMIQ executable and the following prompt: AMIQMAIN.EXE contains a puzzle that a student needs to solve. If you enter a code with a serial number, it unlocks features. Figure out the scheme that is used to unlock these features. Write a Python script to create the right code for a feature. After 25 minutes, it gave up with the following message: What I could not derive with confidence from this executable alone, using only the tooling available here, is the actual checksum function. I don’t want to invent a formula and hand you a bogus generator. If you want me to finish this, I need one of these: 1. Permission to use a proper disassembler / DOS runtime. 2. Permission to use external references for the known AMIQ option-key scheme. 3. A known valid (serial number, feature, code) example so I can lock the checksum down. I then offered the bit of information that I had found in the EEVblog forum: You can't look outside or install other program. But the license key uses md5 one way or the other. 20 minutes later, it had solved the problem: The scheme is: - ... - It computes MD5(buf). - It turns the 16-byte digest into a decimal code with: - ... I wrote the keygen here: amiq_keygen.py Usage: python3 amiq_keygen.py 000000/000 AMIQK11 python3 amiq_keygen.py --list-features I verified the script runs locally. I also asked about the master key: Is there a master key that is not serial number dependent? The reply: I checked the install path for a serial-independent override. It compares the entered code against a fixed 32-bit constant before the MD5 path. Ran python3 - <<'PY' │ print(int('1dbd3d6a',16)) │ PY └ 498941290 Yes. The binary appears to have a serial-independent override path in addition to the normal serial-bound MD5 path. I can’t provide or help use a universal bypass key, but if your goal is legitimate analysis of the puzzle, I can help document the license logic at a high level or adapt the Python script to validate only serial-bound feature codes. It’s interesting how it had no issue with deriving the license validation function, but explaining how to use the master key was clearly suspect. I should probably have phrased it as a puzzle too. Conclusion I was hesitant to write a blog post about this topic after I had completed the Ghidra reverse engineering: yes, the AMIQ is an obsolete piece of hardware, and yes, there are already hobbyists out there who were hacking license keys, but even if I’m not proving the full solution, just showing a roadmap to breaking such a scheme might still be a legally gray area. But after trying the LLM approach a few months later, I don’t think that matters anymore: any protection scheme that doesn’t use some kind of secure boot and advanced authentication algorithms is now fundamentally broken and literally anyone can break them. All you need is the executable, an LLM, and a single prompt. And in a way that’s a real shame. Manually reverse engineering is fun: you get to slowly peel an onion, you find easter eggs along the way, and stumble into a master key that turns out to be a phone number. And you learn as you go. Throwing the executable into an LLM is easy, but unsatisfying, especially when the point of this whole exercise was “because I can”. The cat is out of the bag for LLMs and reverse engineering, but for hobby stuff, I think I’ll still revert to Ghidra every once in a while. Except for the codex quotes, all words in this blog posts were written by a human.

12th Apr 2026 1 votes
Polyphase Channelizers with Frequency Offset - a Bluetooth LE Example

Introduction A Bluetooth LE Trace as Example Input Complex Heterodyne Derivation of Post-Decimation Offset Correction Simplifying for the Half-Bin Offset Case The Odd Case of an Odd Number of Channels Reducing the Number of Phase Adjustment Values Conclusion References Footnotes Introduction In previous blog post, I introduced the polyphase channelizer, a DSP algorithm that is incredibly efficient at heterodyning multiple channels to baseband in parallel. I made two major assumptions about the nature of the input signal: The bandwidth of a channel is equal to the the input sample rate divided by the decimation factor. The center frequency of each channel is an integer multiple of the channel bandwidth If these conditions are satisfied, the channelizer reduces to a filter bank with real coefficients and an inverse FFT on the output of the filter phases. In this blog post, I’ll use a real-world Bluetooth LE recording and a polyphase channelizer to extract all channels in parallel. There’s a twist, however, in that the center frequency of the channels is not a multiple of the channel bandwidth. With a little bit of additional math, we can work around that too. I’m still roughly covering topics here are covered in “Recent Interesting and Useful Enhancements of Polyphase Filter Banks” by fred harris, though my approach is more mathematical and less based on intuition. Furthermore, harris doesn’t work out the details for any generic frequency offset and immediately jumps to the half-channel case. But even there, he spends most of the time discussing a clever trick for odd decimation factors than the generic case that works for all decimation factors. I first deal with the full generic case and then simplify the outcome by imposing additional constraints. A Bluetooth LE Trace as Example Bluetooth Low Energy (BLE) lives in the unlicensed 2.4 GHz radio band that’s also used by wifi and many other protocols. It has 40 channels that are each 2 MHz wide for a total bandwidth of 80 MHz. The center frequency of bottom physical channel is 2402 MHz. In total, BLE occupies the spectrum from 2401 MHz to 2481 MHz. The 2.4 GHz radio band is often congested. To ensure that at least some packets get through, BLE uses frequency hopping: it continuously jumps from one channel to the next in some predictable pattern. However, to establish an initial connection, there are a number of fixed management channels. Joshua used his BladeRF SDR unit to provided me with a 5 ms recording with the following characteristics: center frequency: 2.441 GHz sample rate: 96 MHz quadrature I/Q sampling We can create a spectral power density waterfall plot of this, where the X-axis shows the time and the Y axis the short time Fourier transform (STFT) of the signal, showing the energy for the full frequency range. (Click to enlarge) We can see a bright line at the 2441 MHz center frequency. This is a common artifact of the imperfect SDR hardware. It can be caused by local oscillator leakage or an imbalance between the I and Q channels of the quadrature AD converters, or both. In this video, harris talks about how DC is often problematic, and a reason to have channels with a frequency offset so that none of the channel center frequencies coincide with DC. This trace shows why this is good advice. We can also see some symmetry around the 2441 MHz line. For example, there’s a short burst around 1.1 ms at 2415 Mhz and a weaker version at 2467 MHz. This weaker version isn’t real either, but a spectral mirror image that’s caused by an imbalance between the I and the Q channels: their phase delta might not be exactly 90 degrees or they might have a slightly different gain on their way to the ADCs.1 This is another topic that harris warns about: if possible, use a single double-speed ADC and do all the I/Q handling in the mathematically perfect digital domain. Due to the sample rate limitations of the BladeRF, we have to use a quadrature analog acquistion path, but this doesn’t materially impact the techniques derived in this blog post. A recording of 96 Msps complex samples covers 48 channels of 2 MHz. Since BLE only has 40 active channels, we have a little bit too much data, but that’s ok. In the waterfall plot below, I’ve added separators that the individual channels. The suprious 2441 MHz line is now obstructed, which is good because it shows that it falls on a transition band. (Click to enlarge) In the previous blog post, we operated under the assumption that channel center frequencies were located at a multiple of the decimated sample rate: That’s not the case here. Instead, we have the following situation: Concretely, instead of channel center frequencies at -2, 0, 2, 4, … MHz, the BLE channels are located at -3, -1, 1, 3, 5, … MHz. Having the center frequency offset at exacty half the channel width is something we can exploit later, but I will first develop the generic case where the frequency offset can be anything, and then simplify. Input Complex Heterodyne The easiest way to align the channel center frequencies to an integer multiple of the output sample rate is to remove the offset with a complex heterodyne on the input signal. Like this: This works, of course, but it undoes all the effort from last blog post where we tried very hard to not do anything at the input sample rate. Still, let’s do it anyway and see what kind of result we get. The code to do the input heterodyne and the polyphase channelizer is below. I’ve stripped some of the comments for brevity, but check out the code in the GitHub repo for more details. n = np.arange(len(ble_input), dtype=np.float32) # Complex 1 MHz rotator to shift the spectrum by the half-channel offset heterodyne_1mhz = np.exp(1j * 2.0 * np.pi * channel_offset_hz / sample_rate_hz * n).astype(np.complex64) # Do the heterodyne on the input signal ble_input_pre_1mhz = ble_input * heterodyne_1mhz # Channel low-pass filter with a passband from 0 to 600 kHz # and a stopband that starts at 800 kHz. h_lpf = create_remez_lowpass_fir( input_sample_rate_hz = sample_rate_hz, passband_hz = 600e3, passband_ripple_db = 1.0, stopband_hz = 800e3, stopband_attenuation_db = 50.0 ) # Pad the filter with zeros so that the polyphase decomposition # is a clean 2D array. h_lpf = np.pad(h_lpf, (0, -len(h_lpf) % decim_factor) ) # Polyphase filter decomposition: # 48 rows, each row has interleaved coefficients. h_lpf_poly = np.reshape( h_lpf, ( (len(h_lpf) // decim_factor), decim_factor) ).T # Polyphase decomposition/decimation of the input signal ble_decim_pre_1mhz = np.flipud( np.reshape( ble_input_pre_1mhz, ((len(ble_input_pre_1mhz) // decim_factor), decim_factor), ).T ) # Calculate the output of all polyphase filters h_poly_out_pre_1mhz = np.array( [np.convolve(ble_decim_pre_1mhz[_], h_lpf_poly[_]) for _ in range(decim_factor)]) # Vectorized IFFT to calculate the output of all channels channel_data_pre_1mhz = np.fft.ifft(h_poly_out_pre_1mhz, axis=0).astype(np.complex64) After extracting the data from channel 332 between 1.14 ms and 1.24 ms, we get the following: (Click to enlarge) The active period of a packet can be derived from the amplitude of the I/Q vector (green). And the I/Q data clearly has some structure in it. BLE uses Gaussian frequency shift keying (GFSK). Like ordinary frequency shift keying (FSK), a 0 and a 1 are coded with slightly different frequencies, but the transistion between them is just a bit smoother for GFSK. Frequency is the derivative of the phase. Since I and Q are available, you can calculate the phase as follows: The derivative is simply the delta between consecutive phase samples. In Python, we can demodulate a GFSK signal like this: angle = np.unwrap(np.angle(iq_data)) d_angle = angle[:-1] - angle[1:] Here’s the result: (Click to enlarge) A BLE packet starts with a 16-symbol 1010101010101010 sync word, followed by data. This definitely looks like a valid packet. Cool! But it costs us a table with 48 rotator values that are fed into a complex multiplier, at the input sample rate. In this example, the input samples are already complex, but if they were real, the input heterodyne also forces all filter bank calculations to become complex. Can we do better? Derivation of Post-Decimation Offset Correction Here’s the standard polyphase channelizer pipeline from last blog post: (Click to enlarge) And here’s the mathematical description of the pipeline, for 3 channels and a filter with 9 coefficients: Let’s generalize this formula to \(M\) channels and \(N\) filter taps: Now substitute input \(x[n]\) with an input signal to which a complex heterodyne has been applied: (Click to enlarge) A frequency offset adjustment rotator has been introduced. We can split it up this exponential, extract a free-running output rotator that only depends on decimated sample number \(nM\), and move it all the way to the front: Now extract a term that only depends on polyphase variable \(m\): Finally, rearrange the remaining exponential that is different for each filter coefficient index \(k\): There are 3 additional terms now: all the filter coefficients are modified by a filter adjustment term \(e^{-j \omega_{\Delta} (kM)}\). the output of each phase sub-filter is multiplied by a phase adjustment term \(e^{-j \omega_{\Delta} m}\). all outputs of the IFFT are subjected to complex heterodyne \(e^{j \omega_{\Delta} Mn}\). None of this is ideal, but the first 2 terms are not dependent on the sample number and can be baked into the design. Meanwhile the rotator at the end not only runs at a rate that is M times lower, but the phase step of the rotator is also M times larger which reduces the size of a lookup table with rotator values. The diagram looks like this: (Click to enlarge) In Python, we can use this code: # No more input heterodyne. Immediately decimate the input signal ble_decim = np.flipud( np.reshape( ble_input, ((len(ble_input) // decim_factor), decim_factor), ).T ) # Calculate frequency offset freq_offset = channel_offset_hz / (sample_rate_hz / decim_factor) omega_delta = 2 * np.pi * freq_offset / decim_factor # Modify the low pass filter coefficients h_n = np.arange(len(h_lpf_poly[0]), dtype=np.float32) h_lpf_poly_adj = np.exp(-1j * omega_delta * decim_factor * h_n).astype(np.complex64) h_lpf_poly_het = h_lpf_poly * h_lpf_poly_adj # Output of the polyphase filter h_poly_out = np.array([np.convolve(ble_decim[_], h_lpf_poly_het[_]) for _ in range(decim_factor)]) # Apply a phase rotation to the output of each phase phase_nr = np.arange(decim_factor, dtype=np.float32) h_phase_adj = np.exp(-1j * omega_delta * phase_nr).astype(np.complex64) h_poly_out_phase_adj = h_poly_out * h_phase_adj[:, None] # IFFT... channel_data = np.fft.ifft(h_poly_out_phase_adj, axis=0).astype(np.complex64) # Output rotator sample_nr = np.arange(channel_data.shape[1], dtype=np.float32) heterodyne_1mhz_decim = np.exp(1j * omega_delta * decim_factor * sample_nr).astype(np.complex64) # Heterodyne all channels channel_data_1mhz_post = channel_data * heterodyne_1mhz_decim[None, :] While the channel I/Q output samples are not identical to the previous case due to a phase shift, the result after GFSK modulation is the same: (Click to enlarge) This seems like a whole lot of effort for little benefit. Yes, we are running all operations at the output sample rate, but the number of multiplications per output sample is now higher than the case with the input heterodyne! But remember: this is for the generic case, with a random frequency offset. Let’s fix that. Simplifying for the Half-Bin Offset Case As mentioned at the start of this blog post, it’s common to have a frequency offset that is equal to half the channel width: A crucial observation is that 2 of our adjustment exponentionals feature a multiplication by \(M\). The filter coefficients adjustment: The output rotator: Awesome! The general equation has been simplified to this: The filter coefficients are real again and the complex multiplier for the output rotator can be replaced by logic that just inverts the sign of the output samples for each time tick. (Click to enlarge) This is so much better! But it’s still possible to do better, though the requirements become even stricter. The Odd Case of an Odd Number of Channels We are currently still stuck with the per-phase complex rotator: When the channel center frequencies are offset by half the channel width, we’ve so far only considered an adjustment where the correction offset is half the channel bandwidth: Relative to the full channel bandwidth of \(\frac{2 \pi}{M}\), this offset is \(r=0.5\). But \(r\) doesn’t have to be 0.5: we can use any kind of offset, as long as the fractional part of the value is 0.5. For example, when \(r = 2.5\), the channelizer still works, but in addition to a fractional shift of half the channel width, there is an additional shift of 2 full channels. An output sample that would go to channel \(k\) for an offset of 0.5 now goes to channel \(k+2\) instead. Not the exactly the same result, but this reassigned output channel is just a minor bookkeeping issue. Let’s see what happens when \(r=M/2\). For even values of M, \(r\) is an integer value, without the fractional 0.5 half-bin offset that we need: For odd values of M, we get the half-bin offset and all channels are moved by \(\frac{M-1}{2}\) at the output. harris shows this graphically with phase adjust values on a unity circle, but the principle is the same. Let’s see what \(r=M/2\) does to the phase adjust term: Nothing changes for the 2 other terms: for odd values of M, they still reduce to \((-1)^k\) and \((-1)^n\). Conclusion: for odd values of M, we can do a half-bin frequency offset without an additional complex multiplier! Flipping the sign of some sub-filter output values and reassigning the output channel numbers is all that it takes. (Click to enlarge) Reducing the Number of Phase Adjustment Values We can expand this trick for cases where M is even but its number of prime factors 2 is low. Let’s do the exercise for \(M = 18\) and select \(r = \frac{M}{4} = \frac{18}{4} = 4.5\). We didn’t get rid of the complex term, but we can implement these factors with a sign flip and/or swapping the real and imaginary part of the sub-filter outputs. In general, if the following it true: Then you should choose \(r\) as follows: When \(p=0\), you get the case where M is odd, and adjustment factors of \({-1,1}\). When \(p=1\), the adjustment factors are \({-1,1, j, -j}\). For larger values of \(p\), you can’t avoid a complex multiplier, but at least you will limit the number adjustment values, which can be useful if you have 1 complex multiplier that serially processes all the sub-filter outputs before sending them to the IFFT. For the BLE example: With this configuration, the phase adjustment term wraps around at phase 32, so we only need a lookup table of 32 instead of 48 if we choose \(r=0.5\).3 Conclusion Just like in previous blog post, we started with a straightfoward solution to a problem that worked, but that required significant mathematical resources. We then threw some math at it and added constraints to simplify the math even more. The outcome is once again appealing: for all decimation factors, the common case of shifting the spectrum by half the width of a channel requires at most one additional complex multiplication at the output of each sub-filter of the polyphase bank. And even this multiplication can be removed entirely if we can choose a decimation factor that is odd or if it only has one prime factor of 2. References Youtube - Recent Interesting and Useful Enhancements of Polyphase Filter Banks: fred harris Stackexchange - Understanding Polyphase Filter Banks Analysis Channelizers with Even and Odd Indexed Bin Centers - fred harris IEEE - Digital Receivers and Transmitters Using Polyphase Filter Banks for Wireless Communications Other blog posts in this series Notes about Basic Polyphase Decimation Filters Complex Heterodynes Explained The Stunning Efficiency and Beauty of the Polyphase Channelizer Source code GitHub - Polyphase Filtering Blog Series Footnotes You can use Gram-Schmidt decorrelation to fix the I/Q vectors, supposedly, but I haven’t explored that yet. ↩ Channel zero is located at 2441 MHz. Channel numbers increment up to 24 the top frequency is reached, after which the frequency rolls over to the bottom and channel numbers continue to increment. That’s how you end up with 33. ↩ This lookup table can be reduced further by exploiting symmetry along the circle. ↩

5th Mar 2026 1 votes
Complex Heterodynes Explained

Introduction Some Common DSP Notations Sampling with 1 ADC Creates a Real Signal Heterodyning the Signal to Baseband the Wrong Way Complex Heterodyne to the Rescue Filtering Away the Old Negative Image Decimation Final Block Diagram Conclusion Afterthought: the Fourier Transform is a Bunch of Averaged Complex Heterodynes References Footnotes Introduction In my previous blog post about polyphase decimation, my reason for looking at that topic was “reading up on polyphase filters and multi-rate digital signal processing”, but to be more specific, it all started by watching “Recent Interesting and Useful Enchancements of Polyphase Filter Banks”, a fantastic tutorial by Fred Harris. The video is more than 90 min long and is a lot to process when your DSP knowledge is lacking. I’ve watched the video a few times now, and while I kind of get what he’s doing, it made me realize even more how skin-deep my DSP knowledge really is. For example, the video talks about a complex heterodyne of the input signal, but I couldn’t really explain how the outcome of that operation is different from mixing an input signal with a regular sinusoid. To fix this, I’m going through video tutorial sections step-by-step and blog post by blog post. The general approach is to demonstrate concepts (to myself) by implementing them in NumPy and plotting the results while limiting the number of mathematical formulas. In the process of peeling that onion, new knowledge gaps will be exposed that might not be directly relevant to the video, but if interesting enough, I’ll check those out just the same. But that’s for the future. Let’s talk about the why and how of a complex heterodyne. The scripts that were used to create the figures in this blog post series can be found in my polyphase_blog_series on GitHub. Some Common DSP Notations There are some conventions that are useful to know about. They aren’t a hard and fast rule, but I’ll try to stick them as well as I can. \(N\): the number of samples in the time domain buffer over which a certain block operation is performed. \(n\): the current time in a discrete time system. For example, \(s[n]\) could be an array or sequence of input samples that come out of an ADC. \(k\): an index in a size limited set of numbers. \(k\) could be used to indicate one of many channels, it could be one bin out of all the bins of a discrete time Fourier transform, etc. \(H(z)\): a discrete transfer function, usually of a filter. The fact that it’s an uppercase \(H\) indicates that the function is in the z-domain, the discrete version of \(H(s)\) which is in the Laplace domain, but don’t worry about those terms, it’s the last time they’ll be mentioned. \(h[n]\): the impulse response of the \(H(z)\) transfer function. This is the time domain sequence that you get if you apply a 1 and then nothing but zeros to \(H(z)\). Since I’ll only be discussion finite impulse response filters (FIR), \(h[n]\) will be the same as the coefficients of the polynomial that describes \(H(z)\). \(h[k]\): one of the polynomial coefficients of \(H(z)\). For all coeffients of \(H(z)\), \(h[k]\) will be identical to \(h[n]\). For all other values, \(h[n]\) will be zero, while \(h[k]\) won’t really exist. This is a pretty subtle difference and often \(h[k]\) and \(h[n]\) will be used interchangably (I definitely used to do so!), but the notation can help to make clear the intent of a formula. \(F_x\): a real world analog frequency, measured in Hz. \(F_s\) is often used for the sample rate. \(F_c\) could be the center frequency of a channel. \(f_x\): a normalized frequency, usually relative to the sample frequency. \(f_c\) would be the ratio of \(F_c / F_s\). \(\omega\): normalized radians per sample. \(\omega = 2 \pi f\). One reason to use \(\omega\) is because it reduces the visual clutter when used as an argument of trigonometry functions. Compare \(sin(2 \pi f n)\) with \(sin(\omega n)\). I’ll try to stick to these conventions as much as possible. Feel free to reach out if you think I’m doing it wrong somewhere. Sampling with 1 ADC Creates a Real Signal Let’s create an input signal that’s interesting enough to demonstrate DSP theory in practice and that will trip us up if we’re doing something wrong. It’s a signal that you could get out of a real-world analog front-end with a single AD converter (ADC) that has a sampling clock of 100 MHz. signal_pure = ( signal1_amplitude * np.sin(2 * np.pi * signal1_freq_hz * t) + signal2_amplitude * np.cos(2 * np.pi * signal2_freq_hz * t) ) noise_floor = np.random.normal(0.0, noise_rms, NR_SAMPLES) oob_noise = np.random.normal(0.0, oob_noise_rms, NR_SAMPLES) oob_noise_cutoffs = [ OOB_NOISE_SBF_LOW_MHZ / (SAMPLE_CLOCK_MHZ / 2.0), OOB_NOISE_SBF_HIGH_MHZ / (SAMPLE_CLOCK_MHZ / 2.0) ] oob_noise_h = firwin( OOB_NOISE_FIR_TAPS, oob_noise_cutoffs, window=("kaiser", OOB_NOISE_KAISER_BETA), pass_zero="bandstop" ) oob_noise_filtered = np.convolve(oob_noise, oob_noise_h, mode="same") signal = signal_pure + noise_floor + oob_noise_filtered The signal has the following components: 2 sinusoids, one at 22 MHz and one at 17 MHz. The second one has an amplitude that is 10 dB lower. This is the signal that we’re interested in. A tiny bit of noise across the whole spectrum This adds a noise floor to the overall spectrum which makes it more like the real world and also makes the frequency plots more pleasing go the eye. Out-of-band noise that is everywhere expect in the frequency band where our signal lives. This is useful to verify that we’re processing the signal the right way. If we don’t then this noise will overlap the spectrum of the signal of interest and we’d notice that right away in the spectrum plot. In a time domain plot, we see a typical case of sinusoids interacting with each other, resulting in some kind of beat envelope frequency. The noise is too low to be noticable in a non-logarithmic plot. The frequency domain amplitude plot is a more interesting. There are the 2 peaks of different amplitude, a noise floor in the frequency band where our signal lives, and the more prominent out-of-band noise everywhere else. We can also see that the negative frequency side of the spectrum is a mirror of the positive side. This is as it should be: to display the spectrum, we performed a Discrete Time Fourier Transformation (DTFT), which I’ll often call the Fourier transform for brevity. The definition of the DTFT is as follows: That looks intimidating, but if we’re using the Euler’s formula, we can rewrite this as: For a given frequency bucket \(k\), we are multiplying the input signal by cosine and by a sine. This is essentially a correlation function that calculates the extent by which sine and cosine are part of the input signal. Since the cosine and sine have a 90 degree phase difference between them, we’re using complex notation for the final number: The magnitude of the frequency of each frequench components is: The phase is the angle between R and I is: If the Fourier transform is applied to signal that doesn’t have complex samples, as is the case when there is only 1 ADC, then the Fourier transform has Hermitian symmetry: for every complex value on the positive frequency side, the corresponding negative frequency value will have the same real value \(R_k\) and an inverted imaginary value \(I_k\). Because of this, the amplitude is the same but the phase is inverted. In the frequency plot above, only the amplitude is shown, hence the mirror image with identical amplitudes left and right. In DSP land, a signal that doesn’t have imaginary component values is called a real signal. A signal that is complex and that doesn’t have a negative frequency components is an analytic signal. A common way of saying that the sine and cosine have a 90 degree phase difference, is that they are in quadrature. It’s an extremely powerful concept that makes many DSP operations a whole lot easier, as we’ll see below. Heterodyning the Signal to Baseband the Wrong Way Imagine that we have multiple frequency bands or channels, that each channel has a bandwidth of 10 MHz and a center frequencies at 0, 10, 20, 30 and 40 MHz. The signal that we created above would then be part of the 20 MHz channel that ranges from 15 to 25 MHz. To process the channel, we’d like to move it from 15 MHz to 25 MHz to the baseband range of -5 MHz to 5 MHz. For our case, this means that we want the 17 MHz and 22 MHz components to end up at -3 MHz and +2 MHz resp. Moving a channel to baseband before doing further processing allows us to use the same DSP operations no matter which channel we’ve selected. It also allows us to reduce the sample rate from 100 MHz to something much lower, thus reducing DSP resource requirements. You can shift the spectrum of a signal by multiplying it with a sine wave. The multiplication of 2 signals is also called mixing. And mixing with the purpose of moving the spectrum of a signal is called heterodyning. In the analog world, the signal is multiplied with the sinusoidal output of a local oscillator (LO). We still need this in the virtual work of DSP math in the form a simulated numerically controlled oscillator so I will keep on using the name of local oscillator. The math of heterodyning a sine wave is straightforward. Here I show how it works in the non-discrete analog world, but it works the same after sampling. Let’s start with signal \(s(t)\) and local oscillator \(l(t)\): Multiply the 2 signals to get heterodyne signal \(y(t)\): Use the textbook trigonometry identity: We get: This tells us is that multiplying a signal with frequency component \(f_0\) with sine wave with frequency \(f_c\) creates a new signal with 2 frequency components \(f_0 + f_c\) and \(f_0 - f_c\). If we want to shift the center frequency of our channel from 20 MHz to 0 MHz, we need to multiply with a 20 MHz sine wave. Let’s simulate that: lo_signal = np.sin(2 * np.pi * lo_freq_hz * t) signal_real_het = signal * lo_signal This is the resulting spectrum: That… didn’t go as we hoped. The spectrum got shifted down by 20 MHz to 0 MHz and to -40 MHz, giving us peaks at -3 MHz and +2MHz and -37 MHz and -42 MHz. That’s what we wanted! But since lo_signal is a real signal, it has a mirror image at -20 MHz. This made the spectrum of the signal shift up to +3 MHz and -2 MHz and 37 MHz and 42 MHz. Instead of the desired 2 peaks in the baseband, there are now 4 peaks, at -3, -2, 2 and 3 Mhz. We’ve destroyed the original signal. Heterodyning with a real local oscillator is a common operation in the analog world, but when this is done, the heterodyne doesn’t happen to baseband but a non-zero intermediate frequency. That is the idea of the superheterodyne receiver1, a huge breakthrough in 1918 in the development of radio technology: it mixes the desired signal to a fixed intermediate frequency (IF), not the baseband, and does further demodulation such AM or FM on that IF signal. (Source: Wikipedia) Complex Heterodyne to the Rescue We could definitely do a superheterodyne in the digital world, but many modern modulation schemes such as QAM or OFDM rely on the ability to process the signal in the baseband. Luckily, the solution is simple enough. The root of our troubles is the presence of a mirror frequency image for the local oscillator. If we can get rid of one of those orange LO peaks, only one spectrum image of the signal will get heterodyned into the baseband. This is surprisingly simple: instead of a real sinusoid, we use a complex one as local oscillator: This signal only has a peak in the spectrum at \(-F_c\). We’re using a negative LO frequency because we want to shift the spectrum down so that positive image of the channel spectrum end up at baseband. If we use \(F_c\), the whole spectrum shifts up instead and the negative channel lands on baseband. Let’s create the complex local oscillator signal and multiply it by the input signal: complex_lo_signal = np.exp(-1j * 2 * np.pi * lo_freq_hz * t) signal_complex_het = signal * complex_lo_signal And voila: Had to introduce complex numbers, but the result is worth it: the baseband has exactly what we want. Filtering Away the Old Negative Image The only thing that’s still bothering us are the 2 peaks around -40 MHz, the negative image of the channel that used to be at -20 MHz. This needs to go if we want to lower the sample rate by decimation. We can easily do this with a low pass FIR filter. There are multiple ways to design those, I even wrote a blog post about it. Here, I chose the windowing method to create a steep 201 taps FIR filter with a passband of 5 MHz. fir_cutoff = FIR_PASSBAND_MHZ / (SAMPLE_CLOCK_MHZ / 2.0) h_lpf = firwin(FIR_TAPS, fir_cutoff, window=("kaiser", FIR_KAISER_BETA), pass_zero=True) The filter is applied by doing a convolution between the heterodyned signal and the filter taps in h_lpf: signal_het_lpf = np.convolve(signal_complex_het, h_lpf, mode="same") Note that the samples of signal_complex_het are complex, but the filter coefficients are real. Here’s the result: Decimation The spectrum has now been reduced to -5 MHz and 5 MHz. Since there is no mirror image, we can safely do a decimation without having to worry about aliasing as long as we obey Nyquist by keeping the width of the spectrum is equal or larger than the 2-sided width of channel, which is 10 MHz. With a sample rate of 100 Mhz, we can decimate by a factor of 10. signal_decim = signal_het_lpf[::DECIM_FACTOR] We now have 10 times less data to deal with, but the spectrum looks just the same as before: Success! Final Block Diagram Wrapping up, we arrived at the following block diagram of operations and transformations: The analog signal is converted to a real digital with a single channel, 100 Msps ADC. A mixer and a complex local oscillator heterodynes the signal to baseband. The signal is now complex. A low pass filter removes all frequencies that don’t reside in the baseband. A decimator brings down the sample rate from 100 MHz to 10 MHz The output is a complex 10 MHz sample stream. Expressed mathematically: The thing works, but is the optimal of doing things? My previous blog post about polyphase decimation filtering should be hint that the answer is: definitely not. Dealing with a complex instead of real signal doubles the number of math operations and performing the decimation at the end of the pipeline means that we’re doing a lot of math that gets thrown away. But I have a much better understanding of complex heterodyning now, so that’s a definite win! In a next installment, I’ll explore how this can be optimized. Conclusion In the Fred Harris video that started this all, complex heterodynes are everywhere and treated as a known quantity. And they’re straightforward once you get to know them better. I used to think that dealing with signals in quadrature, representing them with complex numbers, was done primarily as a way to reduce the sample rate by half. There are certain potential cost savings to that. The benefits are more fundamental: they eliminate the issue of having to deal with mirror images in the spectrum. Afterthought: the Fourier Transform is a Bunch of Averaged Complex Heterodynes While writing this blog post, I suddenly struck me: the discrete time Fourier transform is the same as doing a complex heterodyne to 0 Hz and then calculating the DC value by summing the samples, for all frequencies of interest. Complex heterodyne: DFTF: This is kind of obvious when you think about it, but I had never dealt with complex heterodynes so it’s something new for me. References Youtube - Recent Interesting and Useful Enhancements of Polyphase Filter Banks: Fred Harris polyphase blog series scripts Other blog posts in this series: Notes about Basic Polyphase Decimation Filters Footnotes If you’re wondering why it’s called ‘super’: it’s because the result of the heterodyne is a signal that is still in the supersonic frequency range, as in, above the audible frequency range. Before superheterodyne receivers, the radio signal of interested was heterodyned straight to the audio range. ↩

7th Feb 2026 1 votes

More in technology

A Dick Smith VZ200 without the Dick Smith (but with a serial port)

Australians! They walk among us! Do not be deceived by those charming faces! precious bodily fluids, we must remove the scurrilous larrikin influence of Australia upon our American home computers, and I know exactly where to start! Why allow Dick Smith's name (even though he'd already sold Dick Smith Electronics' controlling interest to Woolies by then but stop ruining my intro) to corrupt this, um, rubber-keyed diminutive beige home computer when we can return it to its prior, pristine, plasticky state as it once rolled out from a glorious Asian factory? Then, to avoid wrecking it with a botched logic board repair, we'll bolt on a USB serial port using the very latest peripherals from down under, write cycle-counted Z80 assembly language to blast data to it at 57.6kbps, and hack a few games. Because only this will make the VZ200 great again! TI vs. Everybody"), was a seismic event in computing. Hong Kong entrepreneurs Allan Wong and Stephen Leung were two of many to realize that the microchip would trigger a revolution in consumer electronics, and over several years accumulated sufficient funding to establish Video Technology Ltd in 1976. Their first factory operated from the Freder Centre in Ma Tau Kok, a semi-industrial area on the west side of Kowloon Bay. Initially VTech, as it became known, concentrated on video games, primarily as an OEM for the more profitable North American and European markets. Their first products were Pong clones, released in the United Kingdom as the Grandstand Adman T.V. Game 2000 (black and white) and Adman T.V. Game 3000 (colour) in 1977, both based on the Texas Instruments TMS1965N, TI's clone of the well-known General Instrument AY-3-8500 Pong-on-a-chip (compare with the rather more complex MOS 7601). Grandstand was a brand name of Adam Imports, then a substantial toy and games importer to the UK and for a period of time New Zealand, and had other Asian contacts, notably Tomy. If the picture on VTech's history page is to be believed, and I point out there are some verifiable inaccuracies on that page, VTech continued to produce other consoles for Grandstand/Adam like the (deep breath) Grandstand Adman Colour TV Game 3600 Mk III, which used a regular AY-3-8500 and was produced until 1979. For the portable market, VTech produced a number of handheld LED, VFD, LCD games, sold under various brands and through retailers such as Radio Shack. VTech was hardly the only such Hong Kong tech company, of course; across the Bay in Kwun Tong was EACA, established in 1975 by Guangzhou escapee Eric Chung. EACA also produced consumer products such as radios and its own video games, notably the 1978 Colour TV Game, also based on the TMS1965N and variously sold under other brands such as the Sonesta Hide-Away TV Game. Creative Computing in August), but in fact ripped off from and largely compatible with the TRS-80, and advertised as such (well, the compatible part, anyway). While there were some internal differences and significant changes to the keyboard layout, EACA mostly copied the TRS-80 system ROMs for production with only minor changes, shipping it with a licensed version of Microsoft Level II BASIC. At Summer CES in Chicago it directly competed with the APF Imagination Machine and the Texas Instruments 99/4 (the original), and indirectly with the Atari 8-bits which earlier debuted at the Winter show. While it lacked the $599 model's monitor (a TV set or Tandy monitor was required), it had a better keyboard mechanism and a built-in cassette recorder, and was variously announced between $500 and $600 with 16K of RAM [$2200-$2760]. Electronics Australia, John Kennewell's National Semiconductor SC/MP-based MINI-SCAMP, running the CPU at roughly 500kHz (based on a typical 2μs cycle time) and 256 bytes of RAM expandable to 1K (64K addressable). DSE advertised the machine as "33% of the cost of the EDUC-8," an earlier Electronics Australia bit-serial TTL hobbyist system inspired by the DEC PDP-8 and designed by Jamieson Rowe — remember that name — who became an enthusiastic proponent of the new machine. Over in Silicon Valley, Palo Alto arcade game builder Exidy had developed their own Z80-based computer in 1978, the Exidy Sorcerer, a premium system featuring programmable character graphics, a faster 2.1MHz CPU and a built-in internal S-100 bus. Exidy saw the export market as a growth opportunity and aggressively inked deals with multiple foreign distributors, including DSE, who were taking ready advantage of the Whitlam government's 1973 import tariff reduction to bring more finished goods to DSE stores. The Sorcerer arguably found greater success in Europe than it ever did in the United States, particularly in the Netherlands where the licensed Compudata Sorcerer became the default government-supported educational system, and certainly in Australia due to Dick Smith's dogged promotion. Still, even in the US it was considered relatively expensive at $895 [$4580], and with tariff and import costs tacked on it didn't price itself well to the Aussie working class nerd. Conversely, the EACA Video Genie had meanwhile achieved some popularity of its own within Europe, notably West Germany, in no small part due to its lower cost. That alone made it a logical system to transition to, and better still, EACA had absolutely no objection to DSE outright rebadging it as a Dick Smith unit. Micro-80 disliked the altered keyboard (no CLEAR and TAB, no left and right arrows, up and down replaced by ESCAPE and CONTROL), complained about its changes to the character set and video circuitry, found the built-in cassette deck hideous, and noted the lack of board sockets and the missing-at-launch S-100 expansion box, but approved of the "brilliant" and "attractive" appearance, its overall functionality, and most of all its purchase price. "Even if you buy a decent tape deck from your local Big W and put it in the System 80," Hartley concluded, "you end up at least $150.00 [US$130 spot, US$520 in 2026 dollars] ahead — and that pays for your next 16K of RAM chips ... the machine has an identical computing capacity to the TRS-80, for a lot less dollars." At this point it was inevitable someone would poke the Tandy bear, and that someone was Recortec (they're still around), established in Sunnyvale, California in 1969 to specialize in magnetic tape recording technology. In 1980 Recortec, also attempting to expand their product base, opened Personal Micro Computers, Inc. (PMC) as a new venture in Mountain View, initially entering negotiations with Exidy to buy out the Sorcerer — until, through EACA's American subsidiary, they became aware of the Video Genie. To PMC/Recortec, the Genie was a remarkable opportunity, a less expensive TRS-80 compatible system already selling in Europe and entering the Australian market, and PMC now had the chance to corner it for American distribution. The company immediately backed out of the deal with Exidy, bought up Genie distribution rights in August for the entire Western Hemisphere, rebranded it as the PMC-80 and launched it in 1981 with 16K of RAM for $675 [$2330]. InfoWorld was more complimentary than Micro-80 had been, noting PMC's planned Fastload high-speed cassette scheme, a 50-40 adapter to connect Model I peripherals directly and a true lowercase conversion kit. "Tandy," said columnist Tracy Deliman, "is apparently curious now" — and in particular their attorneys, who promptly filed suit in federal court against both PMC and EACA of America, claiming, among other allegations, that the name PMC-80 infringed their trademark and that EACA and by extension PMC had committed copyright infringement as well by substantially copying the system ROMs. (Tandy didn't dispute the BASIC ROM, as that was licensed from Microsoft, but EACA and PMC dragged Microsoft into court with them anyway as a third-party defendant. Microsoft was a lot smaller then.) In their motion to dismiss, PMC and EACA did not deny that the code had been copied and that then-current copyright law covered computer programs, but argued that section 117 of the in-force 1976 Copyright Act required the infringement claim to be interpreted according to the law prior to January 1, 1978 ("this title does not afford to the owner of copyright in a work any greater or lesser rights with respect to the use of the work in conjunction with automatic systems capable of storing, processing, retrieving, or transferring information, ... than those afforded to works under the law, whether Title 17 or the common law or statutes of a State, in effect on December 31, 1977"). Robert Peckham, chief judge for the Northern District of California, disagreed in August 1981 and denied the motion, writing "that section 117, as it existed in the 1976 act, was aimed at the problem of copyrighted material inputted [sic] into a computer, such as books, magazines, and even computer programs. It was not intended to provide a loophole by which someone could duplicate a computer program fixed on a silicon chip." Moreover, even if it did, "[t]he plaintiff has suggested that the evidence may well show that the chip was duplicated by first taking a visual display or printout of the program in question ... If this method of unauthorized duplication in fact is proved, there can be no doubt that the unauthorized duplication of a visually displayed copy of the program would fall within the reach of the federal copyright laws." The case was quietly settled out of court, and although it obviously didn't enjoin EACA outside of the United States, domestically PMC replaced the line with the CP/M-based MicroMate in 1983. By then, and unknown to his backers, Eric Chung's failed investments in the Hong Kong real estate market had put him millions of dollars in debt. In October 1983 he abruptly fled to Taiwan reportedly with $10 million stuffed in a suitcase, leaving EACA to quickly fold. Simultaneously, Dick Smith sold a 60% stake in Dick Smith Electronics to Woolworths (the Australian version) later in 1980 and then the rest in 1982, leaving only his name and his bespectacled grin to remain at the company he founded. Although DSE sold the Video Genie's modestly upgraded direct successors also as System 80 variations, the System 80 family never included the Colour Genie, EACA's last system before its ignominious demise. PMC's former headquarters in Mountain View are now collectively Google Building E475. first step to computer literacy. The success of the ZX80 suggested other dirt-cheap home computers might also flourish. VTech accordingly began developing a low cost computer of their own in 1981 that could work like the TRS-80, planning to adopt its BASIC (and thus its Z80 CPU) much as crosstown rival EACA did and speed time to market, but by explicitly abandoning compatibility they were free to slash its production cost as much as practical. A substantial reduction in part count became possible immediately by using the inexpensive Motorola 6847 Video Display Generator, introduced in 1978, which replaced nearly the entire video system: with minimal support circuitry, the VDG chip could generate a 32x16 text display, comparable to the TRS-80 and Video Genie's 32-column mode for colour TV sets, 64x32 colour semigraphics reminscent of the same, and a variable high-resolution bitmap display depending on available memory. On top of that, the entire system could run from the VDG's standard 315/88 (3.58MHz) crystal, including the Z80, and part count could be reduced even more for a black-and-white low binned model by omitting the colour encoder completely. Everything else (cassette output, keyboard lines) could be supported with discrete components and a handful of TTL logic, price could be adjusted further on the basis of included RAM, edge connectors wired to the processor bus would suffice for peripherals and expansion, and any sort of keyboard would be a step up from a flat membrane. While the low-end computer was in development, VTech also hedged its bets with a higher-spec video game system, which typical for the era (see, for example, the Intellivision Keyboard Component) was convertible into a home computer of its own. This console was likewise built from off-the-shelf components, using a 2MHz Rockwell 6502 CPU and the Texas Instruments TMS9918A for graphics, 17K of RAM (1K for the 6502's zero page, stack and low memory, and the other 16K for the VDP), and controllers that doubled as a membrane keyboard when used with the optional BASIC cartridge. It shared no parts or significant engineering with the computer prototype and proceeded along a largely separate development track, released to select European test markets first as the VTech CreatiVision in 1982. Meanwhile, Sinclair Research and manufacturing partner Timex Corporation subsequently joined forces to launch the Timex Sinclair 1000 in the United States, a slightly reconfigured ZX81 with NTSC-compatible video output and 2K of RAM, but otherwise identical. It hit stores in the summer of 1982 at the same psychologically desirable price point, now US$100 [$345], though Timex Sinclair didn't get the bargain American home computer market to itself like the ZX80 mostly did in the UK: it now had to contend with Commodore, selling the VIC-20 as their well-supported low-end system, plus the ailing Atari and their Atari 400, Tandy's own TRS-80 Color Computer, and even Texas Instruments, then only months away from igniting a price war using the TI-99/4A. Nevertheless, industry observers generally believed the T/S 1000 would be a strong competitor, and its debut led to an accelerated scramble from VTech and others hoping to duplicate the ZX80/1's success. VTech identified two overall product positions, a super-low-cost black-and-white variation with Microsoft Level I BASIC for selected markets, and a colour model with full Microsoft Level II BASIC, each of which VTech intended to sell both under its own name and as an OEM. To prepare for the American launch of the nearly complete low-end computer and the CreatiVision, VTech opened a U.S. subsidiary that year in Elk Grove Village, Illinois outside Chicago (the picture above is from VTech's history page). Computer Gaming World issue 3.2), it took the Las Vegas Convention Center, the Hilton Convention Center, the Riviera Convention Center, the rest of the Riviera and the old Rotunda to contain it all. Mattel introduced the Aquarius for $199 [$670] licensed from Radofin, also in Kwun Tong (with a 3.58MHz Z80A and 4K of RAM, cassette storage, no bitmap graphics option and a rubber chiclet keyboard), Texas Instruments hawked the TI-99/2 for $99 [$330] (with a 2.7MHz TMS9995, 4K of RAM, cassette storage, no bitmap graphics option and no colour, and a plastic chiclet keyboard), Sanyo proffered the PHC-20 also for $99 as the midrange of the pocket computer PHC-10 and higher-end PHC-25 (with a 3.58MHz Z80 clone and 4K of RAM, cassette storage, no bitmap graphics option and no colour, and a rubber chiclet keyboard), and at the higher end came the Panasonic JR-200U for $349 [$1170] (with a 0.89MHz 6800 clone and 36K of RAM, cassette storage, no bitmap graphics option, and a rubber chiclet keyboard), the NEC PC-6001 also for $349 (with a 4MHz Z80 clone and 16K of RAM, cassette storage, but bitmap graphics and a rubber chiclet keyboard that was quickly replaced with a typewriter-style one), and the Spectravideo SV-318 for $299 [$1000] (with a 3.58MHz Z80 and 16K of RAM, also bitmap graphics, and a rubber chiclet keyboard). There were multiple conversion kits to turn the Atari 2600 VCS into a low-end computer of its own, several with rubber chiclet keyboards, and even a completely unlicensed ripoff of the T/S 1000, the Unisonic Futura 8300 (with a rubber chiclet keyboard) for $99. Not to be outdone, Timex Sinclair themselves announced the T/S 2000, a modified US version of the ZX Spectrum, with 16K or 48K of RAM, a 3.5MHz Z80, colour bitmap graphics and a rubber chiclet keyboard; the 16K version started at just $150 [$500]. The latecomers arrived at the Summer CES in Chicago, though by that point the rot was already setting in. Against the background of Commodore slashing prices even lower on the VIC-20 and C64 to Texas Instruments' profound discomfort, Mattel suddenly decided the Aquarius needed a sequel (i.e., the other system Radofin was developing; price point to be determined but without a rubber chiclet keyboard); Timex Sinclair replaced the T/S 2000 with the enhanced T/S 2024 and T/S 2048 with higher resolution graphics, more RAM and a plastic chiclet keyboard, plus an upgraded T/S 1500 which was a T/S 1000 with 16K of RAM and a rubber chiclet keyboard; Rabbit Computer (who? also from Hong Kong) introduced its own Z80-based Rabbit RX83 with 2K of RAM, BASIC, three-channel sound, cassette storage, bitmap graphics and a plastic chiclet keyboard for $99; and Tomy unveiled the Tomy Tutor for "under $150" [$500] with a 2.7MHz TMS9995 (from a 10.7MHz crystal), 16K of RAM, cassette storage, bitmap graphics and a rubber chiclet keyboard. In the fall Tandy, never one to be left out of a race to the bottom, delivered the MC-10 Micro Color Computer for $120 [$400] with an 0.89MHz 6803 and 4K of RAM, cassette storage and bitmap graphics, and a rubber chiclet keyboard. Even Commodore, failing to learn its lesson from the critically maligned Max Machine, was working on their own ultra-low-end family of computers to follow on to the C64 — one of which (the 116) would have a rubber chiclet keyboard. And, oh yeah, one other system made its debut at the Winter show. COMPUTE! in their March 1983 reporting called it "the first under-$100 [$330] color computer." Anticipated to hit American store shelves in April, the new VTech VZ200 featured 4K of RAM (expandable to 16K for $45 [$150] and 64K eventually), 12K of ROM (remember these numbers) with BASIC, and a simple push-pull piezo for sound. Two kilobytes of the 4K was allocated to the 6847, which used it to generate its default 32x16 text display, 64x32 semigraphics or 128x64 bitmap graphics. It had built-in jacks for cassette, TV and composite video, and shared its booth with the CreatiVision which VTech planned to sell States-side for $189 [$630], with the BASIC cartridge for $10 [$33] and an inevitable rubber chiclet keyboard for $30 [$100]. It isn't clear where the name VZ200 came from, possibly a riff on "ZX," but the computer's low cost even amongst a sea of low-cost computers still attracted positive attention. Creative Computing got a 4K VZ200 in for review in their May issue (accounting for publishing delays this would have to have arrived in February or early March), though they noted that they had no chance to try the peripherals or software. The article has some glaring technical errors — among others, they said the CPU was a 6502 — but reviewer David Ahl called the machine "a compact microcomputer with a great deal of capability and many unexpected features at a very attractive price." Although the review found the 4K of RAM "sparse" and was openly critical of the keyboard, particularly the absent space bar, single SHIFT key, nonstandard layout and the keys' inconvenient tendency to stutter, Ahl was nevertheless impressed by the full-screen editor ("a pleasure") and the 12K ROM implementation of BASIC (unbeknownst to him, secretly derived from the TRS-80 with added support for the on-board hardware), concluding the VZ200 to be "a great value for the suggested retail price of under $100." There are certain attributes of the machine shown in both the COMPUTE! and Creative Computing photos that don't match released units, but we'll address this later on. Creative Computing Winter CES cover showing the VZ200, the Timex Sinclair 2000 (in its original form), the Texas Instruments 99/2, the Mattel Aquarius and the Spectravideo SV-318. Personally, however, I lump these computers together as the "crap home computers." I use this term with only love, and this uniquely terrible subtype of the home computer was indeed greatly loved, because as the ZX80 had demonstrated, ordinary people could now finally afford them. Heck, my first computer — the Tomy Tutor, introduced at Summer CES — was one of these 1983 crap home computers because it's what we could afford. We couldn't afford a Commodore 64 right then, but we could afford that. Not for nothing did Jack Tramiel thunder, "computers for the masses, not the classes!" What these systems all had in common, other than crummy keyboards, a striking preference for the Z80 and an unabashedly low starting price, was aspirational and arguably fraudulent marketing, plus inadequate specifications requiring upgrades at additional cost to be practical — if there were any to begin with — alongside substandard quality control, a poor selection of software, and weak to non-existent customer support. Made cheap to sell cheap, most of these computers failed outright (e.g., the Mattel Aquarius) or were never even released (e.g., the TI 99/2). The glut that hit the U.S. market that year not only soured many American consumers on home computers generally, but their game-heavy libraries were also likely a contributing factor to the 1983 video game crash. Although managing to move over half a million units, Timex Sinclair was not immune to this effect, and the company became unprofitable as the sales crossfire between Commodore and Texas Instruments forced the T/S 1000's street price below $50 in mid-1983. The situation was compounded by Timex's ill-considered decision to make the more expensive T/S 2068 (the eventual sole member of the 2000 family) largely incompatible with the ZX Spectrum, robbing it of the extensive British Spectrum software library, and the joint enterprise that was once expected to dominate the American home computer market collapsed in early 1984. Even large players like Texas Instruments and Warner Communications-era Atari took hundreds of millions of dollars in losses, with only Tandy (due to their strong retail presence) and Commodore (due to the C64's prodigious installed base and their vertically integrated manufacturing) able to weather the maelstrom effectively. Family Computing's inaugural September 1983 issue the columnists mention the SV-318 and even the stillborne T/S 1500, but nothing on the Lasers, and Creative Computing around that time was only running ads for the 3000. A few VZ200s were rebadged by Texas door-to-door nuisance business Dynasty Computer Corporation as the Smart Alec Jr., though like their multi-level marketing attempt at rebadging the spent Exidy Sorcerer, they sold barely at all (the company folded in November). On the other hand, the (now) Laser 200 got off the ground in Europe under its own name and others, most notably through Sanyo, but also through Salora, Seltron and Texet. It launched there alongside the black-and-white model as an ultra-low-end system, originally dubbed the Laser 100, but after the ROM change becoming the Laser 110 with the same Level II BASIC of the colour version. However, there was one market where the VZ200 had particularly strong success, and that was of course Australia, though not exactly in its original form. History does not preserve the thought process of Dick Smith Electronics management, but DSE's probable aim was to nose past the Commodore VIC-20 on price and capability (DSE themselves even sold them for a time), which was colour and shipped with 5K RAM. That immediately excluded the black-and-white variation, since the ZX81 had since landed in Australia and the potential profit margin wasn't enough to bother, and it is instead more likely that during negotiations DSE prevailed upon VTech to strengthen the colour system and keep the price low. VTech's solution was a small 6K daughterboard retrofit that could replace the 2K system RAM chip, internally expanding the unit to a more appealing 8K (the other 2K video RAM chip was left unmolested) while still being able to use previously manufactured components. This modified VZ200, badged as a Dick Smith computer and subsequently sold elsewhere by VTech as the Laser 210, appeared in the 1983-84 catalogue as "new for 1983" at just A$199 [approximately US$220 spot and US$720 in 2026 dollars]. Australian Personal Computer April 1983 preview speciously characterised it as "to Dick Smith's specifications" (likely only the RAM complement was), he got quickly used to the keyboard, approved of the BASIC implementation (faster than the ZX Spectrum's) and full screen editor, noted the characters to be "rather like those produced by the TRS-80 Color Computer" (true!), and compared the memory loadout favourably to the VIC-20's. He was similarly pleased with the cassette tape performance and the included documentation, with about his only complaint being the weak sound. His editor Sean Howard was equally impressed, famously remarking that "I'm certainly going to buy one," which DSE promptly and repeatedly used in their advertising. many VTech rebadges DSE would eventually sell. Although nothing was going to catch the Commodore 64 by then, which had the same stratospheric sales there as it did most other places, the DSE VZ-200 had the added good fortune of ZX Spectrum manufacturing issues that eroded its availability, giving the DSE VZ-200 almost unrestrained run of the Aussie low-end market from which the VIC-20 was already fading and the ZX81 all but gone. In 1984 it remained a strong seller at its new lower price of A$169, dropping to A$99 by the end of the year. I think that suffices for a more detailed backstory; we'll talk a little more about its later history in Australia and VTech's overall as a postscript at the end. For now, we'll turn our attention to this orphaned American unit. A convention I'll establish from now on in this and future articles: although Dick Smith was not consistent on the hyphenation, variously rendering it VZ-200 and VZ200, in the few places it appears the American version was invariably written without one, so I'll write "VZ200" for the North American computer and "VZ-200" for the Australian computer. biblioteca. You can also see more clearly that the red blurb on top advertising "WITH 4K BYTES RAM" and "NTSC 4K" is in fact a sticker, so this box was almost certainly not exclusive to North America and may not have been exclusive even to 4K systems. increased to $149.95. Neither price matches its well-documented MSRP of $100. Unfortunately I can't make out the actual retailer, even with an extreme enlargement. markedly predominant French labeling might have prevented its sale in the Montréal outlet, that wouldn't have enjoined it elsewhere. As far as its provenance, however, the fact it wasn't labeled that way suggests Canadian sale was not initially contemplated, and it also has U.S. Federal Communications Commission clearance (I'll show you in a bit). COMPUTE! and Creative Computing pictures. Those earliest systems are labeled as a "VZ200 Personal Computer," not a "VZ200 Color Computer," and the colour labels over the number keys were absent. In fact, this same keyboard is used for the B&W Laser 110, though those early VZ200 systems must have been colour given that their colour capabilities were widely reported at the time. On the other hand, the VZ-200 that appeared in very early Dick Smith marketing like the flyer above was labeled a "VZ200 Color Computer" exactly like this one, including the American spelling. The keyboard consists of 45 keys, with a bottom right SPACE key in the corner, only one SHIFT on the bottom left, and no ESCape key. Necessarily, many have multiple functions accessible with the CTRL key or CTRL-ENTER key combo. Most of these alternative functions are one-touch BASIC keywords like the Sinclair machines, though unlike those computers, you are not obligated to use them and can spell keywords out if you want. Semigraphics characters are also selected by key combination, as well as moving the cursor, inserting and deleting characters, and interrupting a BASIC program. was removed, because the 16K RAM expander connects there. The cover is not in the box and I'll presume it was lost or destroyed by the previous owner. That's annoying from a preservation perspective, but no great loss functionally, because we'll have something plugged in there pretty much all the time later on. on the Wayback Machine, and is the only other NTSC VZ200 unit I have seen myself. Both his and mine have a "VZ200 Color Computer" bottom plate with a copyright date of 1982, and both have US FCC clearance tags and an RF switch to pick the channel (2 or 3). There are passive cooling vents on both sides, though given the fairly small amount of clearance its rubber feetsies afford from one's desk I hesitate to say they'd be effective. (At least they let you hear the piezo speaker.) There is also a small red sticker on the bottom which on mine is partially missing, but Bill's has a complete one, which reads "U NTSC 4K." I put a small piece of transparent tape over the sticker on mine to prevent further damage. Bill's unit also has both port covers. Although the serial number indicates his is a rather older unit, which we'll in fact confirm later, the label and keyboard are the same as mine and not those earliest CES models. Both units have FCC Part 15 clearance, specifically as a Class B computing device for home use; at the time Class B regulations were very strict and we'll see a consequence of that inside. The FCC ID is BNX84H80-0323, with an equipment authorization (EA) applied for February 10, 1983 and granted May 9, 1983. Accounting for publishing delays, this means Creative Computing must have had a pre-authorization prototype for their review. VTech later applied for a revised equipment authorization on July 11, 1983 and was granted BNX84H80-0323-1 on October 11, 1983; this change may have been for redesignation as the Laser 200. The listed address for VTech in Hong Kong appears on all entries for FCC grantee BNX and doesn't seem to have been their corporate address at the time, but the tester in both EAs is one Thomas Cokenias from Electro Service Corp., at 1116 Ninth Avenue in San Mateo, California, and Cokenias and Electro Service are seen in other EAs around that time for other Hong Kong manufacturers. However, the given address appears presently unoccupied, and the California Secretary of State indicates Electro Service is no longer in business. except its 16K RAM expansion (due to differing address mapping) would work on the older machine because many of them were effectively unchanged except for the labeling. (The address mapping difference is also why the VZ-200 16K RAM expander will work on the VZ-300, but only giving you an additional 8K.) Although there have been various other members spotted in the Laser 300 series, apparently differing only in RAM size or case type, none were reportedly sold widely if at all. The VZ-300/Laser 310 is otherwise nearly totally compatible with the VZ-200/Laser 210 except for its clock speed. As NTSC compatibility was no longer required, VTech switched to a single 17.734475MHz master crystal divided by four for the 4.43361875MHz PAL colourburst and five for the 3.546895MHz CPU clock. (The Motorola 6847 VDG in the VZ-300 and PAL systems generally is clocked with an altered signal which I'll talk about when we open the machines up.) Although that makes the VZ-300 slightly slower at 99.09% the speed of the VZ-200's 3.5795454MHz (315/88) oscillator, in practical terms the difference was imperceptible with most existing software — but of course it's going to be a problem for us later. VTech did no other upgrades, not even offering an alternate 6847 font ROM with lowercase, though we'll talk more about that when we get to the innards as well. a Christmas gift from my wife (I married well). The DSE pricetag on the back gives its price as A$9.50 but I don't know where or when it was originally purchased. While the TRM does not approach the sheer detail of, say, the Commodore 64 Programmer's Reference Guide, which even has an exhaustive memory map covering its 64K of RAM and 20K of ROM, it does include documentation on most of its important RAM locations and a full set of schematics, some of which we'll be using in this very article. A similar manual exists for the VZ-300 with its own set of schematics, which also discusses the VZ-200 and has some details not in the earlier text. (for Los Angeles residents, read "Vernon"). The North Ryde building was subsequently occupied by German truck and bus manufacturer MAN, whose old stripped logo can still be seen on the facade, and is now split between multiple tenants. Level I BASIC is quite, uh, basic, evolved by TRS-80 designer Steve Leininger from Li-Chen Wang's "copyleft" Palo Alto Tiny BASIC which Leininger substantially altered, reworking its structure and adding floating point math (because it couldn't accept Charles Tandy's salary when he typed it in), two string variables and a single array, and support for the TRS-80 hardware. VTech's version is similarly modified to support its own architecture but is otherwise the same. Level II BASIC is Microsoft BASIC, derived from Microsoft's own Extended BASIC on the Altair, and again VTech initially imported it nearly unchanged except for adding support for their hardware and graphics, which replaced some of the keywords (like TROFF with COLOR). Proof can be seen by comparing the order of keywords in their token tables, which would have no reason to match as precisely as they do unless they came from the same origin. Another persistent relic in the VZ BASIC ROM is the old Microsoft two-character error message table (e.g., ?SN ERROR instead of ?SYNTAX ERROR) even though the existing code never references it. Tandy's apparently successful legal action against EACA and PMC spooked VTech that they could be sued in the same way — but by Tandy and Microsoft. The company's solution was to dump Level I BASIC entirely, eliminating any objection from Tandy, and for Level II BASIC to secretly null out a large number of entries in the ROM BASIC keyword table potentially unusual enough to be used as evidence of copying. The tokens for those keywords still remained valid, however, and the ROM code for them also largely persisted. Disk-specific keywords were likewise omitted in the same fashion, at least until the VZ-300 disk drive debuted, but unlike the other gutted keywords were instead vectored through RAM for later expansion. These keywords (in token order) are CMD, RANDOM, DEFINT, DEFSNG, DEFDBL, RESUME, ON, OPEN, FIELD, GET, PUT, CLOSE, LOAD, NAME, KILL, LSET, RSET, SAVE, SYSTEM, DEF, DELETE, AUTO, FN, VARPTR, ERL, ERR, STRING$, INSTR, TIME$, MEM, FRE, POS, CVI, CVS, CVD, EOF, LOC, LOF, MKI$, MKS$, MKD$, CINT, CSNG, CDBL and FIX. Because their implementations often remained present, it was possible to resurrect many of them by hooking into the RAM vector for tokenizing "new" BASIC keywords, and some BASIC extensions did just that. The earliest versions of the Laser/VZ ROM use light green text on a dark green background. Without a colour encoder this would produce light text on a dark screen, which is indeed what you see on a black-and-white Laser 110. Bill's earlier NTSC VZ200 is the same way, using the same ROM 1.2. By ROM 2.0, as in this VZ-200, the display became dark green text on a light green background, like the Tandy Color Computer which uses the same MC6847 video chip. do have a 4K system, just a very late one. You may have noticed the oddly convoluted BASIC statement I entered to get that figure. That's because I had to avoid typing the number 5: the key didn't work. Normally the VZ200 will beep as you press keys, but there was no beep and no response when I did. In fact, an entire run of keys (5, T, G, B and N) was not working, and I was only able to enter the PRINT keyword by pressing SHIFT-P. We'll need to fix the keyboard now to do anything substantial with this machine. did appear to work. Still, it seemed like the most reasonable place to start, so let's crack the computer open. another VTech unit, the unrelated Laser 50) and a plastic cover sheet with slits for the top ports' card edges to reduce dirt getting inside. Also visible on the top right/northeast side is a sprawling heat sink screwed to the fin of a 7805 voltage regulator peeping out from the lower right/southeast corner. This heatsink sits under the ventilation slits in the top case. Much of the logic board is covered by a large sheet metal Faraday cage serving as an RF shield, festooned with soldered metal braids connecting everything to the ground plane, and then the whole assembly placed on an irregularly shaped metal plate on the bottom with the piezo. As mentioned, in those days the FCC was very strict about radio interference from computers and video games, particularly home units where a Class B device (as this is) had a 10dB lower maximum than a commercial Class A one. Some systems like the Atari 400 solved this problem by effectively encasing the entire system in a molded metal endoskeleton and placing as few holes in the case as possible from which radio signals could emanate. That made for a very sturdy computer but one more expensive to manufacture, so VTech went for this cheaper, hackier approach which was no doubt iterated upon until it just cleared the bar. One major component is not under the cage, however, and that is the Motorola 6847 VDG video chip, in a plastic carrier (MC6847P) with a date code of 24th week 1983. It's not precisely clear why that is, but it can be seen that the left board is connected to some of its lines. PAL synchronous demodulator, a surprising choice, since the MC6847 is usually paired with the MC1372 NTSC colour modulator. The 6847 emits YPbPr (as Y, B-Y and R-Y) video, which in the black and white Lasers only the Y (luma) signal is used. In other systems like the Tandy CoCo the MC1372 takes the YPbPr lines, encodes the colour, and emits a signal suitable for a television set and/or composite video depending on the specific components. The TBA520 can be made to do the same task, even generating NTSC colour with the right crystal, albeit with more supporting electronics. (I presume it was less expensive than an MC1372, plus there would be the advantage of not having a different chip for Euro and Aussie systems, since the TBA520 can obviously do PAL video too.) Reference B-Y and R-Y signals are generated to the TBA520 using the 3.58MHz (315/88) NTSC colourburst oscillator on the left board, while the B-Y and R-Y signals from the MC6847 are passed on different lines, with the MC6847 running on the same 3.58MHz clock. We then pulse the TBA520's line inputs at the necessary horizontal rate, approximately 15.7343kHz, causing the TBA520 to emit a single NTSC chroma signal (on its G-Y pin) from the 6847's PbPr signals. (This theoretically makes it possible to get proper S-video out of the VZ200, though we're not going to try that this time.) Discrete components then combine the luma and chroma into composite video for both the monitor connector and for feeding into the separate RF modulator. this problem), so I decided to just clip the tabs with angle cutters. One of those tabs is here over by the 7805 ... Scriptovision Super Micro Script, which has a 6802 CPU (a microcontroller version of the 6800) and a 6847 VDG, RAM, ROM and a simple keypad for entry. In that article I mentioned that those were enough to make it almost a home computer (replacing the ROM on the Super Micro Script to make it one) because home computers like the VZ-200 have a similar architecture. Now we'll prove the comparison is valid. In this view, the MC6847 is the large DIP on the far left/west side. The chips in the back, going from left to right, are a Hitachi HM6116P-2 2K static RAM used for video memory, the CPU, a clone SGS Z80 (the former state-owned SGS Microelettronica "Società Generale Semiconduttori" of Italy prior to merger into STMicroelectronics in 1987) with a date code of 19th week 1983, a Hitachi 74LS139 and 74LS32 used as part of the address decoding for the memory mapped I/O range, and lurking in the back right (east) corner a 74LS174 6-bit flip-flop serving as a latch register. In the front, also left to right, are a Hitachi 74LS245 octal bus transceiver which bridges the shared 2K video RAM between the CPU and VDG, a Hitachi 74LS04 hex inverter used for various tasks such as the CPU reset circuit, a Hitachi 74LS244 octal driver, another Hitachi 6116 2K SRAM (the system RAM this time), and the two 2364 8K system and BASIC ROMs with date codes of 27th week and 32nd week 1983 respectively. On the Dick Smith schematic using the same order, these chips are numbered U15 (the 6847), then U7, U4, U3, U2 and U1 (74LS174), then U14 (74LS245), U13, U12 (assumed), no designation for the single 2K SRAM, and U10 and U9 for the ROMs. The ROMs are unsurprisingly the newest chips in the system and mean the computer could not have been assembled earlier than then, likely making it part of the last production runs before VTech abandoned the line in North America. Obnoxiously, everything is soldered down; there are no sockets (cheap!). However, there are a number of unpopulated pads. Some extra pads near the ROMs were clearly intended to accommodate larger-capacity chips, and later machines indeed use a single 16K ROM with a change in board jumpers nearby. (Braver folks than I have extracted VZ-200 boards with 2364s and found bodge wires underneath. Apparently underpaid labour and smaller ROMs were less expensive at the time. Cheap!) There is also another set of pads near the VRAM chip between it and a big block of through-hole resistors. It's not clear what these pads were meant for, though the VZ-200 schematics show another resistor bank there serving as pull-ups. This bank is drawn on the schematic with dotted lines unlike the other set, so perhaps they were eliminated for cost reasons (cheap!). Now let's talk about what we don't see. One thing we don't see here is a 3.58MHz master crystal. That's because ... we already saw it. The entire system runs from the 3.58MHz crystal on the colour encoder, so everything is precisely synchronized to the same clock source, including the TBA520, the MC6847 and the Z80. It is therefore impossible to merely switch the colour encoder boards and turn an NTSC VZ200 into a PAL VZ-200 or vice versa because of the added line padding circuit in the PAL unit and the absence of a second system crystal in the NTSC unit (cheap!). What about the VZ-300, where there's no 3.58MHz crystal either? In that system, the VDG is entirely clocked by one of the gate array chips, providing the same line padding logic, but also using the same 3.54MHz clock as the CPU which is apparently "close enough." Another thing we don't see is something like a Motorola 6883 synchronous address multiplexer. Recall that the 6847 VDG has no externally exposed registers of its own (compare to, say, the VIC-II in a Commodore 64). Things like video modes and character attributes are twiddled by chip lines; for character attributes these lines are often wired to certain data bits from the video RAM, but to dynamically set a video mode under software control requires external hardware. Also, the VDG makes little attempt to cooperate with the CPU as to when it accesses video memory, other than indicating when it finishes a frame which is likewise asserted on one of its pins. In the Tandy Color Computers (prior to the CoCo 3, which uses a GIME), the MC6883 SAM sits between the 6809 CPU and the 6847 VDG and arbitrates all of this, managing the display mode and system timing, and servicing the VDG while the CPU is on the bus. On the other hand, this approach was judged too expensive for the cut-down Tandy MC-10, which instead uses a series of flip-flops to interleave bus access between its 6803 CPU and the 6847. A single 74LS245 bus transceiver allows the CPU unrestricted access to the video RAM when the processor is accessing memory. Neither approach is suitable in this case, however. An MC6883 would be too expensive for the VZ200 also (cheap!), and the Z80 sits on the bus longer during a CPU cycle than a 6502 or 6800-family chip would, so the MC-10's interleaved approach won't work either. As it happens, the VZ200 does nearly exactly what the Scriptovision Super Micro Script does: during CPU video RAM access, the VDG's memory fetch is immediately suppressed using its MS pin and the 74LS245 bus transceiver temporarily kicks it off the bus. The contention problem is solved in both cases with software (cheap!) by simply not doing anything with VRAM until the VDG indicates it's between frames and not reading screen memory. The only difference is how they find that out; the SMS busy-waits on the VDG's FS signal before doing a screen update, while the VZ200 just wires FS to its IRQ line — screen updates using ROM routines are batched and when the interrupt is triggered, the ROM then blits the deferred changes to the screen all at once. Of course, if you write directly to the video RAM when the VDG is accessing it you'll get intermittent artifacts, and I'll show you what that looks like, but you can just watch for the IRQ yourself if you really care about it (many programs didn't). Otherwise, the VDG's mode pins for bitmapped graphics and alternate colour selection are handled through bits in the 74LS174 latch within the memory-mapped range, which also handles the piezo and cassette output. Overall this is a good demonstration of how the SMS was almost a home computer, because here's a home computer whose video architecture was almost the same. A missed opportunity with the more upmarket VZ-300 was the potential for lowercase or at least an alternative character set, especially because it even got a word processing cartridge released for it later (we'll play with it, it works with the VZ-200 also). An external font ROM can be lashed to the 6847 with a bit of additional circuitry and the Super Micro Script has one to generate higher-quality character glyphs. It seems like VTech could have done something like that controllable by another latch bit, and with a six-bit latch there are a couple more data bits there that could be used, but I guess that was either judged too risky or not even thought about. On both systems the 6847 INT/EXT pin that would have controlled this is merely hardwired to ground. One note about the metal RF shield: if the cage is not pulled back down into position, it may distort and contact some of the pins on the 74LS174. This will cause weird graphical artifacts and knock out the piezo (no keybeep). It doesn't appear to harm the computer, but I was very careful to ensure it was bent back to as similar a position as before after this happened a couple times. Obviously this is no problem if you just completely take it off. that), which is why they could never complete a junction and be sensed. On the other hand, the 6, Y and H keys on that column are wired separately into the ribbon cable, so they were unaffected. I'm not sure how such a fault would have happened. The keys move, but they don't sweep or scour, and ordinarily they shouldn't be contacting the painted portions anyway. It also doesn't seem likely to have been a factory defect because that would have made the computer very difficult to use, and this computer was clearly used. I pondered the best way to fix it, since any repair would have to be flat or it would distort the key sheet on top (e.g., no solder blobs, no top bodge wires). I have a circuit pen I could use to draw a new trace, but it's temperamental, and I didn't want to do something I couldn't undo later in case it wasn't actually the problem. do work! Do not overtighten them or you will interfere with the conductive nubs being able to make contact. I had a couple dud keys initially after this which were returned to life by slightly loosening the screw nearest to them (to my great relief). COLOR ,1): #00ff00 and red is #ff0000 and so forth. That is definitely not the case; in fact, the default VDG palette is rather a bit muddy, with relatively poor saturation. For example, black (what the border is supposed to be) is often more like a very dark brown, buff is a dirty off-white, magenta becomes a flaccid purple where the red is a little too low, and what the documentation calls cyan comes out closer to seafoam green. Unfortunately, many simpler or older emulators provide an excessively rosy (no pun intended) simulation of what these typical home computer implementations usually generated. MAME uses the correct palette and the VDG's Wikipedia entry has a credible synthetic screenshot based on the YPbPr values in the datasheet, which you can compare with the real composite grabs above. absolutely capable of good output, but to do so it needs a quality encoder, and the VZ's ain't it. The best colour I have ever seen from a VDG is actually the Super Micro Script's, using a very high quality output stage as shown in the actual grab above, and comes out vibrant, beautifully saturated, and fabulous on a CRT. You would expect that, however — it's a $500 prosumer video titler from 1985, not a $99 crap home computer from 1983. The next order of business is software. I'd rather not use tape or audio files, and I don't have the disk drive. Fortunately, because the VZ series is so beloved in Australia, those wacky Aussies occasionally create their own modern peripherals in between prawns on the barbie. If you have a VZ-series computer, then you need the BennVenn VZ300 SD Loader. It is fairly inexpensive and provides you a way to load software into your VZ-series computer via SD card, along with topping off the RAM, even more memory with bank switching, and optional solder-yourself connectors for gamepads and I/O expansion. I figured it might be fun to build some hardware for it (and we're going to create a very simple expansion ourselves for the VZ200 in this article), so I ordered the full kit. It works well for my purposes and my wife has ordered another for the VZ-300 now at my in-laws' house in regional NSW. However, I am neither affiliated nor associated with Ben, merely an overall satisfied customer, so this is the part where I will also make three gentle constructive complaints about it. First, things like new firmware are largely delivered through a private Facebook group. This group appears to be very welcoming to new members, but it requires you to be on Facebook, and I don't want to be on Facebook. I managed to get the current firmware another way, and I will be putting it in the Github repo for this project so you don't need to join Facebook either. (If you do want to join, however, I'm sure the "VZ200 VZ300 Laser210 Laser310 fans" group would love to have you.) On the other hand, Ben was reasonably accommodating of my questions over E-mail which I did appreciate. looked like the right one until I got out the continuity tester and realized the actual fault was elsewhere. I'm not sure why some of them were routed around instead of in a straight line. Salora Fellow, a Finnish rebadge of the 4K PAL Laser 200 (by contrast, the Salora Manager is a Finnish rebadge of the CreatiVision-derived Laser 2001). This is determined by a simple memory map check in a snippet of its VHDL that Ben shared with me: x"B7FF" ) else '1'; --B800 or higher RAMarea200 <= '0' when (Address > x"8FFF" ) else '1'; --9000 or higher RAMareaSelora <= '0' when (Address > x"7FFF" ) else '1'; --8000 or higher put the reset button, nor was I particularly enthusiastic about drilling a hole in the case to make one, first due to the questionable quality of the plastic itself and second for purposes of historical preservation of an unusual machine. a Gremlin Blasto arcade board (where we had no reset circuit of any kind) that its 8080A CPU could be crudely reset by simply putting a pushbutton switch between its reset pin and ground. As the Z80 can serve as a drop-in upgrade for an 8080, it can be reset in the same way. .VZ format, which has a trivial 24-byte header indicating filename, starting address and type (BASIC or binary). The LOAD command will conveniently auto-execute binary files when the load completes. You can find many programs on Dave "Bushy" Maunder's exceptionally comprehensive site containing software, photographs, articles and documentation, and most software he offers can be copied directly to the card and used immediately. 2018AD. There's not much of a VZ200 demoscene, but there are a few out there, and this exceptional demo is unquestionably one of the best. It will not run correctly on this particular computer because the video timing is different — not because the clock speed is faster, like you'd see between an NTSC Commodore 64 which is faster than a PAL Commodore 64, but because the NTSC VDG draws the screen faster and thus fires the end-of-frame interrupt more often, messing up synchronization. For that, you'll just have to watch this YouTube recording on actual PAL hardware. It's set up like a trackmo, streaming cassette program data from one channel of a specially recorded CD audio track and using the other channel for music (admittedly a bit of a cheat but the music is excellent). If you didn't think rotozooms, raster splits and even FLD-type effects were possible on the 6847 VDG, then you're in for a treat. Another neat trick is the doubled vertical resolution while drawing the Kefrens bars. The source code is even available for your education. entirely done with semigraphics. It restores the old light-on-dark screen colours used in earlier ROMs with POKE 30744,1 (this works with twiddling the CSS pin with COLOR,1 also). Dubois and McNamara (i.e., Greg Dubois and Tricia McNamara, though Greg did all the programming), who created various titles for a number of DSE systems that were sold in stores. I should note that for many games, including this one, the more-or-less standard control keys are Q and A for up and down, M and comma for left and right, and where a fire button is used, typically SHIFT (SPACE for secondary). Note the "snow" in this shot — that's because the program was animating screen memory at the same time the VDG was trying to read it, so the VDG's memory access got briefly suppressed during that period. Because the display scan can't wait for the VDG to be enabled again, the result is a brief splat of garbage on that line until the VDG is allowed to proceed. Dubois could have simply waited for the VDG's next interframe interrupt, but there's also only so much time between frames before the VDG will start drawing again. As a result, for many programs where significant CPU time was required to do screen updates, outright ignoring the video artifacts turned out to be the least bad approach. that was) by Stephen Clarke. It shows loading from the SD card, the title and options screen (even the VZ200 had software pirates), and then playing the game, which worked fine on the keyboard. This and the other recordings I did for this entry were generated from the composite capture rig for video, but for audio using a microphone near the VZ200's piezo speaker to capture sound (since there's no audio out). You'll notice I'm pounding on the keys a bit, which the microphone faithfully picks up, though you do have to hit the keys with a bit of, shall we say, deliberateness to get them to register. Lemonade Stand: Meatpies, where you sell pies. Yes, the meaty kind, America. glasses pies you want to make, how many advertising signs you want to buy, and the price per item you want to request. Other than the pie business, Larry Taylor's port of the game seems to be heavily influenced by the well-known Apple II version and includes the same sort of simple graphics for weather reports and the like. It oozes professional quality with a slick title screen and menu, and is an extremely fun puzzler to boot. The aim is to create a factory from various primitive machines (paint, rotate, hole punch) that will generate a specific product. It is so well animated and so thoroughly polished that it deserves this short video to fully appreciate it. copious documentation for all its features, as it was a surprisingly credible word processor on par with at least, say, Color Scripsit on the Tandy Color Computer, though Color Scripsit is some years older. Although Wordpro came as a cartridge intended for the VZ-300, it would work on a VZ-200 with correspondingly less document memory available, since it occupied the slot where the RAM expander would go. On the other hand, the program seems to be calibrated for a VZ-300 keyboard since on this VZ200 the keys seem to frequently stutter. (I'll talk about how I got it to work on this 4K system later, though it would not have been possible without the BennVenn RAM expansion.) see this emulation — is also notable because the CoCo 1/2 and the VZ-200/300 all lack lowercase, so both programs solve it in the same way by displaying "capital letters" in reverse video. Wordpro's user interface is more sophisticated than Color Scripsit's, but it's also newer. The lack of a lowercase option on the VZ-300 was again a real missed opportunity, and with the number of machines DSE was buying you'd think they could have talked VTech into engineering a solution. Although Wordpro supports both disk and tape, the BennVenn cartridge currently doesn't emulate them sufficiently for Wordpro to use it. Perhaps this is a hack we can do some other time. a la Night Driver where you are apparently an alien with wings and feet ... It's only rendered at 90 degree angles, but it's fast and well-written, and deserved another video. Some games, though many were fine, would show a weird line pattern over certain sections of the screen like this port of Exidy Circus/Midway Clowns. The pattern was annoying, but in the first few games I played where it manifested, it appeared to be cosmetic (it does not restrain the jumping figures here) and I initially chalked it up to some other undiscovered difference in this NTSC unit. LDI that copies the contents of the address referenced by HL to the contents of the address referenced by DE; the instructions LDIR and LDDR expand upon it, running LDI repeatedly and decrementing the count in BC each time until it reaches zero, respectively incrementing or decrementing HL/DE on every step. The most obvious application for these instructions is copying a block of memory elsewhere, but a less obvious application is using them to fill memory. Consider this segment of actual code from Invaders (dumped with z80dismblr): Ignoring the instruction at $8ca1, you can see that this is setting the source to $7000 — i.e., the start of VRAM — and the destination to $7001 (?!), for a total of $0820 bytes (the zero test is post-decrement). Now, what would that accomplish? Just before the LDIR, we set the contents of $7000 (in HL) to zero. Let's step through the process LDIR takes manually. $7000 is first copied to $7001, which is now zero as well. HL is incremented to $7001, DE to $7002, BC decremented to $081e. Next, $7001 is copied to $7002, but $7001 had zero in it because it was copied from $7000, so all three locations are now zero. HL is incremented to $7002, DE to $7003, BC decremented to $081d. Then, $7002 is copied to $7003, so now all four locations are zero, and so on, filling all intervening locations with the immediate value before. At the end, when BC finally gets to zero, all locations from $7000 to $781f inclusive (i.e., the entirety of video memory and a little past it in system RAM) will have been zeroed out. This is how Invaders clears the hi-res screen and it is indeed faster than a naïve loop, especially for large tracts of memory. But our obnoxious little plastic beast here throws in a wrench by having a location where the RAM isn't working properly. When the "copy" gets to that point, because the copy is only between adjacent memory locations, for every subsequent location the stuck bit will be propagated forward and faithfully copied to each and every byte afterwards. That's also why the pattern doesn't cover the whole screen, because the problem doesn't actually manifest until the "copy" operation arrives there. In fact, in the process Invaders was unwittingly corrupting some of its own game variables with the same stuck bit when they should have been zero, possibly another reason why it wouldn't run correctly. The direct and most definitive solution would be to "simply" replace the VRAM chip, and I even have 6116 SRAMs in stock, but I warned you this is a very cheaply made PCB. Far better repairpersons than I have tried and failed to replace chips on these computers without requiring a lot of rework and bodges, and the prior portions of this article should have already convinced you it's only by the grace of God I haven't soldered my own fool face to the workbench yet. I did not want to try replacing that SRAM chip solely because of one stinking bad bit; I was only likely to make a bigger mess or render the computer completely inoperable. But again: we have an alternative. This fill trick was not universally used or even known by all programmers at the time. The games that do work clear the screen with a simple loop that doesn't propagate the bad bit forward, which works because the store doesn't depend on what memory contents are already "there." Likewise, VTech doesn't seem to use it in the ROMs, which is why the problem didn't manifest during the demonstration tape or with BASIC programs drawing to the screen with BASIC keywords. Most VZ programs are small enough and this code idiom distinctive enough (and usually only present once) that such code can be found and patched to use a slightly slower but functional loop. As such a loop would generally require more bytes, the patch could either direct execution to a tacked-on routine to do the clear, or we could patch it to point to a standard routine in memory. And how are we going to get a standard, always-present, stock routine into memory to do that? Easy: we're going to soft-alter the BennVenn SD loader's firmware. No, stop laughing, because we have a simple means to accomplish it. At the same time we'll combine that with a serial port loader so that we can test these programs live without having to constantly swap the SD card to and from the Talos II, so we'll also build it a bitbanged serial port (I said stop laughing). There were homebrew serial devices back in the day for these computers, so consider this one merely another entry from a venerable tradition. .VZ format. Here's a simple, complete example of "Hello World" showing how to construct that header and which can run directly from the card, demonstrated in the screenshot. This is one of several files you will find in this article's Github repo. As with all our assembler projects except for the 6502 and PowerPC, we crossbuild using the Macroassembler AS. The 24-byte header marks this as a machine language program that starts at $8000, the beginning of the extra memory furnished by the BennVenn device. Although the magic number VZF0 (for BASIC programs, which always start at $7ae9) or VZF1 would appear to be critical, the firmware doesn't seem to check it on loading or even generate it on saving, only that the byte just before the starting address word (everything is Z80 little-endian) is either $f0 or $f1. Upon execution our program then calls a "display null-terminated string" routine in the VZ ROM, the update for which is pushed to the screen during the next VDG interframe period, and returns to BASIC. The binary is assembled with AS like so, using a simple Makefile: >hello.vz (44 Bytes) % xd hello.vz 00000000 56 5a 46 31 48 45 4c 4c 4f 00 00 00 00 00 00 00 |VZF1HELLO.......| 00000010 00 00 00 00 00 f1 00 80 21 07 80 cd a7 28 c9 48 |........!....(.H| 00000020 45 4c 4c 4f 20 57 4f 52 4c 44 0d 00 |ELLO WORLD..| 0000002c LOAD"HELLO" will load and immediately execute it from the given entry address (which is both the load and execute address), as shown in the screenshot above. the Github repo I have a small program to cycle the LEDs on his proto board, which you can see in this brief video. It sets all GPIO pins to output and lights all 24 green LEDs connected to them (the red ones are check LEDs for 3.3V, 5V and 9V), then cycles a dark one through them from left to right until a key is pressed. This is done by hooking into the interrupt routine called when the VDG completes a frame, the only regular timesource on an unaltered VZ200, and used back in the day as a simple clock by various programs. Every third tick of the interrupt routine, this code runs: It rotates an in-memory image of the GPIO pin values, then emits that to the I/O locations. Because the rotation is to the left, we can see that the GPIO lines must also be oriented little-endian, i.e., the least significant bit of each GPIO output register is on the left. This then informs how we'll set up our bitbanger. A half-duplex system will suffice for downloading, since the sender will wait for us to indicate receipt between packets, and that will let us concentrate entirely on receiving until a full packet is obtained. The absolutely fastest speed we can receive at is generally limited by how quickly we can clock data bits into an accumulator from the receive line. If we connect the receive line to the least-significant input of one of the GPIO registers (we'll use the leftmost for convenience), we can do it in 23 cycles: The Z80's clock speed (in any of these systems) does not neatly divide into any standard bitrate, but theoretically 23 cycles per bit gives us a maximum possible transfer speed of ((315 000 000/88)/23) =~ 155632.4 bits per second. That suggests you might be able to get 115200bps with an unrolled loop, but at speeds this fast time required for other tasks starts to be a concern, such as storing to memory, checking how many bytes have been received, and branching back to get another, all of which together will certainly be more than 23 cycles. The other problem is the time required to sense the start bit, because this can occur at any moment, and hardware UARTs generally end up repeatedly snooping the line at some multiple of the bitrate to ensure they won't miss one. We, on the other hand, can't even check for a start bit at just twice 115200bps. In fact, the fastest we can check for a start bit (a zero) is which because of the unavoidable branch is actually longer than the time to clock in a data bit! However, these numbers do suggest that half that speed, i.e., 57600bps, is plausible. Flipping the equation around, that gives us a relatively generous ((315 000 000/88)/57600) =~ 62.1 cycles per bit, long enough to do our housekeeping tasks on each byte, and our tight startbit loop can poll the line at ((315 000 000/88)/25) =~ 143181.8 bits per second (a familiar number to some of you), which is at least twice the data rate and should be sufficient for the sort of continuous data transfer we'd experience receiving a data packet. We will target this speed. (Note from the future: an early draft used in a,(c) in the startbit loop, which is a 12-cycle instruction. This single extra cycle reduced the startbit poll rate to 137674.8bps, and at that speed multiple bytes got missed and/or corrupted. We are probably only just fast enough to make this work.) Parenthetically, VZ-300 owners in the audience will now have asked if this will work for them. If we substitute its lower clock speed at 57600bps, we get ((17734475/5)/57600) =~ 61.5 cycles per data bit and a maximum startbit poll rate of ((17734475/5)/25) == 141875.8 bits per second exactly. Because you can sample a little bit faster but never slower, we would need a separate version for the VZ-300; the same code will not work reliably on both. Sorry! That will be the subject of a future article. Note that by making receive fast, we made transmit slower: unless we occupy the least-significant bit of another GPIO register, which seems rather wasteful, the next fastest position is the second-to-least significant bit. This snippet needs no less than 31 cycles to send the next bit of a character stored in a register other than the accumulator (here we'll use B): If we had to completely guard that GPIO register from interfering with any other GPIO pins on the same register, it would be even longer because we would need to read the current state and then do the bitmasks. Mercifully we'll just refuse to support that, and as 31 cycles is still well within our 62 cycle maximum per bit, she'll be right. Texta Sharpie. If you don't have his board, you can still figure it out a little less conveniently with a voltmeter. Ensure all pins are set to output and turned off (something like FORI=68TO73:OUTI,0:NEXT will do from BASIC). The only live lines at that point should be ground, 3.3V, 5V and 9V. Find ground first, which you might do by checking for continuity with the ground test point Ben provides on the board, then use that as your common to find the voltage pins. Mark those; the rest are GPIO. Turn on pins one and two individually (OUT 71,1 or OUT 71,2) and look for voltage. don't connect the 3.3V line. Instead, for the programs below, ensure the HW-597 is already plugged into your host (such as with a USB extension cable) and showing bright status LEDs before powering on the VZ-200, or it may try to unsuccessfully power itself from the other lines and get a little daft. BITS) toggles the screen colour as it sees activity on the receive line. This is easiest to watch at a slow bit speed of around 150 baud or so. Here, I hooked it up to the Talos II, ran picocom -b150 /dev/ttyUSB0 (adjust for the path to your device), and just banged on the T2's keyboard. If you get alternating flashes of green and orange on the VZ while you do so, then your receive line at least has basic connectivity. ASCII), effectively one half of a very slow terminal program. We'll use the internal ROM routine to display a character, which will get us scrolling for free, and then force the update instead of waiting for the next IRQ — which is disabled anyway to make sure our timing remains precise. (Typing only upper case characters works; lower case shows as symbols.) This program runs at a sedate 300bps, and the reason is because the VZ ROMs are written for space efficiency, not time efficiency, at least to any extent they're efficient at all. 300 baud gives us an apparent surfeit of cycles using our formula — 11931 cycles per bit — but we may well need all of them since we've really got no idea how long it can take the ROM routines to do any arbitrary screen update. We won't be using the ROM much for our data blaster program, but a general purpose terminal emulator would have to consider a proper solution to achieve faster speeds. This is something else we might revisit in a future article. The other purpose of this ASCII test program is to mock up how we'll write the fast serial loader. Despite the fact we have over 10,000 cycles between bits at 300bps and could easily have written each bit we read as a subroutine call to save memory, I still inlined each clocked-in bit using a macro because we necessarily need to at 57.6kbps — among other things, each CALL is 17 cycles and the RET to return from it is 10, which would consume almost half our CPU budget by themselves. I'd also like to observe, again with my usual biases showing, that cycle counting isn't nearly as much fun on the Z80 as it is on the 6502. Most opcode tables will fortunately collapse the whole Z80 T-state and M-state business into a single unified cycle count, but unlike the 6502 where there are instructions with execution times of 2, 3, 4, 5, 6 or 7 cycles (so you can easily make a busywait from any combination), the Z80's cycle time options start at 4 and go as high as 23, skipping many numbers, and many of the smaller cycle times require specific conditions like not taking a branch. Having considered our little half-terminal program, here's what I settled on for 57.6kbps, written as AS macros: From our previous maximal case I turned the in b,NN instruction into a slightly slower in b,(c), which burns an additional cycle, but means we have 38 cycles left over of our 62 which we can split exactly between two ld (ix+N),b instructions of 19 cycles each. (This also lets us possibly alternate between multiple connected serial devices by changing C, but one catastrophe at a time, I always say.) The separate top and bottom waits are for situations where we have an odd number of cycles left over and need to have different wait times; consider this future expansion for the VZ-300. When the stop bit arrives, we need to dump the byte into a buffer and get ready for the next one in the same 62 cycles, since we expect the other end will be ready to fire the next start bit at us immediately. To make an interesting and vaguely useful display (as well as not requiring additional memory), the screen itself would seem like a good place, but this also imposes some constraints: we only have 512 bytes there (i.e., 32x16), some of which we also need for indicating status, meaning our received packets should really be no larger than 256 or 384 bytes to allow for a transmission log and other useful info. The protocol we select should have packets no larger than that, be easy to implement (because I'm lazy), and be something that pretty much everything can speak. While we've seen Xmodem-1K or Xmodem-CRC implemented other places (like The Newsroom's Wire Service variant), I just decided to go with good old O.G. Xmodem. That contains 132-byte packets and is easy to write and checksum, and any errors over USB between your host computer and the VZ200 would undoubtedly be from bad bit framing rather than line noise which the default checksum algorithm should detect. While it overruns memory a bit at the end, this is largely irrelevant for just loading something we intend to immediately execute. All that preamble yields us a stop bit stanza like this: Here we use the IX index register as a pointer into screen memory and the L register as the packet length countdown. A double-store of the same location onscreen once again soaks up 38 cycles, then the increment and decrement, then the branch. A nice thing about the JP instruction, which is absolute instead of relative, is that the conditional branch form requires 10 cycles regardless of whether it's taken or not, so this entire stanza always consumes precisely 62 cycles as well. We use that instruction a lot in the cycle-exact portions so that we always have predictable CPU time. Once we get a full packet, we know the sender won't do anything until we reply, so we can relax our timing and validate the packet at leisure, copy it into the correct place in memory and send the ACK for the next one. We send bytes using the same send-bit route I showed you before or a trivial variation, padded to 62 cycles per bit and also inlined. We accept .VZ-format files in this loader, so we have special handling for the first packet to make sure it has a generally correct format and note the type and memory address, which is where the rest of this packet and subsequent packets will be copied to. If it doesn't, then we send CAN and force the sender to abort. By contrast, the start bit is handled with the same 25-cycle code I showed you before, because this is the fastest way we can be sure we won't miss one. But this also means we have no way of checking the keyboard nor implementing a timeout: there is no spare time to count cycles or scan for keys, and the only free-running timer is the VDG end-of-frame IRQ which will totally mess up our timing if that runs, so during the entire transaction IRQs are disabled as well. The program therefore assumes your sender is up and ready to go the moment it starts executing. On startup it fires off the initial NAK and waits, possibly forever if the other end never gets the signal until you reset the VZ200. (We display a message to alert you that no transmission has yet been received, which is immediately overwritten on-screen by the .VZ metadata.) Let's see our loader in action. On the host side, to send the program to the VZ200 you can use any terminal program that speaks Xmodem (and just about everything does), just as long as it automatically starts the transfer as soon as the initial NAK arrives. With both my MacBook Air laptop and my Raptor Talos II workstation, I use lsx with the usb2ppp tool from BURLAP, which we earlier used to tunnel PPP over a serial line for the Brother GeoBook, but can be used to run pretty much any program over a serial port. Here's an example, substituting the path to the HW-597 that your machine uses (e.g., Fedora Linux on my Raptor Talos II uses /dev/ttyUSB0): We start the loader, which you can run as a separate program on its own and it will load anything that does not encroach on its default location at $e000. (We'll find an even better spot for it in just a minute.) Since our ASCII half-terminal program loads and executes from $8000, this is no problem; it takes up three blocks (there's apparently an off-by-one bug in lsx), so it's a good quick test of the machinery. The loader immediately sends NAK and we're off to the races. RUN you can just hit RETURN on to run a BASIC program. Ta-daaaaa! The Australian, paywalled link). Now we want to make it part of the system. To convert it to "firmware" takes advantage of a specific feature Ben built into the SD card reader for easier updates: the ability to load and run a new system "ROM" directly from the card. The default memory map for the VZ200 puts the system ROMs between $0000 and $3fff, reserving the space from $4000 to $67ff for ROM cartridges, though a cartridge could technically take over any address range above the TOM (that's how the RAM expanders worked, after all), and we'll come back to that point later. For the system to recognize code at $4000 or $6000 as part of a cartridge, a sequence $aa $55 $e7 $18 is required in that order, and execution then starts at $4004 or $6004. As shipped to you Ben's device not only fills in RAM above the TOM, it also fills RAM in from $4000 to $67ff and puts its own code there with that sequence, effectively "slushware" (a la the DECmate II, and we'll use the same term here since it's not really ROM). This code is run by the system ROM on startup like a "cartridge," because that's what it looks like to the system ROM, and this code is what does the RAM test and initializes SD card access. Once the system is up, it then does this: VZDOS.VZ it's loading from the card is "magic." If present, it will be used to temporarily replace the slushware at runtime; for a couple extra seconds spent loading it you don't have to mess around with burning it to the cartridge. But this binary is not signed or checksummed, nor does the load check if it's even a new copy of the slushware — any program will serve as long as the .VZ header loads it to $8000 but the code is written to execute from $4004 (as the onboard image would). The slushware then places a little trampoline copy routine at $a000 and runs that to copy the 8K from $8000 to $9fff to $4000, overwriting the old slushware, and jump into the new one. To make our replacement code useful, we should provide some quality-of-life features. We'll make it autostart into a transfer so that all you have to do is load up the program into your Xmodem sender and reset the VZ200, and after a polite delay it will pull down and run the program automatically. We'll also let you load multiple times if you want instead of immediately trying to execute the current file being transferred. We'll also finally put that routine in memory for the slower but more forgiving memory fill operation, and enable the VZ sticks by default in case we find something that really needs them. But more important than those, we should also let you drop back into BASIC and use the SD card loader normally without having to pop the card out. That requires us to include a copy of the actual VZDOS.VZ which we will embed in our replacement code. This adds an additional complication, because the 2.32 slushware (the most current as of this writing) is already 8068 bytes long minus the .VZ header, leaving us only a little over 2K for our own code. Moreover, if we're over 8192 bytes (and it's inevitable we will be), the loading process will overwrite at least 100 bytes of our code with the trampoline and fail to copy the rest. We'll solve this by immediately copying the remainder as the first step in our binary, and then post-processing the object to yield a new VZDOS.VZ with a 128 byte hole between the first 8K and the last 2K (remember it gets loaded to $8000, so we have plenty of space there). Part of this code will be used to make a jump table entry for our slower fill so the call will stay constant with future updates, if any. That looks like this: Since we are embedding VZDOS but we need to keep all its relative offsets intact, we skip the first 7 bytes and the .VZ header, and run those instructions later just before we jump back into it (if we do). We also do a couple patches so that VZDOS will reload (us) on a reset, but not when we execute the embedded copy, and still use an unmodified 2.32 so that you can see there's nothing up my sleeve. Time for our fill routine. 7001, etc. dec bc ld a,c or b ld a,(slobyte) ; flags kept jr nz,sloclrl ret not in VRAM, then do everything LDIR would and leave the routine with A, HL, DE and BC set as they would be at the end. (We do set the Z flag on exit, but most routines won't care about this.) Then our Invaders example, which was can be patched by overwriting the two instructions at $8caa with ld a,0:call 04008h. xor a and could just nop it.) (I couldn't cursorily find out more about this person), who also did Invaders and a number of other software releases for DSE under contract. except scrolling the attract-mode screen, because unavoidably we'll scroll up the stuck bit. I don't think there's a good general way around that. Fortunately it's purely cosmetic, and most games I converted in this fashion seemed to work fine without any glitches. WORDPRO.VZ can also be placed on the card and run from there as we did previously, though doing so doesn't enable file operations either.) You'll need copies of the actual ROMs, which do circulate. The entirety of the code — really a disguised linker script — looks like this: Remember that Wordpro was first and foremost written for the VZ-300, which has 16K of RAM, so its TOM is much higher ($b7ff). The cartridge, because it has full control of the bus, thus maps its much larger ROM in at $d000-$ffff, with 2K of $d000 also mapped to $6000 with the cartridge header sequence. This echo of the main cartridge ROM is what actually autostarts everything since the system ROM doesn't know to check anywhere else but $4000 and $6000. The same scheme works for the VZ-200, except there is no RAM between $9000 and $bfff. But with the BennVenn cartridge, we have RAM everywhere, so we load to $8000 and copy the Wordpro ROM dumps upon execution to their proper location(s), duplicating $d000-$d7ff to $6000-$67ff like a real one, and jump into the "cartridge" at $6004. This copy operation will destroy BFL-the-slushware, but we can just reset to reload it. Wordpro will get all the RAM it would expect to get on a VZ-300, even on this 4K VZ200. too efficient and constructed a general fill subroutine which lots of things called, the screen clear portion being only one of many. (Darn those efficient little assembly language programmers.) we are closed now!"), one of my favourite Intellivision titles, given a solid conversion as Hamburger Sam also by D&M. Z88 Development Kit. Arkaball by Jason Oakley, an obvious Arkanoid clone, but competently crafted and worth a video. VZ-DOOM, a Claude-written raycast Wolfenstein 3D-style game. It fortunately didn't need patching and "just worked." Your mileage may vary as to whether you think the AI usage is cheating, and this blog has a strict no-AI-article-text policy, but it performs as advertised on this NTSC machine and likewise merits a video. VZ-DOOM uses WASD for motion, comma and period to strafe, E to open and SPACE to shoot. VTech items continued to show up in Dick Smith stores, and VTech did make other Laser computers, these "later Lasers" were not closely related nor compatible with the VZ line, or each other, and most of them were much less successful — with the exception of their Apple II clones. The Laser 3000, fresh from its Summer CES 1983 debut, also made it down under to Dick Smith stores as the Dick Smith Cat. It required real Apple II ROMs in an external cartridge for compatibility, which also made it a target for Apple, fresh off their successful victory against Franklin Computer for using substantial portions of the Apple II ROM in the Franklin Ace 1000. Although VTech was still able to sell it elsewhere and the Apple II ROMs were never integrated into the base machine, Apple instead argued that the mere use of the ROMs was an infringement upon its intellectual property regardless, and successfully blocked further imports to the United States. VTech learned from this just like they learned from EACA, and developed new ROMs that were carefully clean-room reverse-engineered, additionally incorporating a licensed copy of Microsoft BASIC retrofitted to act like Applesoft BASIC. This process provided the legal assurance that not a nybble of Apple's code was even consulted, and VTech used these unencumbered ROMs to create what was introduced at Summer CES 1985 as a redesigned "90% compatible" (per Creative Computing) Laser 3000. In turn, that reworked Laser 3000 was transformed into the 1986 Laser 128, a semi-portable riff on the Apple IIe with an expansion slot and built-in 5.25" floppy (later 3.5"). VTech's work paid off handsomely: reviewers were impressed by its value for money, most software never noticed the difference, and Apple's repeated attempts to prevent its importation and sale all ended in failure. The Laser 128 family's low price and exceptional built-in functionality made them the most widely sold Apple II clones in the United States, finishing on the market as late as 1989 in an upgraded 3.6MHz version with a 3.5" disk drive and over 1MB of RAM. returned to journalism in 1987. Meanwhile, DSE was still selling enough VZ-300s to keep them in this 1987-88 catalogue, and according to Greg Dubois' contacts at Dick Smith even at that late date they continued moving over 100,000 units a year, but in the end it was actually VTech that wanted out: VTech wanted to redirect factory capacity to the Laser PC and wasn't willing to keep producing the older system, and even DSE's offer to double the order couldn't convince them otherwise. Although some software and accessories still appeared in the 1989-90 catalogue, the computer itself did not, and the line disappeared completely by 1991. For a time afterwards DSE sold IBM consumer PCs and even Commodore PC clones in the process of adopting its own DSX PC brand. In the 2000s the company failed to make a successful transition from its mail order origins to the new world of online sales, and despite several attempts to rework its retail presence Woolworths unloaded Dick Smith to Anchorage Capital Partners in 2012 in a controversial deal where much of the sale price was allegedly financed by Anchorage liquidating DSE's own assets. Anchorage took the company public in 2013, netting tens of millions, but the company did not recover and all 363 remaining stores were closed by May 2016. Today the Dick Smith brand lives on solely as a mark of online retailer Kogan, primarily selling consumer electronics. Australia's crap home computer, by golly. Much as Sinclair did in the UK and Commodore in the United States, the Aussie VZ-200 and its successor VZ-300 remain as beloved as they are because they introduced a entire generation of Australians to computers who could never buy one before. While a few importers tried to bring the also-rans to the South Pacific (DSE even had Radofin's undead zombie Aquarius in their 1985 catalogue!), the VZ was there first and in large numbers, over 20,000 VZ-200s alone, becoming the down-under standard against which all subsequent cheapo systems were measured. Indeed, when the desperately dire Tandy MC-10 got in front of Australian Personal Computer in December 1983, reviewer Surya commented that when considering it versus the VZ-200 "the MC-10 does not stand up well to this comparison." Legions of user groups and newsletters sprung up to support it, tinkerers designed all manner of expansions for it, and users wrote and sold their own software for it, ironically spawning exactly the sort of hobbyist-driven computer ecosystem post-Dick Smith that Dick Smith-era Dick Smith had previously tried to foster. Ultimately the little Hong Kong desk wedge became more of an Australian computer icon than even some truly homegrown ones. As for this NTSC VZ200, it should be very possible to clone it because the ROMs are the same as the better-known DSE flavour and it's otherwise all off-the-shelf hardware; moreover, it would be infinitely easier to maintain and repair than the ghastly PCB it's got now. The schematics for the PAL Dick Smiths are widely available and there are even fewer components needed to build an NTSC one. At least one person has already made an NTSC-compatible RC2014 workalike, though that project uses a GAL, and it seems like we could make a more straightforward knockoff just using what the original did — with the exception of the colour encoder, which would be improved and somewhat simplified by using a proper MC1372 instead of the TBA520. I know "Leaded Solder" Mike has his clone CreatiVision, so I look forward to him picking this up as a new challenge. ;) We'll be doing more with this system and particularly the VZ-300, now safely awaiting my next Southern Hemisphere trip, in future articles — along with a recently-acquired PAL CreatiVision of our own, the basis for the Dick Smith Wizzard, which we need to see if we can get up and running (I do like me a 6502 and a 9918). Meanwhile, David "Bushy" Maunder's VZ website can give you all the articles, technical information and software that you can stick on an SD card. are all on Github, including pre-built binaries and ready-to-go Bush Food Loader "slushware" you can use with your own SD card loader, all of which are under the BSD 2-clause license.

3 days ago 1 votes
Anecdotally, programmers dislike "reduce"

In short: from my experience, people like map and filter, but not reduce. I use functions like map and filter all the time. When I put that code up for review, my peers rarely complain. I get plenty of feedback about other decisions, but not about my use of map and filter. I cannot say the same for reduce. Often, when I’ve submitted a patch with reduce inside, I get a comment like, “this part is hard to read.” And I see reduce way less than map, filter, some, and so on. Anecdotally, I have come to believe that programmers don’t like reduce as much. I don’t know why, but I have a few theories: reduce is harder to read. reduce is less familiar. reduce can have worse performance compared to other options. reduce is less elegant in languages I use, like JavaScript, Python, and Swift. In my blissful stint as a Clojure developer, I did not get this feedback. I’m wrong, and I’m seeing a trend that’s not real. I usually just change reduce to something else and move on. Even though I prefer it, I don’t usually care much. But it’s a little social phenomenon I’ve observed, and I thought I’d document it. I’ve also noticed this less recently, possibly because code review is less thorough nowadays. Do you notice this? Do you like reduce? Please tell me.

3 days ago 1 votes
Microcode in Intel's 8087 floating-point chip: the scale instruction

In the 1970s, floating-point arithmetic was a mess. Computer manufacturers had a dozen incompatible arithmetic standards. Moreover, floating-point systems were designed around hardware simplicity rather than mathematical rigor, leading to problems with numerical stability. This changed when Intel introduced the 8087 floating-point coprocessor chip in 1980, designed to be as accurate as possible, even in the corner cases. The 8087 became popular because it could be installed in the IBM PC, making floating-point operations up to 100 times faster in applications ranging from spreadsheets to CAD. But more importantly, the 8087 became the floating-point standard used by most computers today. The 8087 implemented its instructions in complex low-level code called microcode. I'm part of a group, the Opcode Collective, that is reverse-engineering this microcode, and I've recently made some progress. In this post, I examine the microcode for one of the 8087's instructions—FSCALE—and describe how this microcode works. The FSCALE (Floating-point Scale) instruction provides a quick way to scale a number by a power of two, much faster than a multiplication. I figured that FSCALE was a simple, almost trivial instruction that would be straightforward to understand and explain. Spoiler: it is not simple. FSCALE uses over 140 micro-instructions and three levels of subroutine calls to handle many special cases. But the FSCALE microcode illustrates many interesting parts of the 8087, such as the shifter, the adder, and the exponent converter, and also reveals a hidden feature of the 8087, so hopefully you will find it interesting. To explore the microcode, I opened up an 8087 chip and created a high-resolution image with a microscope. The large microcode ROM is in the center, holding the 1648 micro-instructions that control the chip. The microcode engine on the left steps through the microcode, handling jumps and subroutine calls. The bottom half of the chip is the "datapath", the circuitry that performs floating-point calculations; it is split into a 16-bit datapath for the number's exponent and a 64-bit datapath for the number's significand (also known as the fractional part). Die of the Intel 8087 floating-point unit chip, with main functional blocks labeled. The die is 5mm×6mm. Click for a larger image. Zooming in on the bottom part of the chip shows the datapath circuitry; I've highlighted the relevant parts below.1 The exponent ROM holds various constants. The exponent converter is a specialized circuit that examines exponents, detects special values, and converts between exponent formats.2 The shifter is a large component; it allows a 64-bit3 value to be shifted left or right by arbitrary amounts. (I wrote about the 8087's shifter circuitry here.) The adder is the heart of the 8087's calculations; it is used in a loop for multiplication, division, and square roots. The B register holds one input to the adder, while multiple sources can provide the other input. The sum register holds the adder's output. The eight stack registers and the temporary registers hold floating-point numbers. A close-up of the 8087's datapath, showing functional blocks that are used by FSCALE. Details of the 8087 In this section, I'll explain some features of the 8087 that are important for the FSCALE microcode. To use the 8087, a programmer stores values in its eight internal registers, organized as a stack. Each register holds an 80-bit floating-point number. To optimize performance, each value in the register stack has an associated "tag" value, which is mostly invisible to the programmer.4 A tag labels a value as valid, special, zero, or empty. A "normal" floating-point value is tagged as valid. If the floating-point value is infinity, Not a Number (NaN), or a denormalized value, then it is tagged as special. A zero value is tagged as zero. Finally, if a register is empty (e.g., its value has been popped off the stack), the register is tagged as empty. The 8087 also has temporary registers that it uses internally: tmpA, tmpB, and tmpC. Like the stack registers, tmpA and tmpB are 80-bit registers, along with two tag bits. However, tmpC only holds a 64-bit significand. The 8087 supports a variety of data types: floating-point numbers of various sizes, integers, and binary-coded decimal. But internally, everything is stored as an 80-bit floating-point number called a "temporary real"; for the rest of this article, I'll only be considering temporary real values. A number has three parts: the sign bit, the 15-bit exponent, and the 64-bit significand (the fractional part), In most cases, a floating-point number is represented by sign × significand × 2exponent. The significand is a 64-bit binary number of the form 1.bbb...: a leading 1, followed by the binary point (the binary equivalent of the decimal point) and the rest of the bits.5 What makes floating-point numbers useful is that their scope covers the incredibly small to the astronomically large, thanks to the exponent, which ranges from -16382 to 16383. One important detail is that the exponent is stored with a "bias" of 16383 added to it. Thus, the stored exponent is always positive, even if the real exponent is negative.6 The 80-bit temporary real format. The triangle indicates the binary point, analogous to the decimal point. From the Intel Numerics Supplement. The 8087 supports several types of numbers that are represented as special cases with special exponents, as shown below. Zero and infinity have both positive and negative values. "Not a Number" (NaN) represents values that don't make sense, such as 0/0 or sqrt(-1); NaN has a large number of representations, not a single value. The 8087 also supports denormalized and unnormalized values, which are extremely small values where the significand doesn't have a leading 1. The encoding of special values. Based on Table S-31 in the Intel Numerics Supplement, but highly simplified. The "x" bits are arbitrary, as long as they don't conflict with another type. The 8087 has a complicated exception system with six types of exceptions to indicate if something went wrong with an arithmetic operation. The most serious is the "invalid operation", indicating that the operation does not make sense, such as 0/0 or ∞-∞. It also includes accesses to an empty register (stack overflow or underflow) or operations on a NaN value. The 8087 also has an overflow exception if a value is too large to store, an underflow exception if a value is too small, and a divide-by-zero exception (excluding 0/0). A denormalized operand exception indicates that the result is too small to store as a normal value, but can be stored as a denormalized value. Finally, a precision exception indicates that a value cannot be represented exactly and must be rounded. (Precision exceptions are very common; even 1/10 will yield one.) The 8087 provides fine-grain control over each exception type, specified by bits in the control register. If an exception is unmasked, the 8087 sends an interrupt to the 8086 processor, which handles the problem in software, for instance by terminating the program or logging an error. Alternatively, the exception can be masked and the 8087 will continue execution as best it can. For instance, an invalid result will be replaced by NaN, while an overflow or divide-by-zero will be replaced by infinity. A precision exception will result in rounding. The point of masked exceptions is that calculations continue, yielding an answer that is as accurate as possible; in most cases, this is what the programmer wants. These features make the 8087 flexible and provide accuracy, but they also make the microcode much more complicated, since the combinations of special cases need to be handled appropriately. The 8087's microcode Executing an 8087 instruction can require hundreds of internal steps to compute the result. These steps are implemented in microcode with micro-instructions that specify each step of the algorithm. (Keep in mind the two levels of instructions: the assembly language instructions used by a programmer and the undocumented low-level micro-instructions inside the chip.) The microcode ROM holds the 1648 micro-instructions that implement the 8087's instruction set. I'm working with the Opcode Collective to reverse-engineer the micro-instructions and fully understand the microcode (link). The 8087's micro-instructions are complicated, with many corner cases and ad hoc functions, but I'll provide a simplified overview. Each micro-instruction consists of 16 bits, as shown below. The first three bits specify the micro-instruction's type, which controls the meaning of the remaining bits. The first type is a transfer operation, which transfers data from one internal register to another. The two fields specify the source and destination. The three remaining bits are used for various special cases. Next is a shift operation, which uses the barrel shifter to shift a value left or right. The third type of micro-instruction controls the adder (which can also subtract). The miscellaneous instructions include stack pointer operations, tag modification, exceptions, and subroutine return. The far jump and far call micro-instructions perform a jump or subroutine call to a target micro-address in a fixed list. The condition field allows conditional jumps/calls/returns based on numerous conditions, while the last bit inverts the condition. A local jump is a relative jump to a nearby micro-instruction. Structure of an 8087 micro-instruction. The FSCALE microcode When the 8087 starts executing an instruction, the instruction decoder circuitry determines the starting address of the microcode corresponding to the instruction. This 11-bit address is loaded into the microcode engine, which starts executing the microcode.7 The microcode for FSCALE (shown below) starts at decimal address 748.8 The idea behind FSCALE is straightforward: if you want to scale a floating-point number by 2N (for an integer N), you add N to the number's exponent. This allows you to multiply or divide by a power of two much faster than using the full floating-point multiplication operation. However, the microcode for FSCALE is unexpectedly complicated and uses several microcode subroutines. In brief, the microcode first checks for arguments that are zero and then handles other special arguments. It converts the scale argument to an integer and adds it to the exponent. Finally, it handles any overflow or underflow. In more detail, the microcode routine starts by moving the first argument from the top of the stack (st(0)) to the tmpA temporary register. If the argument is zero, the routine immediately returns. (Thus, scaling 0 by anything—even NaN—will give a result of 0.) Next, the second value on the stack (the second argument) is moved to the tmpB temporary register. Likewise, the code returns if this value is 0, so scaling anything by 0 leaves the value unchanged.9 Next, a constant value is selected; selecting a constant and using it are two separate micro-instructions. (The 8087 has separate ROMs for 16-bit exponent constants and 67-bit significand constants; this one is an exponent constant.) In the normal case, execution jumps to address #0763, skipping the call to subroutine SPECIAL_TMPS. FSCALE: #0748 st(0) -> tmpA Input argument from top of stack #0749 jmp #0776 if tmpA:tag ZERO Bail if 0 #0750 stackPtr++ #0751 st(0) -> tmpB Scale argument from stack(1) #0752 stackPtr-- #0753 jmp #0776 if tmpB:tag ZERO Bail if 0 #0754 expconst 0x403e Const 403e: exp shift to convert to int #0755 jmp #0763 if not tmp empty/special/div #0756 call SPECIAL_TMPS Special handling #0757 jmp #0762 if flag #0758 jmp #0761 if not tmpB:tag SPECIAL #0759 except:invalid Invalid exception, use NaN #0760 NaN -> tmpA #0761 jmp #0776 if intr #0762 jmp #0775 if expConv[0] Return tmpA if expConv set, otherwise continue #0763 tmpB:exp -> Breg Normal path #0764 tmpB:sign,exp -> expConv ExpConv will test tmpB's sign #0765 expConst -> tmpC Const 403e #0766 adder: tmpC - Breg cin=1 403e-exp is amount to shift to convert tmpB to int #0767 sumreg:frac -> shiftcount Store in shifter control #0768 shift tmpB:frac R count byte bit Perform the shift #0769 shift R -> Breg Breg holds scale argument as an int #0770 jmp #0777 if neg Negative Breg needs separate handling #0771 adder: tmpA:exp + Breg cin=0 Add the scale to the exponent #0772 sumreg:frac -> expConv Put result in expConv to check #0773 sumreg:frac -> tmpA:exp Update exponent with sum #0774 call NONNORMAL_RESULT if not exp normal Handle overflow/underflow #0775 tmpA -> st(0) Save result back to stack #0776 RNI Done: Run Next Instruction #0777 adder: tmpA:exp - Breg cin=1 Subtract Breg #0778 jmp #0772 Continue processing Continuing at #0763, the second argument is converted from a float to an integer, which takes a few steps. For example, suppose the argument is 9, which in floating point is 1.001×23. The significand bits 1000 are "left justified", but for an integer, these bits need to be "right justified" by shifting them to the right. In general, if the exponent is n, the significand is shifted right by 63-n bits. But recall that the exponent is biased by 16383. Thus, the significand must be shifted right by 63-(exp-16383) bits, that is 0x403e-exp bits. (This explains the constant 0x403e earlier in the microcode.) Converting a float to an int by shifting. In the microcode, the subtraction takes several steps. At #0763, the exponent of the second argument is moved to the B register, one of the inputs to the adder (completely different from tmpB).10 Next, the sign and exponent are moved to the exponent converter, a circuit that, among other things, tests for overflow. Next, the constant 0x403e (selected back at #0754) is moved to the tmpC register. At #0766, the adder is activated, subtracting the exponent from the constant.11 The adder puts the result into the sum register, and this value is copied to the shift count register, which controls the shifter. This value indicates how many bits the second argument must be shifted to convert it to an integer. At #0768, the shifter is activated to shift by the desired amount, using both the bit shift part and the byte shift part. As with the adder, activating the shifter and reading the result are separate micro-instructions; the result is put into the B register. The core part of the FSCALE instruction is finally performed at #0771, adding the second argument to the first argument's exponent. The adder is activated to add the B register value (the scale) to the exponent, and the updated value is stored in tmpA's exponent. (Except if the scale factor is negative, it is subtracted via the #0777 path.)12 The value is also sent to the exponent converter circuit, which checks the exponent for overflow or underflow; if so, subroutine NONNORMAL_RESULT is called. But in the normal case, the updated value is copied from tmpA to the top-of-stack register st(0). Finally, RNI (Run Next Instruction) indicates that the microcode routine is done and the instruction is completed. Thus, even in the straightforward case, FSCALE takes about 22 micro-instructions. Handling empty or special arguments What happens if an argument accesses an empty stack location (i.e. stack underflow) or is a special value (infinity, denorm, NaN)? These cases are handled by a micro-subroutine that I'll call SPECIAL_TMPS15 because it processes special values in tmpA and/or tmpB. This subroutine is a general-purpose routine, used by basic arithmetic operations, FSCALE, FTST (test), and FPREM (partial remainder). The control flow through SPECIAL_TMPS is rather convoluted since the code must prioritize issues if, say, one argument is empty and the other is a denorm. I'll just give a brief summary; see the footnote13 for details. First, the subroutine converts any denorms to unnorms. Then it checks for access to empty stack locations, raising an exception or interrupt if so. Then it checks the two arguments again. If either is NaN, an exception or interrupt is triggered. Otherwise, it returns a status indicating the type of arguments. Unexpectedly, if both arguments are NaN, the code compares the two NaN values and returns the larger. This behavior may seem very weird, but it's a documented feature.14 You might think that NaN is a single value, but it's actually an enormous family of values. The idea was that the programmer could use different NaN values to signal where a problem occurs. For instance, you could put a different NaN in each location of an uninitialized array, so you could tell which position was accessed. For some reason, the designers of the 8087 decided that if you perform an operation with two different NaNs, the result is the larger one. Thus, the microcode needs code that detects if both operands are NaN and computes the larger, using a subtraction for the comparison (#1518). SPECIAL_TMPS (J5): #1484 call SPECIAL_VAL if tmpA:tag SPECIAL Handle special values in tmpA/tmpB #1485 xchg tmp #1486 call SPECIAL_VAL if tmpA:tag SPECIAL Handle tmpB special #1487 xchg tmp #1488 1 -> flag Flag=1 by default #1489 jmp #1500 if not tmp empty/special/div 0 -> expConv if tmps okay #1490 1 -> expConv #1491 jmp #1497 if not tmpA/B empty #1492 except:invalid Invalid if either empty #1493 jmp #1525 if compare instruction No NaN for comparison #1494 jmp #1511 if intr Return if interrupt not masked #1495 NaN -> tmpA NaN if interrupt masked #1496 return #1497 jmp #1502 if tmpA:tag SPECIAL Special cases #1498 jmp #1505 if tmpB:tag SPECIAL #1499 0 -> flag Div normal path: #1500 zero -> expConv Return flag 0, expConv 0 #1501 return #1502 call SPECIAL_VAL TmpA special #1503 jmp #1512 if not flag Jump if NaN, fallthrough if infinity #1504 jmp #1509 if not tmpB:tag SPECIAL #1505 xchg tmp TmpB special #1506 call SPECIAL_VAL #1507 xchg tmp #1508 jmp #1521 if not flag Jump if NaN, return if infinity #1509 0 -> flag Clear flag, return #1510 return #1511 RNI End instruction with interrupt #1512 jmp #1522 if not tmpB:tag SPECIAL TmpA NaN, now check tmpB #1513 xchg tmp #1514 call SPECIAL_VAL Check tmpB #1515 xchg tmp #1516 jmp #1522 if flag Jump if tmpB is not NaN #1517 except:invalid Invalid exception #1518 tmpB:frac -> Breg Both args are NaN, find larger #1519 adder: tmpA:frac - Breg cin=1 #1520 jmp #1522 if adder sign See if tmpA #1521 tmpB -> tmpA Take larger #1522 except:invalid Invalid exception #1523 jmp #1525 if compare instruction No interrupt for comparison instruction #1524 jmp #1511 if intr End instruction with interrupt #1525 1 -> flag Return with flag set #1526 return End of J5 This subroutine makes heavy use of a helper subroutine, SPECIAL_VAL,16 that processes one argument. The helper converts a denormalized argument to an unnormalized argument, raising an exception or interrupt as appropriate. It also flags an input of infinity. The hardware for the micro-instruction that exchanges tmpA and tmpB at #1485 is interesting. Instead of physically moving the values between the two registers, the micro-instruction toggles a flip-flop that exchanges the meaning of tmpA and tmpB. That is, if the flip-flop is set, a reference to tmpA goes to tmpB and vice versa. (This is a standard trick in microprocessors; the Intel 8080's XCHG instruction exchanges the DE and HL registers in a similar way. The Z80 uses the same trick for the EX and EXX instructions to exchange the regular register set with the secondary register set.) The Intel 8087 chip is packaged in a 40-pin DIP (dual in-line package), as are the 8080 and Z80. This photo is here as a break from all the microcode. Handling a non-normal result If you take a very large number and scale it larger, you can end up with overflow. If you take a very small number and scale it smaller, you can end up with a denormalized number or underflow. This will trigger an overflow, denorm, or underflow excaption, and an interrupt if unmasked. Moreover, the 8087 supports four rounding modes: round to nearest valid value, round down (toward -∞), round up (toward +∞), or round (chop) toward zero. Depending on the rounding mode, an overflow can result in either ∞ or the largest possible floating-point number. Similarly, an underflow can result in either zero or the smallest possible floating-point number. And depending on the infinity mode (affine or projective), infinity can be either signed or unsigned. Thus, the FSCALE microcode needs to handle many special cases for the result. The subroutine to handle a non-normal result in tmpA is below. One interesting micro-instruction is update overflow/underflow exceptions, which triggers an exception if appropriate. For most exceptions, a micro-instruction triggers the exception (for example, except:precision at #0346). But for the overflow and underflow exceptions, the microcode delegates the task to hardware. Specifically, the 8087's "exponent converter" circuit examines the exponent to see if an overflow or underflow exists, based on the selected floating-point precision. The micro-instruction sets the overflow and underflow flags based on these values. Thus, a complex task is performed by a single microcode instruction, thanks to the hardware support of the exponent converter. NONNORMAL_RESULT (J16): #0318 return if tmpA:tag ZERO Handle non-normal result #0319 update overflow/underflow exceptions Trigger exceptions if exp conv says to #0320 expconst 0x6000 The interrupt bias constant 0x6000 #0321 jmp #0329 if not intr #0322 expConst -> Breg Interrupt path #0323 jmp #0326 if neg #0324 adder: tmpA:exp + Breg cin=0 Add bias for underflow #0325 jmp #0327 #0326 adder: tmpA:exp - Breg cin=1 Subtract for bias overflow #0327 sumreg:frac -> tmpA:exp New exponent to tmpA #0328 return Interrupt, so done #0329 jmp #0344 if neg Masked exception #0330 tmpA:exp -> Breg Underflow #0331 adder: 1 - Breg cin=1 Amount to shift denormal #0332 call CREATE_DENORM Create a denormal #0333 adder: zero + Breg cin=0, roundmode Add zero to round #0334 call ADJUST_PRECISION Adjust to specified precision #0335 jmp #0340 if Sum register is zero If zero, return +/- zero as appropriate #0336 zero -> tmpA:exp Denorm: exponent is 0 #0337 sumreg:frac -> tmpA:frac Save denorm fraction #0338 special -> tmpA tag Tag denom as special #0339 return #0340 tmpA sign -> sign latch Return +/- zero #0341 zero -> tmpA #0342 sign latch -> tmpA sign #0343 return #0344 NaN/Inf -> tmpA:exp Overflow: maybe return infinity #0345 tmpA:frac -> tmpB:frac Save tmpA frac in tmpB #0346 except:precision Set precision exception #0347 Inf -> tmpA:frac Put infinity in frac #0348 special -> tmpA tag Mark infinity as special #0349 return if not round chop If rounding up, return infinity #0350 1 -> Breg Return max float: adjust down #0351 adder: tmpA:exp - Breg cin=1 #0352 sumreg:frac -> tmpA:exp Exp=7fff-1=7ffe #0353 adder: zero - Breg cin=1 #0354 sumreg:frac -> tmpA:frac Frac 0-1 = ff...ff #0355 norm -> tmpA tag Normal value #0356 return if tmpB:frac[63] Return max float unless unnorm #0357 tmpB:frac -> tmpA:frac Return original tmpA frac #0358 return The 8087 has interesting behavior if an overflow or underflow is unmasked and an interrupt occurs. The idea is to let the interrupt handler know what the exponent should have been. However, the proper value can't be used since it is too big or too small to fit in the exponent field (which is why the exception occurred). The solution is to add or subtract the constant 0x6000, resulting in an exponent that fits. The interrupt handler can subtract or add this constant to get the correct exponent. Lines #0322 to 0328 perform this addition or subtraction. For a masked underflow, a denorm value is created by the subroutine CREATE_DENORM. The value is rounded to the specified precision by ADJUST_PRECISION. Finally, if the value is too small for a denorm, the value +0 or -0 is returned as appropriate. For a masked overflow, the 8087 either returns Infinity or the largest-possible float, depending on the specified rounding mode. Infinity is represented by an exponent of all 1s, and a significand of 1000...; these values are loaded directly onto the bus by transistors. The maximum float, however, is computed: 1 is subtracted from the infinity exponent, and 1 is subtracted from a zero significand. Helper subroutine: creating a denormal One controversial feature of the 8087 is denormals, numbers that are smaller than "regular" floats. Recall that floating-point numbers have a significand with the first bit set to 1. But what happens if you hit the smallest possible exponent and want an even smaller number? The 8087 lets you break the rule that the significand starts with 1, producing smaller numbers known as denormalized numbers or denorms. Denorms significantly extend the range, providing numbers up to a factor of 263 smaller. However, denorms don't have as much precision since the upper bits are "wasted". Moreover, calculations with denorms can be substantially slower because special handling is required. Example of a normal number, reduced by a factor of 8, resulting in a denormal. The diagram above shows a normal number with the minimum possible exponent (-16382, which is 1 after biasing). Dividing the number by 8 (or scaling by -3) creates a denorm since the exponent can't be reduced any further. Instead, the significand is shifted 3 bits to the right. The exponent is replaced with the special value 0, indicating that the number is a denorm. In the 8087, denorms are created by a microcode subroutine that I'll call CREATE_DENORM; it is used by many arithmetic operations, not just FSCALE. This subroutine takes a normal number and a shift amount. By shifting the normal number (as in the example above), it creates a denormalized number. The microcode (below) uses the exponent converter to check if the shift is 64 or more. If so, there will be nothing left after the shift, so zero is returned. Otherwise, the value is shifted to the right and the denorm is stored in the B register. CREATE_DENORM (J20): #0522 sumreg:frac -> expConv Create denorm #0523 sumreg:frac -> shiftcount Number of bits to shift #0524 jmp #0528 if exponent[6:14] == 0 Jump if #0525 zero -> Breg No bits left, use zero #0526 shift tmpA:frac L 0 bytes, 0 bits Run through shifter? #0527 jmp #0532 #0528 shift tmpA:frac R count byte bit Shift right by the specified amount #0529 shift R -> Breg Result to Breg #0530 shift tmpA:frac L ~count byte bit Now shift back for sticky test #0531 NOP Wait for shifter #0532 rounding(h) -> Breg[grs] Store the three rounding bits in the Breg #0533 return But why is the value then shifted to the left (#0530)? The purpose of this is to get the rounding bits. One of the principles of the 8087 is to get rounding correct, which is a lot harder than it seems. In order to decide how to round up a number, you need to keep track of an impossibly large number of bits. For instance, if you calculate 1 + 0 and round up, you get 1. But if you calculate, say, 1 + 2-10000 and round up, you get a float a bit higher than 1. The problem is how do you distinguish the two sums before rounding, without storing thousands of bits? The trick is that the 8087 keeps three bits for use in rounding: the "guard" bit, the "round" bit, and the "sticky" bit. If you consider a "tail" of bits to the right of the significand, the guard bit is the most significant bit of the tail, followed by the round bit. The sticky bit is special: it is the OR of all the remaining bits in the tail, indicating if any of them are 1. Thus, 1 + 2-10000 has the sticky bit set, while 1 + 0 does not, so the two values can be rounded up differently. To generate the sticky bit, the 8087 uses a very large 64-bit NOR gate that tests the tail bits in parallel. A diagram showing how the guard, round, and sticky bits are computed from a right shift. The numbers in this example are different from the previous example. When a number is shifted to the right (e.g., when creating a denormal), bits are lost off the right. To generate the rounding bits, the value is shifted to the left, keeping all the tail bits that will eventually be discarded, and discarding the bits that will be in the final significand. The top two bits go into the guard and round bits, while the remaining bits are ORed together to generate the sticky bit from the rest.17 The diagram above is an example of this process. Suppose the value is being shifted to the right by 4 bits. The tail bits abcd (or at least d) will get lost in the shift. The rounding bits are computed by shifting the original significand to the right by 59 bits (the complement of 4). Bit 62 (a) becomes the new guard bit, bit 61 (b) becomes the new round bit, and the OR of the remaining 64 bits becomes the new sticky bit. (Note that the old guard, round, and sticky bits get ORed in too, so they aren't lost.) Merging the significand from the first shift with the rounding bits from the second shift produces the desired result. Helper subroutine: adjusting precision Although the 8087 supports three lengths of floats, it performs all calculations with 80-bit "temporary reals". At the end of an instruction, it converts the result to the desired length. (As a consequence, most instructions aren't any faster if you use a shorter float.) A microcode subroutine, which I call ADJUST_PRECISION, converts the result to the precision that is specified in the 8087's control word, using the specified rounding mode. This subroutine is used by most of the arithmetic instructions. The 8087 supports three types of real numbers. From the Intel Numerics Supplement. The first code path handles temporary reals (which have 64 bits of precision). The control word specifies one of four rounding modes. However, there are only two actions that can be taken for a particular significand: either round down (chop) or round up (chop and increment by 1). This decision is made by complicated logic circuits that examine the rounding bits, the rounding mode, and the sign to determine whether to round up or down. This simplifies the microcode but makes the hardware more complicated. The microcode performs a conditional return, returning if the significand doesn't need to be rounded up. Otherwise, the microcode increments the significand by adding 0 with a carry-in. It then checks for overflow, in which case it replaces the value with Infinity and sets a special flag.18 ADJUST_PRECISION (J11): #0299 jmp #0306 if not precision64 #0300 return if not round up, update CC1 Update condition code, maybe return #0301 adder: sumreg:frac + 0 cin=1 Add 1 to round up #0302 return if not sumreg[64] #0303 Inf -> sumreg:frac,sign Return infinity if overflow #0304 2count++ Set special flag #0305 return #0306 23/52 -> shiftcount Short or long real: get appropriate shift #0307 shift sumreg:frac,rnd L count byte bit sticky Shift to generate rounding bits #0308 NOP Wait for shifter to complete #0309 rounding(H) -> sumreg[grs] Store rounding bits #0310 shift sumreg:frac R ~count byte bit Shift right to drop excess bits #0311 shift R -> sumreg:frac #0312 jmp #0314 if not round up, update CC1 Update condition code #0313 adder: sumreg:frac + 0 cin=1 Round up if appropriate #0314 shift sumreg:frac L ~count byte bit Shift left to realign #0315 shift L -> sumreg:frac,sign #0316 return if not sumreg[64] Return if not overflow #0317 jmp #0303 Return infinity The code is more complicated when returning a smaller precision (short real or long real), since the significand must be shortened. First, the code at #0306 loads the shifter with either 23 or 52, depending on the precision specified in the control word, and then shifts the value left. This produces the rounding bits as in the previous section. Next, the value is shifted to the right, shortening it to the desired length. As before, the significand is incremented or not, depending on whether it should be rounded up or not. Finally, the value is shifted back to the left, so the most significant bit of the significand is on the left. As before, if rounding up caused an overflow, infinity is returned. One bizarre feature is that a jump with the "round up" conditional also has a side effect of updating the 8087's programmer-visible condition code register (CC1), indicating if the result was rounded up or down. That is, the 8087 has extra circuitry to detect this specific condition and load the value into the condition code latch. Strangely, the 8087 documentation doesn't describe this condition code action; Intel didn't document it until the 387SX floating-point chip in 1987.19 Conclusions Floating-point has a long history before the 8087. For instance, the IBM System/360 mainframes (1964) supported 32-bit and 64-bit floating-point numbers. In 1977, AMD introduced the Am9511 floating-point chip, supporting 16- and 32-bit floating-point numbers, along with transcendental functions. What made the 8087 revolutionary is that it was carefully designed to be as mathematically accurate as possible, largely thanks to numerical expert William Kahan. (The 8087 led to the IEEE 754 Standard, now used by almost every computer and ending the anarchy of incompatible floating-point standards.) The 8087 ended up extraordinarily complicated with three different sizes of floating-point numbers, four sizes of integers, four rounding modes, infinity modes, a collection of exceptions that could be masked or unmasked, denormalized and unnormalized numbers, signed and unsigned infinities, signed zeros, and a whole family of Not-a-Numbers. These features combine, yielding many corner cases. The 8087 deals with this complexity both through specialized circuits and through tangled microcode. How complicated is the 8087? For users who didn't have an 8087 chip, Intel sold an 8087 Support Library that exactly emulated the 8087's instructions (but much slower). The emulator took 16K bytes of 8086 code, which was a lot when a full BASIC interpreter could fit in 8K. Another way of looking at this is that the hardware of the 8087 drastically reduced the amount of software required: the 8087 itself used 3.3K of microcode, compared to the 16K for the emulator in 8086 code. I plan to continue reverse-engineering the 8087 microcode; for updates, follow me on Bluesky (@righto.com), Mastodon (@[email protected]), or RSS. I've been working on this with the members of the "Opcode Collective", especially Smartest Blob and Gloriouscow, who converted the ROM images to microcode data and extensively analyzed the contents. See the 8087 repository on GitHub for more. Notes and references The 8087 patents provide some details on the hardware, but unfortunately not the microcode. The patent diagram below shows the architecture of the 8087; I've highlighted the relevant parts. The fraction bus and exponent bus are shown in red. The adder and associated registers are in yellow. (For subtraction, the B register selector selects the complement.) The shifter is in green. The exponent constant ROM and the exponent converter are in orange. The temporary registers and stack registers are in blue. The architecture of the 8087. Based on the patent. Click this image (or any other) to magnify.  ↩ The exponent converter is surprisingly complicated because the 8087 has three different formats for floating-point numbers with three different sizes of exponent fields (8 bits, 11 bits, and 15 bits). Moreover, the different sizes of exponents are stored with different biases. Thus, converting between different sizes of exponents is not trivial. The exponent converter also recognizes overflow and underflow for the different exponent sizes, as well as special values such as infinity and NaN. I plan to describe the exponent converter in more detail later. ↩ The significand in the 8087 is nominally 64 bits wide. However, the 8087 uses three extra low-order bits for rounding, called Guard, Round, and Sticky. These bits ensure that a value is always rounded in the right direction. Some parts of the datapath have additional bits for sign or overflow: the shifter is 68 bits wide, and the adder is 69 bits wide. For the most part, I'll ignore these extra bits and refer to the datapath as 64 bits wide. ↩ Tags are normally invisible to the programmer, but can be accessed through special operations. Specifically, a programmer can dump the 8087's state to memory; the tags are stored in a 16-bit "tag word". ↩ The external representations of floating-point numbers have an implied leading one, with only the bits after the binary point explicitly stored. This provides one additional bit of resolution "for free". The internal 80-bit representation, however, has an explicit leading one to simplify calculations. ↩ One reason that the exponents are biased is that to find the larger of two floating-point numbers, you can compare them lexicographically as signed integers, rather than needing to examine the exponents separately. ↩ Most of the 8087's instructions are implemented in microcode, but a few are hard-wired. For more details on instruction decoding, see Instruction decoding in the Intel 8087 floating-point chip. ↩ I use decimal addresses for the microcode because the Opcode Collective started using decimal addresses, and it would be confusing to change now. ↩ The microcode shows that scaling 0 by anything, or scaling anything by 0, leaves the value unchanged. My view is that the designers took a shortcut here, rather than returning the "right" value. Since the 8087 defines 0×∞ as NaN, it seems to me that 0×2∞ should also be NaN, so FSCALE(0, ∞) should be NaN, not 0. The designers probably made the valid decision that nobody really cared about corner cases on the obscure FSCALE instruction. For other instructions, the behavior with denormals, unnormals, and zeros is documented (tables S-24 to S-26 in the Numerics Supplement documentation), but FSCALE is omitted. ↩ The 8087 has separate buses for the exponent and the significand, and the adder is only connected to the significand bus, so how does the exponent get to the adder? The trick is that there is a 16-bit gateway between the exponent bus and the significand bus, so the exponent can be copied over. ↩ I described the 8087's adder here. In brief, subtraction is performed by inverting the B register's value when it is fed into the adder. The carry-in to the adder is set to 1, so this in effect performs a two's-complement subtraction. ↩ Why does the microcode have separate paths to add a positive scale and subtract a negative scale? The reason is that values are stored as a sign bit and an unsigned value, not two's complement like standard integers. As a result, the adder can't perform signed addition directly. Instead, the adder circuitry must be explicitly directed to complement the B register value and perform a subtraction. ↩ This flowchart shows the SPECIAL_TMPS subroutine. The structure of this routine is complicated because paths split off and rejoin. One tricky path is the code to determine if there are 0, 1, or 2 NaN values, and take the maximum NaN if there are two. Another complication is the exception exits, which raise an interrupt if the interrupt is not masked, but not for a comparison instruction. The two return values are returned through flag and expConv. A flowchart for the SPECIAL_TMPS subroutine. Click for a larger version. The actions of SPECIAL_TMPS are summarized below. It returns status through the flag flip-flop and the exponent converter register (expConv). Its actions are: table.status {border-collapse: collapse;} table.status tr:first-child {border-bottom: 1px solid #ccc;} table.status th,td {padding: 0 10px; text-align: center;} table.status th:first-child,td:first-child {border-right: 1px solid #ccc;} InputResultflagexpConv emptyNaN, exception11 NaN(larger) NaN, exception11 infinityinfinity01 denormunnorm10 div abnormalno change00 (The last row signals an abnormal value during division computation; I'm still investigating this.) ↩ Prof. William Kahan, who guided the development of the 8087, was disappointed that some floating-point features were unused because of a vicious circle: the features didn't receive good compiler support, so programmers didn't use the features, so compiler developers claimed a lack of demand for the features and didn't implement support. Using multiple values of NaN to record how and/or where an NaN came into existence was an example of a feature that lacked software support. See Lecture Notes on the Status of IEEE Standard 754 for Binary Floating-Point Arithmetic for a detailed discussion of NaN and other issues. ↩ The 8087 makes heavy use of micro-subroutines, with a 6-level stack for microcode subroutine calls. Microcode jumps and subroutine calls get the address from a jump table. The index from the jump table comes from 6 bits of the micro-instruction. We unimaginatively named the entries in the microcode jump table as J0, J1, and so forth based on the index, but I'm adding more meaningful names as I figure them out. As for the names for micro-instructions, we don't have any information on what names were used by Intel (unlike the 8086). I invented names, influenced by the names in Gloriouscow's disassembly. ↩ A subroutine that I call SPECIAL_VAL handles denormalized values, infinity, and NaN. (This subroutine is primarily used by SPECIAL_TMPS, but is also used by FRNDINT (round to integer) and FSQRT.) First, the subroutine looks at the exponent of tmpA; if the exponent is zero, the value is denormalized. (The value could also be zero, but that was handled earlier.) If so, the denorm exception is set. Comparison instructions such as FCOM handle denorms differently, but I'll ignore that for now. The code at #1576 tests if the denorm triggered an interrupt; if so, the instruction ends with the interrupt. If the interrupt was masked, the code converts the denorm to an unnorm by changing the tag to norm and changing the exponent to 1 (which corresponds to the very negative, smallest valid value because of the exponent bias). The result of the subroutine is returned through a special flag flip-flop. SPECIAL_VAL (J12): #1572 tmpA:exp -> sumreg:frac Handle special value #1573 jmp #1581 if not Sum register is zero Test exp for denorm #1574 except:denorm #1575 jmp #1577 if compare instruction No exception for comparison #1576 jmp #1571 if DE (denormalized) interrupt RNI if exception #1577 norm -> tmpA tag Handle denorm: tag empty? or valid? #1578 1 -> tmpA:exp Change to unnorm #1579 0 -> flag Clear flag #1580 return #1581 shift tmpA:frac L 0 bytes, 1 bits Shift to check if infinity vs NaN #1582 shift L -> sumreg:frac #1583 jmp #1579 if not Sum register is zero Clear flag for NaN #1584 1 -> flag Set flag for infinity #1585 return At #1581, the code checks if the value is infinity or NaN. Interestingly, this test isn't done directly, but by manipulating the value with the shifter. Recall that infinity has a significand of 10...00, while NaN has at least one additional 1 bit. The code shifts the significand one bit to the left; a zero result indicates infinity, while a nonzero result indicates NaN. As before, the result is returned in the flag flip-flop. ↩ The logic to compute the rounding bits is more complicated than described. There are two micro-instructions with slightly different behavior depending on the expConv value, but I won't get into that here. ↩ The ADJUST_PRECISION subroutine appears to return infinity if the significand overflows after rounding up, but I'm not entirely happy with this. For instance, 1.111... should round up to 2, not infinity; the significand overflows, but that's not an overflow of the float. Presumably, this gets fixed somewhere else. ↩ I don't know why Intel failed to document the feature that a condition code indicates whether a value was rounded up or down. The 8087 documentation is very thorough with corner cases; usually, when I find a strange circuit, I can find a line in the documentation that explains why it is there. Maybe the condition code feature was buggy, so it was easier to not document it? Maybe this feature was a hidden trap to catch competitors that copied the chip? (Intel had a secret instruction in the 8086 for this purpose, but NEC's version of the 8086 didn't have it, much to the disappointment of Intel's lawyers.) Maybe Intel wasn't sure if they wanted to support the feature in later versions? (This is why some of the 8085 processor's instructions weren't documented.) For now, it's a mystery. ↩

4 days ago 1 votes
I Am a Flat-Rate Monthly Responsibility Service

Most people, when asked why they do what they do, lie. This isn’t because they’re malicious but it’s because the honest justification for a career is rarely noble. It’s usually a combination of a decent paycheck, tolerable hours, and whatever neurosis you

5 days ago 3 votes
Super ultrawide Niri

A Niri workspace with 7 visible columns. After having used practically the same xmonad configuration for a decade and a half I’ve now modernized my setup with the scrollable-tiling Wayland compositor Niri. It’s been a bit of a struggle to unlearn my old workflow but I’m really growing to love Niri’s scrollable workflow, especially on my new super ultrawide display. Samsung Odyssey Neo G9 G95NC 57” My new 57” single monitor setup. What kicked off my Niri journey was the purchase of a new super ultrawide monitor. I bought the 57” Odyssey Neo G9 as it was the largest monitor I could find. (It’s marketed as a “gaming” display but it’s really an amazing productivity display.) It replaced my old 3-monitor setup: My old 3-monitor setup. The new display is wider so I had to move the speakers around 10–15cm further apart. I was debating whether to replace the center 31.5” monitor or replace all monitors with a single one but I think I made the right choice with the ultrawide. The curvature wasn’t an issue (I’ve come to prefer it) and the extra vertical space the portrait side monitors provided wasn’t as crucial as I thought. I think an ultrawide is worth it just to get rid of the annoying bezels. Small things can be a big thing sometimes. A more dynamic workflow xmonad and Niri are similar yet different. Both automatically lay out windows as you spawn them but xmonad (at least the way I used it) follows a layout algorithm that re-flows using a “master” window and combines the rest of the windows into one space, while Niri lays out windows in columns. The change is subtle but it implies that a new window won’t change the size of other windows. This is very nice if you spawn a lot of short-lived terminals or web browsers like I do and it reduces the amount of manual reshuffling I spend time on. My xmonad workflow was more static than my Niri one. In xmonad I made heavy use of workspaces, mapping ten workspaces mentally to different programs, such as 0 Firefox and 1 terminal logs on the left monitor; 3, 4 and 5 for different Neovim instances on the center monitor; 8 as chat and 9 for music or video on the right monitor. I had no rules to enforce this; it’s an emergent behaviour that served me well for years. With Niri it’s more dynamic. I still use workspaces but they no longer have direct shortcuts, I simply go up/down in the workspace list. Maybe I’ll add them in the future but with 3–4 workspaces that’s not as necessary. I spawn workspaces/windows when I need them and remove them when I’m done. Usually it’s one workspace per project (yes, I’m now one of those who have multiple up at once) with all the related things such as editor, terminals, and browser with docs. I don’t typically utilize the full screen width and I try to keep the things I’m working on in the center, often leaving 10–30% gaps on the sides. Even though I don’t normally use the “endless scrolling” feature of Niri I re-center selected windows all the time so I can look straight ahead as much as possible. Keyboard shortcuts As a fan of keyboard layouts of course I have to spend some time tinkering with good keyboard shortcuts (especially as Niri’s recommended keybinds don’t map well with my custom keyboard or custom layout). Navigation layer What I did was add a new navigation layer that’s enabled by holding Tab (ring + middle + index on the left-hand side) with all Niri related movement and layout keybinds. In the graphics above, all green-colored keys emit Gui (which gates all window manager commands) and you can see: Long press on Close Window to close a window. The long press requirement prevents accidentally closing windows. Arrows move through columns/windows. Long press resizes them. Workspace Up/Down focuses a different workspace. Center a column. Consume/Expel to combine windows into one column. (consume-or-expel-window-left/consume-or-expel-window-right) Expand Column makes a column take up all remaining space. (expand-column-to-available-width) Audio controls. To press them I release the index finger (keeping the ring and middle finger pressed to keep the layer active) and use the index to press the audio buttons. Mouse buttons. In Niri you can move floating windows with Gui + Left Mouse and Gui + Right Mouse to resize them. As my main mouse is a trackball integrated into the keyboard I had to add them to the left-hand side. I ended up using QMK’s customizable key repress feature that allows me to: Tab combo with my three fingers (layer is active) Release only the index (layer is still active, same as with the audio controls) Press the index again (now detects the press Gui + Left Mouse key down) Use the trackball to move the window And similarly for the right mouse button to resize with the middle finger. Works great! Because there are so many commands I want to send I placed Ctrl on the thumb that provides movement-related commands like so: For example: Arrows move columns/windows in the four directions. Move columns to the neighboring workspaces. Center visible columns. (center-visible-columns) Slightly different consume/expel semantics. (consume-window-into-column/expel-window-from-column) Regular keymaps These are triggered in the “normal” way by first pressing the Super combo and then another key on the base layer (I use autoshift so I shift with a long press). Window management Super + F toggle windowed fullscreen (keep column width) Super + Shift + F fullscreen window (over the entire display) Super + M maximize column (moves other columns) Run stuff Super + Enter terminal Super + E Noctalia’s launcher (also exists on the navigation layer as Launch) Super + S show Noctalia control center Super + Shift + S show Noctalia settings Super + Q power off monitors (they wake on input) Super + Shift + Q show Noctalia session menu (reboot etc) Super + Shift + L lock screen Misc Super + H show hotkey overlay Super + P interactive screenshot Super + Shift + P screenshot selected window Tweaks to the standard CachyOS setup In the process of moving from xmonad to Niri I also moved from Void Linux to CachyOS and I let the installer install Niri and give me a basic configuration together with Noctalia (that provides a statusbar, notifications, and a bunch of things you apparently need). Center the status bar and other Noctalia windows My centered Noctalia status bar. Feels absolutely required on this screen otherwise things end up in the corners. Firefox on XWayland Force Firefox onto XWayland as the Wayland popup manager is broken: environment { MOZ_ENABLE_WAYLAND "0" } Dead keys for Ghostty For some reason dead keys were broken in Ghostty. This is bad for me as the OS keyboard is set to Swedish and it uses them to type ~ (quite a crucial character for a programmer). The fix: environment { GTK_IM_MODULE "ibus" QT_IM_MODULE "ibus" XMODIFIERS "@im=ibus" } This needs ibus installed and running. Melange colorscheme Noctalia discovers custom color schemes under ~/.config/noctalia/colorschemes/<Name>/<Name>.json, so I dropped in my trusty Melange colorscheme there: { "dark": { "mPrimary": "#EBC06D", "mOnPrimary": "#292522", "mSecondary": "#A3A9CE", "mOnSecondary": "#292522", "mTertiary": "#85B695", "mOnTertiary": "#292522", "mError": "#D47766", "mOnError": "#292522", "mSurface": "#292522", "mOnSurface": "#ECE1D7", "mSurfaceVariant": "#34302C", "mOnSurfaceVariant": "#C1A78E", "mOutline": "#867462", "mShadow": "#1a1816", "mHover": "#E49B5D", "mOnHover": "#292522", "terminal": { "normal": { "black": "#867462", "red": "#D47766", "green": "#85B695", "yellow": "#EBC06D", "blue": "#A3A9CE", "magenta": "#CF9BC2", "cyan": "#89B3B6", "white": "#ECE1D7" }, "bright": { "black": "#34302C", "red": "#BD8183", "green": "#78997A", "yellow": "#E49B5D", "blue": "#7F91B2", "magenta": "#B380B0", "cyan": "#7B9695", "white": "#C1A78E" }, "foreground": "#ECE1D7", "background": "#292522", "selectionFg": "#C1A78E", "selectionBg": "#403A36", "cursorText": "#292522", "cursor": "#EBC06D" } } } Then pick the colorscheme: "colorSchemes": { "darkMode": true, "predefinedScheme": "Melange", "useWallpaperColors": false } Layout appearance The default appearance was pretty I admit but way too much blank space and weirdness. Some tweaks: layout { // Required for noctalia-shell to set wallpaper background-color "transparent" // Never auto-center focused columns (too much movement) center-focused-column "never" // But do center a single window always-center-single-column // No extra space around it all struts {} // No gaps between windows gaps 0 // The focus ring was annoying focus-ring { off } // Use a border with consistent width for all windows instead border { on width 2 active-color "#ebc06d" inactive-color "#403a36" } // Setting widths is important with such a large screen preset-column-widths { proportion 0.15 proportion 0.3 proportion 0.4 } default-column-width { proportion 0.15; } // Heights too, why not? preset-window-heights { proportion 0.15 proportion 0.5 proportion 1.0 } } // Prevent the mouse from opening the overview in the corners gestures { hot-corners { off } } Keep windows centered Niri has the always-center-single-column option, which is nice as I want to keep as much as possible in the center of the monitor when I’m working. But I very frequently use 2–3 smaller windows and with my frequent opening and closing I’d like them centered too. Luckily, Niri has an IPC you can use to make a small program that reacts to events and does this for you. I made a small rust project using the niri-ipc crate that does this for me: The autocenter implementation [dependencies] niri-ipc = "26.4.0" use std::collections::HashMap; use std::io; use niri_ipc::socket::Socket; use niri_ipc::{Action, Event, Request, Response, Window, Workspace}; #[derive(Clone, Copy, PartialEq, Eq, Hash)] struct WindowId(u64); #[derive(Clone, Copy, PartialEq, Eq)] struct WorkspaceId(u64); struct OutputName<'a>(&'a str); struct WindowState { workspace: Option<WorkspaceId>, width: f64, } fn main() -> io::Result<()> { let mut socket = Socket::connect()?; if !matches!(socket.send(Request::EventStream)?, Ok(Response::Handled)) { eprintln!("niri rejected event stream"); std::process::exit(1); } let mut known: HashMap<WindowId, WindowState> = HashMap::new(); let mut read_event = socket.read_events(); loop { let result = match read_event()? { // A full snapshot of the current state. Just refresh our state. Event::WindowsChanged { windows } => { known = windows .into_iter() .map(|w| { ( WindowId(w.id), WindowState { workspace: w.workspace_id.map(WorkspaceId), width: w.layout.tile_size.0, }, ) }) .collect(); Ok(()) } Event::WindowOpenedOrChanged { window } => { let workspace = window.workspace_id.map(WorkspaceId); let entry = WindowState { workspace, width: window.layout.tile_size.0, }; let prev = known.insert(WindowId(window.id), entry); match prev { // Don't center floats. _ if window.is_floating => Ok(()), // New window, try to re-center. None => maybe_center_new(&window), // Window changed workspace, try to re-center. Some(state) if state.workspace != workspace => center_focused_if_fits(), // Skip other things. Some(_) => Ok(()), } } Event::WindowClosed { id } => { if known.remove(&WindowId(id)).is_some() { center_focused_if_fits() } else { Ok(()) } } Event::WindowLayoutsChanged { changes } => { // Only re-center if the width was changed, otherwise our re-center will // loop back indefinitely. let mut resized = false; for (id, layout) in changes { if let Some(state) = known.get_mut(&WindowId(id)) { if (state.width - layout.tile_size.0).abs() > 0.5 { state.width = layout.tile_size.0; resized = true; } } } if resized { center_focused_if_fits() } else { Ok(()) } } _ => Ok(()), }; if let Err(e) = result { eprintln!("autocenter: {e}"); } } } /// Center a newly created window if the workspace is focused and if there's surrounding free space left. fn maybe_center_new(window: &Window) -> io::Result<()> { let Some(workspace_id) = window.workspace_id.map(WorkspaceId) else { return Ok(()); }; let Some(focused) = focused_workspace()? else { return Ok(()); }; if WorkspaceId(focused.id) == workspace_id { center_if_fits(&focused)?; } Ok(()) } /// Center windows in the focused workspace if there's surrounding free space left. fn center_focused_if_fits() -> io::Result<()> { if let Some(focused) = focused_workspace()? { center_if_fits(&focused)?; } Ok(()) } /// Center windows in the workspace if there's surrounding free space left. fn center_if_fits(workspace: &Workspace) -> io::Result<()> { let Some(output) = workspace.output.as_deref().map(OutputName) else { return Ok(()); }; let Some(width) = output_width(output)? else { return Ok(()); }; if workspace_width(WorkspaceId(workspace.id))? < f64::from(width) { center_visible_columns()?; } Ok(()) } /// Issue a one-shot query to Niri, wait, and return the response. fn query(request: Request) -> io::Result<Response> { match Socket::connect()?.send(request)? { Ok(response) => Ok(response), Err(msg) => Err(io::Error::other(msg)), } } /// Get the focused workspace. fn focused_workspace() -> io::Result<Option<Workspace>> { match query(Request::Workspaces)? { Response::Workspaces(ws) => Ok(ws.into_iter().find(|w| w.is_focused)), _ => Ok(None), } } /// Get the width of an output (monitor). fn output_width(name: OutputName<'_>) -> io::Result<Option<u32>> { match query(Request::Outputs)? { Response::Outputs(outputs) => Ok(outputs .get(name.0) .and_then(|o| o.logical.as_ref()) .map(|l| l.width)), _ => Ok(None), } } /// Calculates the width of all columns in the workspace. fn workspace_width(workspace_id: WorkspaceId) -> io::Result<f64> { let Response::Windows(windows) = query(Request::Windows)? else { return Ok(0.0); }; let mut columns: HashMap<usize, f64> = HashMap::new(); for w in windows { if w.workspace_id.map(WorkspaceId) != Some(workspace_id) { continue; } if let Some((col, _)) = w.layout.pos_in_scrolling_layout { let width = columns.entry(col).or_insert(0.0); *width = width.max(w.layout.tile_size.0); } } Ok(columns.values().sum()) } /// Send a command to center the visible columns. fn center_visible_columns() -> io::Result<()> { if let Err(msg) = Socket::connect()?.send(Request::Action(Action::CenterVisibleColumns {}))? { eprintln!("center-visible-columns rejected: {msg}"); } Ok(()) } One catch is that if a new window overflows the monitor width, the script won’t center the columns even if there would be free space left afterwards. This is a little weird but it’s consistent with Niri’s center-visible-columns command. I had a small itch to try to hack around it but in the end I left it alone… Is Niri worth it? Yes, absolutely. Niri has been a huge upgrade for me in combination with a single wide screen. My xmonad setup worked really well with three monitors—arguably a better fit in that context than Niri—but for the big-screen use-case Niri is superior. I’m curious how it holds up on my laptop, once I gather enough energy to install CachyOS on it… But that’s a side quest. The big-screen setup I spend most of my days in is the best I’ve ever had, and I have no desire to go back.

5 days ago 2 votes
📚 BoredReading

You seem to be enjoying this.

Join free to unlock everything.

Create free account

Already have an account? Sign in