Booting Fuchsia’s Zircon on a homebrew RISC-V SoC, on an Artix-7 FPGA

#FPGA#RISC-V#Zircon#Fuchsia#Vivado#bazel#Cocoapuffs

Fuchsia’s Zircon kernel now boots on Cocoapuffs, my homebrew RISC-V system-on-chip on an AMD Artix-7 FPGA. As far as I know, this is the first Fuchsia boot on programmable hardware. Below you will find the boot video and the full chain, from a SERV serial loader through OpenSBI to userspace. All of the code, from RTL to the boot chain, lives in filmil/cocoapuffs-fpga. The rest of this post chronicles the five bugs that stood in the way.

Here is the entire sequence as a screencast: programming the FPGA, uploading the OpenSBI + boot-shim + Zircon image over serial (with the roughly 24-minute upload fast-forwarded), and the board booting:

Program → upload → boot on the cocoapuffs NOEL-V board, end to end on the serial console: OpenSBI, the boot shim, physload/physboot, the Zircon kernel's init, into userspace and userboot, ending at the expected no-bootfs __builtin_trap. (Earlier raw serial capture, through the physboot handoff: hw_zircon128_capture.log.)

The boot chain

The full path from power-on to userspace has a lot of links:

SERV ihex loader ──▶ OpenSBI (M-mode) ──▶ linux-riscv64-boot-shim
      ──▶ physload ──▶ physboot ──▶ Zircon kernel ──▶ userboot (userspace)
  • SERV receives the firmware image over the serial line as Intel HEX and jumps to it. This is the only way onto the board: no JTAG bitstream reload between software iterations.
  • OpenSBI is the M-mode supervisor firmware. It parses the devicetree, sets up the SBI runtime (timer, IPI), and drops to S-mode.
  • The linux-riscv64-boot-shim is a Zircon phys executable that repackages a devicetree and ZBI into the handoff format expected by the kernel.
  • physload / physboot decompress and relocate the kernel, lay out physical memory, and hand execution over to the kernel proper.
  • The Zircon kernel runs its init hooks and, finally, launches userboot, the first userspace process.

Every one of these stages had at least one board-specific surprise.

Background

The resulting programmable hardware design is named “Cocoapuffs”, for reasons that are now lost to time. It is the end product of a project that I started long ago with a colleague from work. It even has a logo, straight from the pen of the creative genius that is my daughter:

Cocoapuffs logo

Full disclosure: while FPGA bringup is a personal project, Fuchsia is my day job. This work was done entirely on my own time, on my own equipment, and using only publicly accessible information about Fuchsia.

The complete article documenting the bringup, warts and all, is here. This design is only the tip of the iceberg of the entire toolset I developed to make it possible. I hope to tell the story about those tools one day. Fragments of that other story can be found in my bazel registry.

It has been a long time in the making, but I am now finally able to confirm that the FPGA-based system-on-chip of my own design is capable enough to boot a real production kernel. My design loads Zircon up, and boots into userspace. At the moment, for unrelated reasons there is not much going on in userspace itself, but the boot process does complete as one would expect. Why not Linux instead? I thought it would be more interesting to boot Fuchsia instead of, say, Linux. Not many people have brought up Fuchsia on a brand new device.

The boot video above is slightly abridged to shorten the boring parts. In contrast to the post about the 32-bit RISC-V design, which showed the entire synthesis and programming run in about 6 minutes, synthesis and place-and-route for a modern 64-bit core is serious business and as a result takes a long time. It runs for about 1.5 hours depending on design particulars, and watching it is not very illuminating, so I fast-forwarded it a bit.

If you are curious to learn more about Zircon, you can check out the official Zircon documentation, or if you prefer reading on paper, you can check out my technical report on Zircon. In either case, let me know if you have feedback.

The main component of the design is the NOEL-V RV64 soft core running on an Artix-7 (xc7a200t) FPGA. The design is a homebrew SoC I call cocoapuffs: a GRLIB noelvsys subsystem, a small SERV core that acts as the serial boot loader, DDR3, an APBUART, and the usual GRLIB plumbing, all built with Bazel rules for Vivado. The board is the AX7A200B from Alinx, an exceptionally reasonably priced and well-made FPGA development board.

The complete source for the Cocoapuffs SoC (the RTL, the RISC-V firmware and boot chain, and the Bazel build that ties it all together) is open source at github.com/filmil/cocoapuffs-fpga.

Zircon is not Linux. There is no earlycon, no forgiving BIOS, and the RISC-V port is comparatively young. When something goes wrong before the console comes up, you get silence. Most of this post is about turning that silence back into information. As a concession to the somewhat constrained space on the xc7a200t device, the boot process is somewhat unorthodox. (One might say that the xc7a200t is roomy as far as FPGA devices go. While that is mostly true, remember that CPU designs are notorious for mapping poorly onto FPGA fabric, generating congestion in deeply pipelined logic and spending FPGA resources like there is no tomorrow.)

The rest of the article documents the final stretch to getting to a working Zircon boot.

The bringup ladder

I like to think of a bringup like this as a ladder: each rung is a bug that hides the next one. You only ever see the lowest unsolved rung, so progress looks like a series of identical-looking hangs that each turn out to have a completely different cause.

Rung 1: DDR3 that lies about its size

The AHB decoder maps a 1 GB window at the DDR3 base, but the physical part is only 128 MB and the window aliases every 0x0800_0000. Declaring 1 GB of RAM in the devicetree let the kernel place its data ZBI up high, where it silently wrapped back onto physical address 0 and overwrote the firmware. The fix was to tell the truth in the devicetree (one 128 MB memory@0 node) and to place the ZBI at an address that lives inside the first alias:

memory@0 {
    device_type = "memory";
    reg = <0x0 0x00000000 0x0 0x08000000>; // 128 MB, no aliasing games
};

Rung 2: Caches off, so sc.w never succeeds

GRLIB comes out of reset with the L1 caches disabled. The store-conditional (sc.w/sc.d) path in this core requires a D-cache hit, so with caches off, every store-conditional fails and any lock-based code spins forever. A three-instruction boot stub before OpenSBI turns the caches on via the GRLIB cache control CSR (0x7C1):

li   t0, 0x1CF     # dsnoop | dflush | iflush | dcs=11 | ics=11
csrw 0x7C1, t0     # GRLIB CCTRL: enable I$ and D$

Rung 3: The console goes silent because the panic path faults

This is my favorite one. The kernel would come up far enough to hit an early assert, start printing a backtrace… and then go completely silent. The backtrace walker in GetBacktraceCommon checked for a null frame pointer before the loop, but not inside it. On this core the chain terminated with a null FP mid-walk, so the walker read from (0 - 16) + 8 == 0xFFFF...F8, faulted, re-entered the panic handler, and recursed until the console never emitted another byte. A one-line guard inside the loop restored the panic output, and that panic output is what told me about all the later rungs.

Rung 4: A phantom memory node

An early memory@c0000000 node described the 4 KB boot BRAM as if it were 1 MB of general RAM. The PMM dutifully placed page bookkeeping near the top of that window, the AHB signaled an error on the write, and you got a store access fault (scause=7) deep inside PmmNode::InitArena. Deleting the bogus node resolved the fault; the kernel only ever needed the real DDR3.

Rung 5: No PLIC, no boot

The RISC-V kernel refuses to come up without an interrupt controller: interrupt_get_max_vector() returns 0 and an assert fires. NOEL-V’s GRLIB PLIC lives at 0xF800_0000. Teaching the devicetree about it got us past the assert and, at last, all the way to “starting user space”.

Debugging without a console

For long stretches there was no working serial console, so the ground truth came from the RISC-V debug module. NOEL-V exposes an AHB-mapped debug module that you can reach over JTAG (BSCANE2 to an OpenOCD ahb_read config). With that I could halt the hart, read CSRs (scause, sepc, sstatus), and, crucially, dump the kernel’s debuglog ring buffer straight out of RAM. When the wire is silent, the debuglog still has the whole story:

$ openocd -f ahb_read.cfg -c 'ahb_dump 0x858400 3072'
...
starting user space
userboot: entry point                     @  0x2097aa
userboot: stack mapped                    @

I also built a small AHB transaction recorder peripheral: it passively snoops the core’s AHB bus into a BRAM ring and, on a button press, dumps the captured trace over the UART. When a hang is a bus problem rather than a control-flow problem, being able to inspect the last few thousand transactions is worth a great deal. (It also caused a bug of its own: more on that below.)

Wherever possible I reproduced issues in simulation first using NVC and xsim, because a synthesis and place-and-route cycle on this design takes about an hour, and nobody wants an hour-long edit-compile-debug loop.

Two console gotchas worth their own section

By this point the kernel booted all the way to userspace, but only the debuglog dump over JTAG proved it. On the serial line, the boot always cut off in the exact same place. This turned out to be two separate problems stacked on top of each other.

The debug recorder was eating the UART

The AHB recorder I mentioned earlier was wired to steal the UART transmitter when it dumps. To make things worse, the boot loader automatically triggered a dump about 30 seconds after handing the CPU over to the core. That trigger fired on the first boot after programming the FPGA (right around the physboot-to-kernel transition), so a fresh program-then-boot sequence never showed the kernel on the wire. (A second boot without reprogramming was completely clean, which is exactly the kind of Heisenbug that wastes an entire afternoon.) Dropping the recorder from the UART mux gave the core the serial line for the whole boot.

Userspace output needs the PLIC, not the hart

Even with the recorder gone, the console died at the last kernel init hook, and userspace never appeared over serial. The underlying cause is a great illustration of how Zircon’s console actually works:

  • Early boot writes to the UART synchronously: busy-wait on “TX ready” and push a byte. This works reliably regardless of interrupt state.
  • At runtime, the console is driven by the debuglog drainer thread, which performs a blocking, interrupt-driven UART write: it enables the TX interrupt and then sleeps until that interrupt fires.

The TX interrupt never fired because the devicetree routed the UART’s interrupt to the hart-local controller (cpu0_intc) instead of the PLIC:

/* wrong: a hart-local line is never delivered as a peripheral IRQ */
interrupts-extended = <&cpu0_intc 1>;

/* right: the APBUART is PLIC source 1 in this noelvsys */
interrupts-extended = <&plic0 1>;

On RISC-V, a peripheral interrupt must arrive as an external interrupt through the PLIC. With the wrong routing, the drainer thread blocked forever the moment the TX FIFO filled up. The log kept accumulating in the ring (which I could still inspect over JTAG), but nothing after the first FIFO buffer ever reached the wire.

Reparenting the UART interrupt to the PLIC was necessary, but it turned out not to be sufficient. With <&plic0 1>, the PLIC initializes and the kernel switches over to interrupt-driven TX (UART: IRQ driven TX: enabled). Yet the APBUART’s TX interrupt on source 1 still never gets serviced: a deeper GRLIB-PLIC delivery issue that I still need to root cause. The pragmatic fix that gets the entire boot, userspace included, onto the wire is to sidestep the interrupt entirely with the kernel.debug_uart_poll=true boot argument. That flag instructs the kernel to poll the TX register just like early boot does. With that in place, starting user space, userboot, and the final trap all print cleanly on the serial console, with no JTAG required.

One full proven boot

Zircon now boots the whole chain on real hardware: OpenSBI, the boot shim, physload/physboot, the kernel’s full init, and finally into userspace and userboot. The image I boot is a bootfs-less “eng” ZBI, so userboot correctly runs to its expected no-bootfs terminus (a __builtin_trap after “no ‘/boot’ bootfs in bootstrap message”), matching the exact end state of the QEMU reference for that image. That is “one full proven boot”: every stage runs, in order, on silicon I can hold in my hand.