ET-SoC-1 test drive: from a laptop simulator to silicon

I built the ET toolchain and platform on a Mac and ran the hello worlds on the sys_emu simulator. Then I ran a first real workload, a scalar fp32 matrix multiply, on an ET-SoC-1 card in the AI Foundry lab. Below are the numbers, what broke on the way, and what came next.

Checked on three cards (26 September 2026). This page's SGEMM claims were re-measured under a pre-registered plan on aifoundry2, aifoundry3 and aifoundry1 card 1, three passes each. Of 5 claims tested here, this page counts 1 held and 4 corrected; the hub’s scoreboard, 5 “a test behind it failed”. Every output was exact again; the n ≥ 512 times and the 127 GFLOP/s repeat on all three cards; the corrections are the n = 64 time (0.55 ms, not 0.31), the speedup's last digit (28.5×, not 28.7×), the launch-to-launch spread (up to 7% at n = 512 on 32 shires) and the device time per run (0.13–0.39 s). The table now gives the three cards' mean. Record: docs/reports/data/2026-09-25-claims-v3.

Terms used on this page

The ET-SoC-1 is Esperanto's 1,088-core RISC-V accelerator, now open-sourced by AI Foundry. Kernels run on 32 compute shires of 32 minion cores; each minion runs two hardware threads (harts), so a full launch is 2,048 harts. sys_emu is the chip's functional simulator, part of et-platform: it checks results, not speed. More in the hub's glossary.

Scalar fp32 SGEMM on one card, n = 1024

127 GFLOP/s

2,048 harts on 32 shires, plain fmadd.s (no SIMD, no tensor unit yet); 127 on each of three cards

Speedup, 1 → 32 shires

28.5×

n = 512: 68.0 ms → 2.39 ms, three cards' mean (28.3–28.6× per card)

Mismatches vs host reference

0 of 1.6M

all silicon runs, fp32 vs double; 0 of 14.2M in nine runs on three cards

Device held per run

0.13–0.39 s

mostly runtime init (0.13–0.18 s); the card is shared

Mac setup from scratch

≈15 min

M5 Max; the toolchain build is 7 of it

What I did

  1. Development environment on macOS. The build expects Ubuntu 24.04. I used a Lima VM running Ubuntu 24.04 natively on arm64, with 16 vCPUs and the repo mounted at the same path.
    Two build-time notes
    • The ET RISC-V GCC (aifoundry-org/riscv-gnu-toolchain, branch et, GCC 15.2) has to be built from source, because the release tarball is x86-only. That takes about 7 minutes.
    • The et-platform superbuild then takes about 2 minutes: firmware, sys_emu, runtime, and gp-sdk. Apart from fix 1 below, it needed no changes for arm64.
  2. Hello worlds on the simulator. Each run takes about 40 s, and about 37 s of that is simulated firmware booting 33 shires.
    Three simulator checks
    • it_test_code_loading, the README/CI hello world: 3/3 passed.
    • gp-sdk hello_world_launcher + print.elf: prints HELLO WORLD!!!!. This needed a fix, below.
    • My own check-in kernel, where every hart writes one cache-line record: 2048/2048 harts from 32 shires.
  3. A first workload on silicon. I wrote a small SGEMM in the style of marty1885/et-testdrive, using just the runtime API and et-common-libs. That mattered because the lab machine's /opt/et is older and has no gp-sdk. (gp-sdk was later built on aifoundry2, outside /opt/et, with scripts/deploy-lab-gpsdk.sh; /opt/et itself still has none.) I verified it on the simulator on the laptop. Then I copied the sources to aifoundry3 (4.5 MB), built them there against its own toolchain, and ran them on the card with a hard 10 s cap.

SGEMM on the card

This computes C = A·B for n×n fp32 matrices, checked against a double-precision host reference.

How the kernel is written
  • The unit of work is a 1×16 strip of C, which is exactly one 64-byte cache line. The minion L1 caches aren't coherent and write back whole lines, so giving each hart whole lines is what makes the parallel result correct.
  • The 16 accumulators stay in FP registers. The inner loop is 17 loads + 16 fmadd.s per step of k.
nShiresHartsTime per launchGFLOP/sMismatches
641640.55 ms0.960 / 4,096
51216468.0 ms3.90 / 262,144
512322,0482.39 ms1120 / 262,144
1024322,04816.9 ms1270 / 1,048,576
The three-card check, 25–26 September 2026: aifoundry2, aifoundry3 and aifoundry1 card 1, three passes each, each row its own process with three launches; the time is the mean over the three cards of each card's mean, and GFLOP/s is 2·n³ over it. Per card: 0.54, 0.54 and 0.56 ms; 68.0, 68.0 and 68.1 ms; 2.38, 2.38 and 2.41 ms; 16.9, 16.9 and 16.9 ms. Time is host-measured kernelLaunch → waitForStream, so launch overhead is included. Launch to launch within a process the time varied by up to 0.3% at n = 512 on one shire, 1.0% at n = 1024 and 3.6–6.9% at n = 512 on 32 shires; at n = 64 single launches ranged from 0.41 to 0.66 ms. Every output was exact, in every pass on every card. The first run, on aifoundry3 on 18 September, gave 0.31 ms (1.7 GFLOP/s), 68.0 ms, 2.37 ms and 16.8 ms; its printout was not kept, and its n = 64 time did not repeat on any card.

How far from the chip's peak?

About the FOSDEM comparison

For scale, the FOSDEM 2026 talk “Zero to matmul with the ET-SoC-1” (512×512 fp32, 650 MHz; its rungs are transcribed in docs/et-soc1-notes.md) climbed from a naive kernel to the tensor unit. The chart puts this page's scalar kernel (16 accumulators per hart) on the same log axis. The tensor unit's inner loop, on cached tiles, ran the same day at 9.5 TFLOP/s in fp32 on aifoundry2 (Matmul efficiency; the same on each of three cards on 26 September), FOSDEM's 10.25 scaled to 600 MHz; a full tensor-unit GEMM has not been run, while FOSDEM's rung is a whole 512×512 matmul.

Needs JavaScript; the table above has this page's rows.

Things worth fixing upstream

1. Fresh et-platform builds fail since 2026-08-05 build

The symptom
  • ThirdParty() pins easy_profiler to the git describe ref v2.1.0-66-gcc0e154 but clones with GIT_SHALLOW TRUE.
  • Upstream easy_profiler got new commits on 2026-08-05, so a depth-1 clone no longer contains cc0e154.
  • The build dies with fatal: invalid reference.

Fix: clone the third-party repos fully, or pin full SHAs.

Patch (1 line)
diff --git a/cmake/ProjectFunctions.cmake b/cmake/ProjectFunctions.cmake
index b237797..7f34554 100644
--- a/cmake/ProjectFunctions.cmake
+++ b/cmake/ProjectFunctions.cmake
@@ -38,7 +38,7 @@ function(ThirdParty name url tag cmake_args)
         ${name}
         GIT_REPOSITORY ${url}
         GIT_TAG ${tag}
-        GIT_SHALLOW TRUE
+        GIT_SHALLOW FALSE  # describe-style tags (vX-N-gSHA) are not reachable in a depth-1 clone once upstream moves
         PREFIX ${CMAKE_BINARY_DIR}/${PROJECT_BASE_DIR}/${name}
         CMAKE_ARGS
                    -DCMAKE_INSTALL_PREFIX=${STAGING_DIR}

2. gp-sdk launchers can’t boot the simulator gp-sdk

The symptom
  • GenericLauncher never tells sys_emu which firmware to preload: SP BL2 and the master/machine/worker minion images.
  • With --device_type=sysemu, every gp-sdk launcher (hello_world_launcher, saxpy_launcher, …) boots the Service Processor on empty memory.
  • It then loops on “Trapping to the same address” and writes gigabytes of log within minutes.
  • The runtime’s own tests avoid this because TestUtils.h sets the paths.

Fix: set them the same way, from the installed CMake packages.

Patch (GenericLauncher.cpp + CMake)
diff --git a/gp-sdk/host/CMakeLists.txt b/gp-sdk/host/CMakeLists.txt
index bffd6eb..df26535 100644
--- a/gp-sdk/host/CMakeLists.txt
+++ b/gp-sdk/host/CMakeLists.txt
@@ -23,6 +23,8 @@ find_package(runtime REQUIRED)
 find_package(esperantoTrace REQUIRED)
 find_package(hostUtils REQUIRED)
 find_package(sw-sysemu REQUIRED)
+find_package(EsperantoBootLoader)
+find_package(EsperantoDeviceMinionRuntime)
 
 find_package(PkgConfig REQUIRED)
 pkg_check_modules(FFTW3F REQUIRED fftw3f)
diff --git a/gp-sdk/host/sdk/CMakeLists.txt b/gp-sdk/host/sdk/CMakeLists.txt
index 2cef11f..9da73db 100644
--- a/gp-sdk/host/sdk/CMakeLists.txt
+++ b/gp-sdk/host/sdk/CMakeLists.txt
@@ -10,6 +10,27 @@ add_library(etsoc_gpsdk STATIC
     src/lib/GenericLauncher.cpp 
 )
 target_include_directories(etsoc_gpsdk PUBLIC include sdk/include)
+
+# Device firmware that sysemu has to preload (same source as the runtime tests).
+foreach(_fw BOOTROM_TRAMPOLINE_TO_BL2_ELF BL2_ELF MASTER_MINION_ELF MACHINE_MINION_ELF WORKER_MINION_ELF)
+  set(_${_fw} "")
+endforeach()
+if (EsperantoBootLoader_FOUND)
+  get_property(_BOOTROM_TRAMPOLINE_TO_BL2_ELF TARGET EsperantoBootLoader::BootromTrampolineToBL2.elf PROPERTY LOCATION)
+  get_property(_BL2_ELF TARGET EsperantoBootLoader::ServiceProcessorBL2_fast-boot.elf PROPERTY LOCATION)
+endif()
+if (EsperantoDeviceMinionRuntime_FOUND)
+  get_property(_MASTER_MINION_ELF TARGET EsperantoDeviceMinionRuntime::MasterMinion.elf PROPERTY LOCATION)
+  get_property(_MACHINE_MINION_ELF TARGET EsperantoDeviceMinionRuntime::MachineMinion.elf PROPERTY LOCATION)
+  get_property(_WORKER_MINION_ELF TARGET EsperantoDeviceMinionRuntime::WorkerMinion.elf PROPERTY LOCATION)
+endif()
+target_compile_definitions(etsoc_gpsdk PRIVATE
+    GPSDK_BOOTROM_TRAMPOLINE_TO_BL2_ELF="${_BOOTROM_TRAMPOLINE_TO_BL2_ELF}"
+    GPSDK_BL2_ELF="${_BL2_ELF}"
+    GPSDK_MASTER_MINION_ELF="${_MASTER_MINION_ELF}"
+    GPSDK_MACHINE_MINION_ELF="${_MACHINE_MINION_ELF}"
+    GPSDK_WORKER_MINION_ELF="${_WORKER_MINION_ELF}"
+)
 target_link_libraries(etsoc_gpsdk 
     PUBLIC 
         deviceLayer::deviceLayer
diff --git a/gp-sdk/host/sdk/src/lib/GenericLauncher.cpp b/gp-sdk/host/sdk/src/lib/GenericLauncher.cpp
index 988aecb..2390811 100644
--- a/gp-sdk/host/sdk/src/lib/GenericLauncher.cpp
+++ b/gp-sdk/host/sdk/src/lib/GenericLauncher.cpp
@@ -31,6 +31,19 @@ emu::SysEmuOptions getDefaultOptions(std::string const& simulator_params) {
 
   emu::SysEmuOptions sysEmuOptions;
 
+  // sysemu runs in-process and only boots the firmware it is told to preload;
+  // without these the Service Processor starts on empty memory and traps forever.
+  auto setIfExists = [](std::string& dst, const char* path) {
+    if (path[0] != '\0' && std::filesystem::exists(path)) {
+      dst = path;
+    }
+  };
+  setIfExists(sysEmuOptions.bootromTrampolineToBL2ElfPath, GPSDK_BOOTROM_TRAMPOLINE_TO_BL2_ELF);
+  setIfExists(sysEmuOptions.spBL2ElfPath, GPSDK_BL2_ELF);
+  setIfExists(sysEmuOptions.masterMinionElfPath, GPSDK_MASTER_MINION_ELF);
+  setIfExists(sysEmuOptions.machineMinionElfPath, GPSDK_MACHINE_MINION_ELF);
+  setIfExists(sysEmuOptions.workerMinionElfPath, GPSDK_WORKER_MINION_ELF);
+
   sysEmuOptions.runDir = std::filesystem::current_path();
   sysEmuOptions.maxCycles = kSysEmuMaxCycles;
   sysEmuOptions.minionShiresMask = kSysEmuMinionShiresMask;

3. A0 errata 1.29 and the compiler question

Errata 1.29 (“VPURF timing path”): reading a VPU register one cycle after it was written can return stale data. The workaround is a guard instruction, e.g. fmv.x.w x0, fN after loads. The FOSDEM matmul loop has one for this reason.

  • GCC 15.x doesn’t insert the guard. This applies to both the lab’s 15.1 and the laptop’s 15.2. sys_emu -vpurf_warn flags every fmadd.s in this SGEMM, twice (524,291 “type A” warnings on the compute shire at n = 64; the firmware adds 68 more), and gp-sdk’s own saxpy_vector.elf as well.
  • Silicon results were still correct for all 1.6 M outputs, and for all 14.2 M outputs of the three-card check (nine runs of the four table rows on three cards).

Open questions: are the lab cards A0 steppings, and is the compiler meant to handle this? For now: check new kernels with -vpurf_warn, and treat “wrong only on silicon” as a prime suspect for this bug.

Smaller notes

Four smaller notes
  • Esperanto’s trace decoder dt2json, referenced in the gp-sdk docs, isn’t open source. A 60-line decoder on top of the header-only et-trace is enough to read kernel et_printf output.
  • GenericLauncher::loadKernel() leaves std::cout in hex mode. My first “800/800 harts” result was really 2048/2048: 800 is 2,048 in hex.
  • When the runtime starts sys_emu, it turns on the coherency, scratchpad, FLB, and tensor-store checkers by default. The errata checker is opt-in: -vpurf_warn.
  • On the card, runtime init (opening the device) takes 0.13–0.14 s on aifoundry2 and aifoundry1 card 1 and 0.18 s on aifoundry3 in the three-card check (about 0.19 s in the first session). Peak host RSS is about 2.2 GB (2.0 GiB) on all three cards for the duration of a run, which comes from the runtime’s mapped buffers.

Reproduce

Reproduce this

All commands run from a checkout of the project repository. The first two blocks use the Lima VM on the Mac (on a Linux machine with /opt/et, drop the scripts/vm prefix). The last block runs from the same checkout, on the Mac or on Linux: it copies the sources to aifoundry3, builds them there against its /opt/et and runs each table row as its own timeout 10 process. The table is the host's “launch r: X ms, Y GFLOP/s” lines; the three-card check ran the same four processes through tools/claims-v3/lat/, and its logs are in docs/reports/data/2026-09-25-claims-v3/raw/<card>/lat/p*/sg/.

# one-time: Lima VM + toolchain + et-platform (~15 min)
scripts/create-vm.sh

# simulator, on the laptop
scripts/vm make run-hello        # 2048-hart check-in
scripts/vm make run-sgemm        # SGEMM n=128 on sys_emu
scripts/vm make run-sgemm SGEMM_N=64 SGEMM_ARGS='--shires 0x1' SIM_PARAMS=-vpurf_warn
grep 'type A' build/sgemm/run/sysemu.log | grep -c ' S0:'   # 524,291 (150 MB log)

# real card: copy sources, build on aifoundry3, run each table row with a hard cap
scripts/deploy-lab.sh aifoundry3 workloads/sgemm
ssh aifoundry3 'cd ~/nekko/build/sgemm && for a in "-n 64 --shires 0x1" "-n 512 --shires 0x1" "-n 512" "-n 1024"; do timeout 10 host/sgemm_host $a --reps 3; done'

Of the steps planned that day (8-wide .ps SIMD with register blocking, A and B staged in the L2 scratchpad, then the tensor unit), the tensor unit was done the same day; SIMD and scratchpad staging were not taken further. What came next:

← All ET-SoC-1 measurement reports