AI Foundry lab · Fix log · 25 September 2026
aifoundry1 is fixed
Both ET-SoC-1 cards in aifoundry1 answer again. The et_soc1 kernel module was rebuilt by DKMS from the fixed driver source that was already on the machine (et-platform 353f20e), for all three kernels in /boot, and reloaded. It now reports version 0.20.0 and srcversion 47D26A305A0428B29FB7FC4, the same as aifoundry2 and aifoundry3, and a read-only management query succeeds on both cards. The interrupted 24 September kernel update was finished as well, after freeing space on the nearly full root file system. The machine was not rebooted, and no running user work was touched.
Update, 27 September 2026. Card 1 has since run real work: the whole three-card version-3 check (25–26 September, as aifoundry1-c1) and the gathers and scatters after it (E48, 26 September 06:57–09:22). Card 0 cannot hold a load: in about ten minutes of short smoke launches on 25 September its die reached 98–102 °C, and after the last one it read 115–117 °C with nothing running, drawing 66–71 W at 600 MHz, until its firmware dropped it to 300 MHz and it cooled. Its cooling (a fan, heat sink or airflow) is at fault, which its 300 MHz idle normally hides; it takes no sustained work and was left out of the check (amendment A4 of the check's record).
- Cards
- Both answer:
dev_mngt_service -m DM_CMD_GET_MODULE_FIRMWARE_REVISIONSsucceeds on card 0 and card 1; each card's management queue counted its first command. - Driver
et_soc1[0.20.0] 47D26A305A0428B29FB7FC4loaded now and installed for 7.0.0-30, -31 and -34; loaded at boot by/etc/modules-load.d/et_soc1.conf.- Kernel update
linux-image-7.0.0-34fully installed: initramfs (with ZFS) written, GRUB updated; the next boot starts 7.0.0-34, whoseet_soc1is the fixed one.- Still open
- The pool is 95% full: 6.4 GB free after clearing caches, the rest is user data (who should delete what); card et0's PCIe link still counts ~1 corrected error per second (its kernel log is rate-limited since 25 September); the two cards run different firmware; the next boot is untested; card 0 overheats under load (update of 27 September).
1. The cards now
A read-only management query per card, run as an ordinary user with nobody else on the cards (15:02 PDT):
| aifoundry1 card 0 (01:00.0) | aifoundry1 card 1 (02:00.0) | aifoundry2, aifoundry3 | |
|---|---|---|---|
| Firmware release | 1.4.1 | 1.2.0 | 1.3.1 |
| BL1 / BL2 | 0.21.2 | 0.18.0 | 0.20.0 |
| PMIC firmware | 1.6.1 | 1.3.0 | 1.5.0 |
| Master / worker / machine minion | 0.24.0 | 0.22.0 | 0.23.0 |
The two cards in aifoundry1 run different firmware, and neither matches the lab's other two cards: card 0 is one release newer than aifoundry2/3, card 1 one older. Anything measured on these cards should record which card and firmware it ran on. At 15:02 only the management path had been exercised. Since then card 1 has run the three-card check and card 0 has been found to overheat under load (update of 27 September).
2. What was done, in order
The procedure is section 1 of the troubleshooting report; every command and its output is in the transcript (docs/reports/data/2026-09-25-aifoundry1/fix-transcript.txt in the repository). Backups are in /root/et-soc1-fix-20260925/ on aifoundry1.
Every step, in order
- Checks (15:00). No process held any
/dev/et*node (fuser), the module's reference count was 0, no CI job was running, 154 MB free on/, Secure Boot off. Another user was logged in but held no card. - CI paused, backups taken (15:01). The GitHub Actions runner
actions.runner.nekkoai-hf-hackathon.aifoundry1-et-soc1was stopped.dkms status,modinfo, the three installedet-soc1.ko.zstfiles and/etc/modules-load.dwere copied to the backup directory. - DKMS source replaced (15:01).
dkms remove et-soc1/0.20.0 --all; the old source (et-platform09531e5c1, with the Makefile typo$(ET_MODULE_VERSION=)) moved to the backup directory;/usr/src/et-soc1-0.20.0refilled fromgit -C /usr/src/et-platform archive 353f20e et-driver(its Makefile reads$(ET_MODULE_VERSION),VERSION0.20.0). - Built for every kernel (15:01).
dkms add, thendkms installfor 7.0.0-30, 7.0.0-31 (running) and 7.0.0-34. Each reads[0.20.0] 47D26A305A0428B29FB7FC4, byte for byte the value on aifoundry2. - Module reloaded (15:01).
modprobe -r et_soc1 && modprobe et_soc1: the four device nodes came back (0666), both cards registered with the driver,/sys/module/et_soc1/versionreads0.20.0. The driver's remove path does not reset the cards; their firmware kept running. - Verified (15:02). The version files deviceLayer reads (
/sys/bus/pci/devices/0000:0{1,2}:00.0/driver/module/version) both read0.20.0; the management query succeeded on both cards (section 1). - Boot and CI (15:02).
/etc/modules-load.d/et_soc1.confadded so the module loads at boot, as on aifoundry2; the runner restarted and active. None of its three workflows (benchmark.yml,benchmark-board.yml,sync-huggingface.yml) builds or reloads the driver. - Space freed (15:03). Only system data that is re-created on demand:
apt-get clean(281 MB of package cache),journalctl --vacuum-size=50M(118 MB of archived journal), and ten disabled snap revisions (old copies kept for rollback)./went from 155 MB to 553 MB free. No user data was touched. - Kernel update finished (15:03).
dpkg --configure -a(exit 0): DKMS autoinstall for 7.0.0-34,update-initramfswrote/boot/initrd.img-7.0.0-34-generic(it contains the ZFS modules), andupdate-grubrewrotegrub.cfg, whose default entry is now 7.0.0-34.linux-image-7.0.0-34andlinux-modules-7.0.0-34readii,/var/lib/dpkg/updatesis empty,apt-get checkis clean.
Rollback of the driver (puts back the broken module): the commands in the troubleshooting report's section 1, with B=/root/et-soc1-fix-20260925.
3. Still open
- Card 0 overheats under load (found on 25 September after this fix; update above): 98–102 °C in short smoke launches and 115–117 °C idle afterwards. Its fan, heat sink and airflow need checking; until then it takes no sustained work.
- The disk: see section 4. After clearing caches there is 6.4 GB free; the rest of the pool is user data whose owners have to decide.
- Card et0's PCIe link still counts about one corrected receive error per second at root port 00:01.0; since 25 September its kernel log is rate-limited, and the error counters still count (troubleshooting report, section 4). It does not stop the card; a reseat at the next maintenance window is the fix.
- The next boot is untested. It will start 7.0.0-34 with a freshly written initramfs, as aifoundry2 and aifoundry3 are set to.
et_soc1for that kernel is the fixed build and loads frommodules-load.d. After the reboot:cat /sys/module/et_soc1/versionshould print0.20.0. - Different firmware on the two cards (section 1). Whether to bring them to one release is for the lab to decide.
- An older
esperantoDKMS driver is still registered for 23 kernels (not loaded). It is harmless but can be removed (dkms remove esperanto/0.20.0 --all, and/etc/udev/rules.d/50-esperanto.rules) as a separate change.
4. The disk: who should delete what
aifoundry1 has one ZFS pool, rpool, 452 GB. At 15:00 it was 96% allocated and every file system on it (/, /home, /root, /var) had only 155 MB available, because ZFS holds back the last ~3% of a pool. Only caches and logs that are re-created on demand were cleared, in two steps (15:03 and 15:11): the apt package cache (281 MB), archived journal files (118 MB), ten disabled snap revisions, root's Hugging Face download chunk cache /root/.cache/huggingface/xet (5.4 GB; the downloaded models in hub/ were kept) and root's ccache (0.5 GB). Every file system now has 6.4 GB available. Nothing in a user's home directory, and no model, checkpoint, container or project directory, was touched.
The remaining space is user data. From zfs userspace and du (sizes and dates only; no file was opened), 25 September 15:10 PDT:
Every directory, size and owner, as a table
| Where | Size | Owner (account) | What is in it | Newest file |
|---|---|---|---|---|
/home/rehan | 176 GB | rehan | models/ 123 GB: GGUF model files (the largest gemma-4-26B Q8_0 27 GB, two Qwen3.5-35B Q4_0 20–21 GB each, DeepSeek-V2-Lite Q8_0 and Q4_0 17 + 9 GB, Qwen3.6-27B 16 GB); et-jobs-deploy/ 14 GB; hf-hackathon-week2/ 11 GB; .local 8 GB | 19 Jul (models) |
/home/saqib | 135 GB | root (no saqib account; the directory name suggests Saqib) | models/ 120 GB: GGUF files, several of them 8B Llama variants at Q8_0 (Llama-3-8B-Instruct twice, once saved with ?download=true in its name, Llama-3.2-8B, llama-7b), Qwen3.6-35B, Qwen3VL-8B, rwkv7-7.2B; hacks/ 10 GB; llamaCpp/ 4.5 GB | 21 Jul |
/home/roman | 66 GB | roman | justin/ 52 GB (a Hugging Face cache with gpt-oss-20b ONNX, an ONNX Runtime bundle of 10.5 GB, lerobot-0.1.tar 8.9 GB, legacy-0.6.tar 3.1 GB); src/ 13 GB | 21 Oct 2025 |
/root | 19 GB | root (shared) | et-jobs-deploy/ 9.8 GB (GGUF test models), the Hugging Face model cache 5.4 GB, the CI runner 1.7 GB, martin_experiments/ 0.9 GB, src/ 0.7 GB | 24 Aug (cache) |
/home/sbn | 11 GB | root (no sbn account) | one ONNX Runtime bundle of 10.5 GB, the same file name as the one in /home/roman/justin | 8 Jan |
/var/lib/docker | 8.0 GB | root | 5 images, 7.8 GB of them used by no running container; 3 stopped containers (shelley, exited 5 weeks ago; two hello-world) | |
| everyone else | < 5 GB each | ruben 4.4, marin 1.7, ubergarm 1.2, jason 0.3 GB, others less |
Who should delete: rehan (123 GB of models, the only recent user among the large owners) and whoever owns /home/saqib (120 GB of models; by name and size its GGUF files include Llama-3-8B-Instruct at Q8_0 twice) account for 243 GB, more than half the pool. roman should decide on /home/roman/justin (52 GB, untouched since October 2025). The lab admin should decide on /home/sbn (11 GB, which may duplicate a file in justin/), the unused Docker images (7.8 GB, docker image prune -a after removing the stopped containers) and the GGUF copies in /root/et-jobs-deploy. Deleting only the models nobody uses would leave the machine with well over 100 GB free; a shared read-only model directory (one copy per model, readable by everyone) would stop the duplication from coming back.
Related
- What is broken on aifoundry1: the diagnosis, evidence and procedure.
- The transcript and backups:
docs/reports/data/2026-09-25-aifoundry1/in the repository;/root/et-soc1-fix-20260925/on aifoundry1.