AI Foundry lab · Fix log · 25 September 2026

aifoundry1 is fixed

Both ET-SoC-1 cards in aifoundry1 answer again. The et_soc1 kernel module was rebuilt by DKMS from the fixed driver source that was already on the machine (et-platform 353f20e), for all three kernels in /boot, and reloaded. It now reports version 0.20.0 and srcversion 47D26A305A0428B29FB7FC4, the same as aifoundry2 and aifoundry3, and a read-only management query succeeds on both cards. The interrupted 24 September kernel update was finished as well, after freeing space on the nearly full root file system. The machine was not rebooted, and no running user work was touched.

Update, 27 September 2026. Card 1 has since run real work: the whole three-card version-3 check (25–26 September, as aifoundry1-c1) and the gathers and scatters after it (E48, 26 September 06:57–09:22). Card 0 cannot hold a load: in about ten minutes of short smoke launches on 25 September its die reached 98–102 °C, and after the last one it read 115–117 °C with nothing running, drawing 66–71 W at 600 MHz, until its firmware dropped it to 300 MHz and it cooled. Its cooling (a fan, heat sink or airflow) is at fault, which its 300 MHz idle normally hides; it takes no sustained work and was left out of the check (amendment A4 of the check's record).

Cards
Both answer: dev_mngt_service -m DM_CMD_GET_MODULE_FIRMWARE_REVISIONS succeeds on card 0 and card 1; each card's management queue counted its first command.
Driver
et_soc1 [0.20.0] 47D26A305A0428B29FB7FC4 loaded now and installed for 7.0.0-30, -31 and -34; loaded at boot by /etc/modules-load.d/et_soc1.conf.
Kernel update
linux-image-7.0.0-34 fully installed: initramfs (with ZFS) written, GRUB updated; the next boot starts 7.0.0-34, whose et_soc1 is the fixed one.
Still open
The pool is 95% full: 6.4 GB free after clearing caches, the rest is user data (who should delete what); card et0's PCIe link still counts ~1 corrected error per second (its kernel log is rate-limited since 25 September); the two cards run different firmware; the next boot is untested; card 0 overheats under load (update of 27 September).

1. The cards now

A read-only management query per card, run as an ordinary user with nobody else on the cards (15:02 PDT):

aifoundry1 card 0 (01:00.0)aifoundry1 card 1 (02:00.0)aifoundry2, aifoundry3
Firmware release1.4.11.2.01.3.1
BL1 / BL20.21.20.18.00.20.0
PMIC firmware1.6.11.3.01.5.0
Master / worker / machine minion0.24.00.22.00.23.0

The two cards in aifoundry1 run different firmware, and neither matches the lab's other two cards: card 0 is one release newer than aifoundry2/3, card 1 one older. Anything measured on these cards should record which card and firmware it ran on. At 15:02 only the management path had been exercised. Since then card 1 has run the three-card check and card 0 has been found to overheat under load (update of 27 September).

2. What was done, in order

The procedure is section 1 of the troubleshooting report; every command and its output is in the transcript (docs/reports/data/2026-09-25-aifoundry1/fix-transcript.txt in the repository). Backups are in /root/et-soc1-fix-20260925/ on aifoundry1.

Every step, in order
  1. Checks (15:00). No process held any /dev/et* node (fuser), the module's reference count was 0, no CI job was running, 154 MB free on /, Secure Boot off. Another user was logged in but held no card.
  2. CI paused, backups taken (15:01). The GitHub Actions runner actions.runner.nekkoai-hf-hackathon.aifoundry1-et-soc1 was stopped. dkms status, modinfo, the three installed et-soc1.ko.zst files and /etc/modules-load.d were copied to the backup directory.
  3. DKMS source replaced (15:01). dkms remove et-soc1/0.20.0 --all; the old source (et-platform 09531e5c1, with the Makefile typo $(ET_MODULE_VERSION=)) moved to the backup directory; /usr/src/et-soc1-0.20.0 refilled from git -C /usr/src/et-platform archive 353f20e et-driver (its Makefile reads $(ET_MODULE_VERSION), VERSION 0.20.0).
  4. Built for every kernel (15:01). dkms add, then dkms install for 7.0.0-30, 7.0.0-31 (running) and 7.0.0-34. Each reads [0.20.0] 47D26A305A0428B29FB7FC4, byte for byte the value on aifoundry2.
  5. Module reloaded (15:01). modprobe -r et_soc1 && modprobe et_soc1: the four device nodes came back (0666), both cards registered with the driver, /sys/module/et_soc1/version reads 0.20.0. The driver's remove path does not reset the cards; their firmware kept running.
  6. Verified (15:02). The version files deviceLayer reads (/sys/bus/pci/devices/0000:0{1,2}:00.0/driver/module/version) both read 0.20.0; the management query succeeded on both cards (section 1).
  7. Boot and CI (15:02). /etc/modules-load.d/et_soc1.conf added so the module loads at boot, as on aifoundry2; the runner restarted and active. None of its three workflows (benchmark.yml, benchmark-board.yml, sync-huggingface.yml) builds or reloads the driver.
  8. Space freed (15:03). Only system data that is re-created on demand: apt-get clean (281 MB of package cache), journalctl --vacuum-size=50M (118 MB of archived journal), and ten disabled snap revisions (old copies kept for rollback). / went from 155 MB to 553 MB free. No user data was touched.
  9. Kernel update finished (15:03). dpkg --configure -a (exit 0): DKMS autoinstall for 7.0.0-34, update-initramfs wrote /boot/initrd.img-7.0.0-34-generic (it contains the ZFS modules), and update-grub rewrote grub.cfg, whose default entry is now 7.0.0-34. linux-image-7.0.0-34 and linux-modules-7.0.0-34 read ii, /var/lib/dpkg/updates is empty, apt-get check is clean.

Rollback of the driver (puts back the broken module): the commands in the troubleshooting report's section 1, with B=/root/et-soc1-fix-20260925.

3. Still open

4. The disk: who should delete what

aifoundry1 has one ZFS pool, rpool, 452 GB. At 15:00 it was 96% allocated and every file system on it (/, /home, /root, /var) had only 155 MB available, because ZFS holds back the last ~3% of a pool. Only caches and logs that are re-created on demand were cleared, in two steps (15:03 and 15:11): the apt package cache (281 MB), archived journal files (118 MB), ten disabled snap revisions, root's Hugging Face download chunk cache /root/.cache/huggingface/xet (5.4 GB; the downloaded models in hub/ were kept) and root's ccache (0.5 GB). Every file system now has 6.4 GB available. Nothing in a user's home directory, and no model, checkpoint, container or project directory, was touched.

The remaining space is user data. From zfs userspace and du (sizes and dates only; no file was opened), 25 September 15:10 PDT:

Every directory, size and owner, as a table
WhereSizeOwner (account)What is in itNewest file
/home/rehan176 GBrehanmodels/ 123 GB: GGUF model files (the largest gemma-4-26B Q8_0 27 GB, two Qwen3.5-35B Q4_0 20–21 GB each, DeepSeek-V2-Lite Q8_0 and Q4_0 17 + 9 GB, Qwen3.6-27B 16 GB); et-jobs-deploy/ 14 GB; hf-hackathon-week2/ 11 GB; .local 8 GB19 Jul (models)
/home/saqib135 GBroot (no saqib account; the directory name suggests Saqib)models/ 120 GB: GGUF files, several of them 8B Llama variants at Q8_0 (Llama-3-8B-Instruct twice, once saved with ?download=true in its name, Llama-3.2-8B, llama-7b), Qwen3.6-35B, Qwen3VL-8B, rwkv7-7.2B; hacks/ 10 GB; llamaCpp/ 4.5 GB21 Jul
/home/roman66 GBromanjustin/ 52 GB (a Hugging Face cache with gpt-oss-20b ONNX, an ONNX Runtime bundle of 10.5 GB, lerobot-0.1.tar 8.9 GB, legacy-0.6.tar 3.1 GB); src/ 13 GB21 Oct 2025
/root19 GBroot (shared)et-jobs-deploy/ 9.8 GB (GGUF test models), the Hugging Face model cache 5.4 GB, the CI runner 1.7 GB, martin_experiments/ 0.9 GB, src/ 0.7 GB24 Aug (cache)
/home/sbn11 GBroot (no sbn account)one ONNX Runtime bundle of 10.5 GB, the same file name as the one in /home/roman/justin8 Jan
/var/lib/docker8.0 GBroot5 images, 7.8 GB of them used by no running container; 3 stopped containers (shelley, exited 5 weeks ago; two hello-world)
everyone else< 5 GB eachruben 4.4, marin 1.7, ubergarm 1.2, jason 0.3 GB, others less

Who should delete: rehan (123 GB of models, the only recent user among the large owners) and whoever owns /home/saqib (120 GB of models; by name and size its GGUF files include Llama-3-8B-Instruct at Q8_0 twice) account for 243 GB, more than half the pool. roman should decide on /home/roman/justin (52 GB, untouched since October 2025). The lab admin should decide on /home/sbn (11 GB, which may duplicate a file in justin/), the unused Docker images (7.8 GB, docker image prune -a after removing the stopped containers) and the GGUF copies in /root/et-jobs-deploy. Deleting only the models nobody uses would leave the machine with well over 100 GB free; a shared read-only model directory (one copy per model, readable by everyone) would stop the duplication from coming back.