Skip to content

Detect Intel Arc Battlemage GPUs and fix the SYCL stack path - #7443

Open
SergiioB wants to merge 19 commits into
Osmantic:mainfrom
SergiioB:intel-battlemage-detection
Open

SergiioB wants to merge 19 commits into
Osmantic:mainfrom
SergiioB:intel-battlemage-detection

Conversation

@SergiioB

@SergiioB SergiioB commented Oct 7, 2026 •

Copy link
Copy Markdown

Summary

detect_gpu() silently misses Intel Battlemage cards, so the installer falls back to CPU-only mode on current Arc hardware. Verified on a real dual-GPU Battlemage G31 host (2 x 0xe223, 24 GB each): detection reported "No GPU detected" and selected CPU tier T1.

Root causes fixed in this PR:

  • Device IDs: the Arc match table only covered Alchemist/DG2 (0x56xx/0x569x). Battlemage uses 0xe2xx.

  • lspci gate: detection required an lspci name containing Arc, but Battlemage enumerates as Battlemage G31 [Intel Graphics], and lspci may not be installed at all. Detection now relies on sysfs vendor/device IDs only; lspci is used opportunistically for a display name.

  • Multi-GPU: the Intel path returned on the first matching card (GPU_COUNT=1, single-card VRAM). It now counts all matching cards and sums lmem_total_bytes.

  • scripts/detect-hardware.sh had no Intel path at all (Jetson/NVIDIA/AMD/Apple only). Added sysfs detection incl. lmem_total_bytes, xe/i915 driver state, multi-GPU counting, and ARC/ARC_LITE tiers + descriptions + model picks.

  • gpu-database.json had zero Intel entries (no known_gpus, no heuristic classes, no bandwidth defaults), so classification always resolved to cpu. Added B60/B580/B570/BMG-G31, A770/A750/A580/A380, intel_arc/intel_arc_lite heuristic classes, and an intel bandwidth table + sycl default.

  • classify-hardware.sh lacked intel/sycl in OVERLAY_MAP, so a classified Arc host still produced no GPU overlay.

  • Compose overlay order: both resolvers preferred docker-compose.arc.yml (oneAPI 2025.0 build-from-source, ~10-20 min, toolchain predates Battlemage) over docker-compose.intel.yml (the pinned prebuilt server-intel image that phase 08 actually pulls). Preference flipped; arc.yml remains the fallback.

Test plan

  • bash tests/test-intel-arc-detection.sh — new mock-sysfs suite: dual Battlemage, Alchemist, missing lspci/product_name/lmem_total_bytes, iGPU rejection. All pass.

  • bash tests/test-amd-igpu-dgpu-selection.sh — unchanged, still passes.

  • bash tests/test-tier-map.sh — 165 passed, 0 failed.

  • classify-hardware.sh on 0xe223/24 GB → ARC, backend intel, overlays base+intel.

  • Real-hardware install.sh on the dual-BMG host (installer re-run pending; host currently mid-install from the CPU path).

Audit repair: cap the B570 PCI aperture fallback at its actual 10 GiB capacity in both detectors, and keep optional lspci name lookup failures from aborting strict-shell detection. Regression cases cover both defects. Portable checks pass; real Battlemage GPU validation remains required before merge.

Integrated audit follow-up

Audited candidate f0ff0266787ebd38d12e3a2064a60d4e82fccafc integrates the reviewed software batch while keeping this hardware PR independent of the other hardware-gated changes. Clean merge including prior d5e35d8 repair: B570 fallback cap and nonfatal optional lspci naming. No new source defect found on software base.

Fresh focused validation:

  • 9 Intel cases,6 AMD cases,165 tier-map assertions passed

Merge remains held for actual hardware/setup acceptance: Actual Battlemage host installer using repaired exact candidate: correct GPU count/memory/driver readiness, SYCL pinned overlay/image, usable inference/model fit, correct multi-GPU placement and fallback behavior. Portable sysfs fixtures do not prove any real Arc capability.

Source, fixture and CI results do not waive this gate.

Validated main 39fe4a69dc40ded4a7d529a2ad5f55267391a67d is now an ancestor. This final alignment changes no source files or tree relative to audited candidate 08098df948e96dd0aa65355523f2c9f250ca6f23; CI is rerunning on the published head.

sergio and others added 8 commits October 7, 2026 08:54
detect_gpu() only matched Alchemist/DG2 device IDs (0x56xx/0x569x) and
additionally gated on an lspci marketing string ("Intel ... Arc"), so
Battlemage cards (0xe2xx, e.g. B580/B570/Arc Pro B60; lspci reports them as
"Battlemage G31 [Intel Graphics]") were silently missed and the installer
fell back to CPU-only mode.

Verified on real hardware: a dual Battlemage G31 host (2 x 0xe223, 24 GB
each) was detected as "No GPU detected -> CPU-only tier T1".

Changes:

- installers/lib/detection.sh: detect Arc purely from sysfs vendor/device
  IDs (lspci only names the card now), add the Battlemage 0xe2xx family,
  and count/sum VRAM across all matching cards instead of returning on the
  first one (single-card behavior unchanged).
- scripts/detect-hardware.sh: add Intel Arc detection to the diagnostics
  tool (previously Jetson/NVIDIA/AMD/Apple only), including lmem VRAM,
  xe/i915 driver state, multi-GPU count, and ARC/ARC_LITE tier mapping.
- config/gpu-database.json: add Intel known_gpus (B60/B580/B570/BMG-G31,
  A770/A750/A580/A380), intel heuristic classes (>=10 GB -> ARC,
  4-10 GB -> ARC_LITE), an intel bandwidth table, and an sycl default.
- scripts/classify-hardware.sh: map intel/sycl backends to
  docker-compose.intel.yml and add sycl to the bandwidth default table.
- resolve-compose-stack.sh + compose-select.sh: prefer the pinned prebuilt
  image overlay (docker-compose.intel.yml, same image phase 08 pulls) over
  docker-compose.arc.yml (oneAPI 2025.0 source build; its toolchain predates
  Battlemage support). arc.yml remains as fallback.
- tests/test-intel-arc-detection.sh: mock-sysfs tests covering dual
  Battlemage, Alchemist, missing lspci/product_name/lmem, and rejection of
  integrated iGPUs. Registered in tests/ci-suite.txt.
…_total_bytes is absent

i915 exposes lmem_total_bytes in sysfs, but the xe driver — which binds
Battlemage by default on current kernels — has no per-device VRAM file
(verified on kernel 7.0, dual 0xe223). Fall back to the largest PCI BAR
aperture from the device's resource file, which is the local-memory
window (32 GiB BAR on a 24 GB Arc Pro B60). A card with neither source
still detects with VRAM=0.
lsmod lives in sbin, which user-level PATHs often lack (hit on the
Battlemage test host), leaving driver_loaded=false while xe is bound.
config/backends/ had amd/apple/cpu/nvidia but no intel.json, so the
capability path logged 'Could not load backend contract for intel' and
fell back to defaults. Uses the same pinned ghcr.io/ggml-org/llama.cpp
server-intel image that docker-compose.intel.yml and phase 08 pull.
The Arc validation grep chain ran inside a command substitution without a
guard, so on Battlemage (lspci name 'Battlemage G31 [Intel Graphics]' — no
'Arc'/'B###' token) the pipeline failed and set -e killed phase 02 with no
error output. Guard with || true and match the Battlemage name.

Also adds config/backends/intel.json, previously missing: the capability
path logged 'Could not load backend contract for intel'.
Phase 02 only built GPU_TOPOLOGY_JSON for NVIDIA and AMD; on a multi-Arc
host (verified: 2x BMG-G31) phase 03's assign_gpus.py got an empty
topology and aborted with 'ERROR: no GPUs found in topology'.

Adds detect_intel_topo() (same gpus[]/links[] schema as the AMD path;
links stay empty since SYCL has no P2P rank and llama.cpp picks devices
by ONEAPI_DEVICE_SELECTOR index) and calls it when GPU_BACKEND=intel and
count>1. Shared the per-card VRAM reader (lmem_total_bytes, else largest
PCI BAR) and the Arc device-ID predicate as lib functions so the topology
scan and detect_gpu can't drift apart.
The first version assigned GPU_TOPOLOGY_JSON before the variable's '{}'
initialization later in the phase, so the topology was discarded and
phase 03 still saw gpu_count=0.
…r minimums

- render-runtime-configs.py rejected --gpu-backend sycl (argparse choices
  lacked intel/sycl), aborting phase 06 on every Arc host.
- preflight-engine had no ARC/ARC_LITE entries in the tier rank, RAM, or
  disk maps, so an Arc host inherited the tier-2 50GB disk floor even
  though the SYCL model is ~3-6GB. ARC -> 16GB RAM / 30GB disk, ARC_LITE
  -> 8GB / 20GB.
- Preflight treated intel/sycl as 'unknown backend' warning; now a pass.

Also scale the xe BAR-aperture VRAM fallback by 3/4: the aperture is the
next power of two above real VRAM (32 GiB BAR on a 24 GB Arc Pro B60), so
the raw value over-reports and inflates model-fit checks.
@gabsprogrammer

Copy link
Copy Markdown
Member

Thanks for the PR; if possible, please get all the checks to green so we can merge it.

The awk "%.1f" format always emitted a decimal (24.0); assign_gpus.py
consumers and the detection test expect integral GB values as integers.
Print %d when the GiB value is whole, %.1f otherwise.
@SergiioB

SergiioB commented Oct 8, 2026

Copy link
Copy Markdown
Author

Fixed the failing check.

Root cause: detect_intel_topo() formatted per-card VRAM with awk '%.1f', so the topology JSON emitted memory_gb: 24.0 — the test (and assign_gpus.py consumers) expect 24 for whole-GiB cards.

Fix (7c92215): emit %d when the GiB value is integral, %.1f otherwise — integers for 24/32 GB cards, one decimal retained for fractional sizes (e.g. 11.5 GB).

Verified locally: tests/test-intel-arc-detection.sh all 7 assertions pass; full tests/run-ci-suite.sh run: 250/252 (the two pixel-extension failures are a local umask artifact — group-writable checkout trips owned_regular's 0o022 bit check; they pass on the runner). assign_gpus.py reads memory_gb via float() so both forms are safe downstream.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants