Repository navigation
Conversation
detect_gpu() only matched Alchemist/DG2 device IDs (0x56xx/0x569x) and
additionally gated on an lspci marketing string ("Intel ... Arc"), so
Battlemage cards (0xe2xx, e.g. B580/B570/Arc Pro B60; lspci reports them as
"Battlemage G31 [Intel Graphics]") were silently missed and the installer
fell back to CPU-only mode.
Verified on real hardware: a dual Battlemage G31 host (2 x 0xe223, 24 GB
each) was detected as "No GPU detected -> CPU-only tier T1".
Changes:
- installers/lib/detection.sh: detect Arc purely from sysfs vendor/device
IDs (lspci only names the card now), add the Battlemage 0xe2xx family,
and count/sum VRAM across all matching cards instead of returning on the
first one (single-card behavior unchanged).
- scripts/detect-hardware.sh: add Intel Arc detection to the diagnostics
tool (previously Jetson/NVIDIA/AMD/Apple only), including lmem VRAM,
xe/i915 driver state, multi-GPU count, and ARC/ARC_LITE tier mapping.
- config/gpu-database.json: add Intel known_gpus (B60/B580/B570/BMG-G31,
A770/A750/A580/A380), intel heuristic classes (>=10 GB -> ARC,
4-10 GB -> ARC_LITE), an intel bandwidth table, and an sycl default.
- scripts/classify-hardware.sh: map intel/sycl backends to
docker-compose.intel.yml and add sycl to the bandwidth default table.
- resolve-compose-stack.sh + compose-select.sh: prefer the pinned prebuilt
image overlay (docker-compose.intel.yml, same image phase 08 pulls) over
docker-compose.arc.yml (oneAPI 2025.0 source build; its toolchain predates
Battlemage support). arc.yml remains as fallback.
- tests/test-intel-arc-detection.sh: mock-sysfs tests covering dual
Battlemage, Alchemist, missing lspci/product_name/lmem, and rejection of
integrated iGPUs. Registered in tests/ci-suite.txt.
…_total_bytes is absent i915 exposes lmem_total_bytes in sysfs, but the xe driver — which binds Battlemage by default on current kernels — has no per-device VRAM file (verified on kernel 7.0, dual 0xe223). Fall back to the largest PCI BAR aperture from the device's resource file, which is the local-memory window (32 GiB BAR on a 24 GB Arc Pro B60). A card with neither source still detects with VRAM=0.
lsmod lives in sbin, which user-level PATHs often lack (hit on the Battlemage test host), leaving driver_loaded=false while xe is bound.
config/backends/ had amd/apple/cpu/nvidia but no intel.json, so the capability path logged 'Could not load backend contract for intel' and fell back to defaults. Uses the same pinned ghcr.io/ggml-org/llama.cpp server-intel image that docker-compose.intel.yml and phase 08 pull.
The Arc validation grep chain ran inside a command substitution without a guard, so on Battlemage (lspci name 'Battlemage G31 [Intel Graphics]' — no 'Arc'/'B###' token) the pipeline failed and set -e killed phase 02 with no error output. Guard with || true and match the Battlemage name. Also adds config/backends/intel.json, previously missing: the capability path logged 'Could not load backend contract for intel'.
Phase 02 only built GPU_TOPOLOGY_JSON for NVIDIA and AMD; on a multi-Arc host (verified: 2x BMG-G31) phase 03's assign_gpus.py got an empty topology and aborted with 'ERROR: no GPUs found in topology'. Adds detect_intel_topo() (same gpus[]/links[] schema as the AMD path; links stay empty since SYCL has no P2P rank and llama.cpp picks devices by ONEAPI_DEVICE_SELECTOR index) and calls it when GPU_BACKEND=intel and count>1. Shared the per-card VRAM reader (lmem_total_bytes, else largest PCI BAR) and the Arc device-ID predicate as lib functions so the topology scan and detect_gpu can't drift apart.
The first version assigned GPU_TOPOLOGY_JSON before the variable's '{}'
initialization later in the phase, so the topology was discarded and
phase 03 still saw gpu_count=0.
…r minimums - render-runtime-configs.py rejected --gpu-backend sycl (argparse choices lacked intel/sycl), aborting phase 06 on every Arc host. - preflight-engine had no ARC/ARC_LITE entries in the tier rank, RAM, or disk maps, so an Arc host inherited the tier-2 50GB disk floor even though the SYCL model is ~3-6GB. ARC -> 16GB RAM / 30GB disk, ARC_LITE -> 8GB / 20GB. - Preflight treated intel/sycl as 'unknown backend' warning; now a pass. Also scale the xe BAR-aperture VRAM fallback by 3/4: the aperture is the next power of two above real VRAM (32 GiB BAR on a 24 GB Arc Pro B60), so the raw value over-reports and inflates model-fit checks.
|
Thanks for the PR; if possible, please get all the checks to green so we can merge it. |
The awk "%.1f" format always emitted a decimal (24.0); assign_gpus.py consumers and the detection test expect integral GB values as integers. Print %d when the GiB value is whole, %.1f otherwise.
|
Fixed the failing check. Root cause: Fix (7c92215): emit Verified locally: |
…tch23-hw7443-current
Summary
detect_gpu()silently misses Intel Battlemage cards, so the installer falls back to CPU-only mode on current Arc hardware. Verified on a real dual-GPU Battlemage G31 host (2 x 0xe223, 24 GB each): detection reported "No GPU detected" and selected CPU tier T1.Root causes fixed in this PR:
Device IDs: the Arc match table only covered Alchemist/DG2 (
0x56xx/0x569x). Battlemage uses0xe2xx.lspci gate: detection required an lspci name containing
Arc, but Battlemage enumerates asBattlemage G31 [Intel Graphics], and lspci may not be installed at all. Detection now relies on sysfs vendor/device IDs only; lspci is used opportunistically for a display name.Multi-GPU: the Intel path returned on the first matching card (
GPU_COUNT=1, single-card VRAM). It now counts all matching cards and sumslmem_total_bytes.scripts/detect-hardware.shhad no Intel path at all (Jetson/NVIDIA/AMD/Apple only). Added sysfs detection incl.lmem_total_bytes,xe/i915driver state, multi-GPU counting, andARC/ARC_LITEtiers + descriptions + model picks.gpu-database.jsonhad zero Intel entries (noknown_gpus, no heuristic classes, no bandwidth defaults), so classification always resolved tocpu. Added B60/B580/B570/BMG-G31, A770/A750/A580/A380,intel_arc/intel_arc_liteheuristic classes, and anintelbandwidth table +sycldefault.classify-hardware.shlackedintel/syclinOVERLAY_MAP, so a classified Arc host still produced no GPU overlay.Compose overlay order: both resolvers preferred
docker-compose.arc.yml(oneAPI 2025.0 build-from-source, ~10-20 min, toolchain predates Battlemage) overdocker-compose.intel.yml(the pinned prebuiltserver-intelimage that phase 08 actually pulls). Preference flipped; arc.yml remains the fallback.Test plan
bash tests/test-intel-arc-detection.sh— new mock-sysfs suite: dual Battlemage, Alchemist, missing lspci/product_name/lmem_total_bytes, iGPU rejection. All pass.bash tests/test-amd-igpu-dgpu-selection.sh— unchanged, still passes.bash tests/test-tier-map.sh— 165 passed, 0 failed.classify-hardware.shon0xe223/24 GB →ARC, backendintel, overlaysbase+intel.Real-hardware
install.shon the dual-BMG host (installer re-run pending; host currently mid-install from the CPU path).Audit repair: cap the B570 PCI aperture fallback at its actual 10 GiB capacity in both detectors, and keep optional lspci name lookup failures from aborting strict-shell detection. Regression cases cover both defects. Portable checks pass; real Battlemage GPU validation remains required before merge.
Integrated audit follow-up
Audited candidate
f0ff0266787ebd38d12e3a2064a60d4e82fccafcintegrates the reviewed software batch while keeping this hardware PR independent of the other hardware-gated changes. Clean merge including prior d5e35d8 repair: B570 fallback cap and nonfatal optional lspci naming. No new source defect found on software base.Fresh focused validation:
Merge remains held for actual hardware/setup acceptance: Actual Battlemage host installer using repaired exact candidate: correct GPU count/memory/driver readiness, SYCL pinned overlay/image, usable inference/model fit, correct multi-GPU placement and fallback behavior. Portable sysfs fixtures do not prove any real Arc capability.
Source, fixture and CI results do not waive this gate.
Validated main
39fe4a69dc40ded4a7d529a2ad5f55267391a67dis now an ancestor. This final alignment changes no source files or tree relative to audited candidate08098df948e96dd0aa65355523f2c9f250ca6f23; CI is rerunning on the published head.