Skip to content

[A3, Deploy] YOLO-Masted-Edge Backend Compatibility Update upon v26.08 #232

Description

@skywalker-lt

Project Repo: yolo-master-edge

Updates the backend compatibility estimation table in the v26.08 Release Report. Cells marked ▲ were upgraded from the table's estimate on the strength of a full-model measurement on the yolo-master-edge runtime (branch dev/edge-compatibility-upgrade); each upgraded cell is individually justified in the "Upgraded cells" section below. Rows marked (new) did not exist in the original table.

Architecture ONNX TensorRT NCNN MNN Core ML
Dense (v0.1) ✅ ✅ ✅ ✅ ✅
YOLO26 native ✅ ✅ ✅ ⁽¹⁾ ✅ ✅
ES-MoE ✅ ✅▲ ⁽²⁾ ✅ ✅ ✅▲ ⁽³⁾
Shared-MoE ✅ △ △▲ ⁽⁴⁾ △▲ ⁽⁴⁾ △▲ ⁽⁴⁾
MoA (dense inference) ✅ △ ⁽⁵⁾ ✅▲ ⁽⁶⁾ ✅▲ ⁽⁶⁾ ✅▲ ⁽⁶⁾
MoA (sparse inference) ⊘ ⁽⁷⁾ ⊘ ⊘ ⁽⁷⁾ ⊘ ⊘
MoT (dense inference) ✅ △ ⁽⁵⁾ ✅▲ ⁽⁶⁾ ✅▲ ⁽⁶⁾ ✅▲ ⁽⁶⁾
MoT (sparse inference) ⊘ ⁽⁷⁾ ⊘ ⊘ ⁽⁷⁾ ⊘ ⊘
MoA-MoT hybrid (new) ✅ △ ⁽⁵⁾ ✅▲ ⁽⁶⁾ △ ⁽⁸⁾ ✅▲ ⁽⁶⁾
MoLoRA (merged) (new) ✅ △ ⁽⁵⁾ ✅▲ ⁽⁶⁾ ✅▲ ⁽⁶⁾ ✅▲ ⁽⁶⁾
MoLoRA (routing-preserved) (new) ✅ — — ⁽⁹⁾ △ ⁽⁸⁾ ✅▲ ⁽¹⁰⁾
Latent Mixture (dense) ✅ △ — — —
Latent Mixture (sparse) ⊘ ⊘ ⊘ ⊘ ⊘
MultiTask △ — — — —

✅ validated e2e on the edge runtime (raw-tensor parity vs the reference plus CLI box-level checks). △ partial: runtime-ready or block-proven, but the specific full-model conversion or engine build has not been run; no known blocker. - not attempted and not currently supported. ⊘ not exportable by design (data-dependent sparse dispatch cannot be captured in a static graph); see note 7 for what IS preserved. ▲ upgraded from the release table's estimate, evidence below.

Details of models x backends now supported:

Combination Release Table Now Reasoning and evidence
ES-MoE x TensorRT △ ✅ Not an estimate anymore: the EsMoE-N VisDrone FP16 engine has run on a Jetson Orin since edge v1.1.0. Measured on-device: 17.6 FPS model-only at 640 (28 ms steady per frame), det counts matching the ONNX backend on the same images, full 18-test CLI battery passing against the TRT runner.
ES-MoE x Core ML - ✅ The v0.1-seg-N mlpackage contains ES-MoE blocks and has shipped inside the macOS app since v1.1.0 (converted with an aux-loss silencer for the in-place telemetry that breaks the coremltools frontend; class-count and output-contract checks at conversion, live inference in the app).
Shared-MoE x NCNN / MNN / Core ML - △ Block-level proof, deliberately not claimed as full support: the shared-inverted expert group IS the E=16 expert bank inside the v0_10 mixture trunks that pass full parity on all three backends (see the MoA/MoT rows). What has not been run is a standalone Shared-MoE architecture export, so these cells stop at partial.
MoA (dense) x NCNN - ✅ Full-model conversion via pnnx with trace-time rewrites (sparse-routing arithmetic emulation, batch-axis-safe window attention with windows folded into SDPA heads). Raw det-tensor max diff vs ONNX Runtime: 1.2e-4; CLI det counts within +/-1; verified on Linux x64 AND Jetson Orin (first NCNN enablement on that platform, 18/18 battery on-device). Zero custom runtime layers; shipped binaries load it unchanged.
MoA (dense) x MNN - ✅ Full-model MNNConvert path (IsInf/IsNaN guards lowered at export). Raw parity 1.8e-4 vs ONNX Runtime, CLI det parity, verified on Linux x64 and Jetson Orin (on-device MNN 3.6.1 build).
MoA (dense) x Core ML - ✅ fp32 mlprogram conversion with rank-5 window-partition rewrites (Core ML rank limit) and a scalar nan_to_num fix (macOS compiler rejects rank-0 non_zero). On-device Apple Silicon parity: raw 1.5e-4, box match 1.000.
MoT (dense) x NCNN - ✅ Same pipeline as MoA plus: deformable expert runs via ncnn GridSample (grid_sample unrolled per attention head to respect the missing batch axis), Swin shift via slicing. Raw parity 3.1e-4, CLI dets within +/-1, verified on x64 and Orin.
MoT (dense) x MNN - ✅ Raw parity 3.9e-3 vs ONNX Runtime (within the 5e-3 fp32 gate), CLI det parity; x64 and Orin.
MoT (dense) x Core ML - ✅ grid_sample(align_corners=True) converts to MIL resample and EXECUTES correctly on the Core ML runtime (this was a flagged risk in the release's capability matrix). On-device raw parity 2.1e-4, box match 1.000.
MoA-MoT x NCNN / Core ML (no row) ✅ The combined trunk converts and passes with the same pipelines: NCNN raw 3.1e-4 (x64 + Orin), Core ML raw 1.5e-4 with box match 1.000 on-device.
MoLoRA (merged) x NCNN / MNN / Core ML (no row) ✅ Merged MoLoRA collapses to the base architecture (after the grouped-conv merge fix), so it inherits the base pipelines: NCNN 1.2e-4, MNN 1.8e-4, Core ML 1.5e-4 with box match 1.000, all measured.
MoLoRA (routing-preserved) x Core ML (no row) ✅ Exceeds release report's own support envelope: the YOLO-Master exporter hard-blocks molora_export_mode=routing_preserved for Core ML, but a direct TorchScript trace converts (topk + one_hot + dense accumulate are all Core ML-supported) and passes on-device parity: raw 2.1e-4, box match 1.000.

Notes

  1. YOLO26 on NCNN ships in the anchors layout: the exporter auto-disables the end2end head for NCNN (no in-graph top-k). The edge runtime now decodes BOTH layouts ([1, 4+nc, anchors] and end2end [1, max_det, 6]) on every backend, selected by model metadata with a shape heuristic fallback.
  2. See "ES-MoE x TensorRT" above.
  3. See "ES-MoE x Core ML" above.
  4. See "Shared-MoE" above.
  5. TensorRT for the mixture families: the edge TRT runtime is layout-complete (both det layouts) and tested with ES-MoE engines, but mixture .engine builds have not been attempted. Release v26.08 ships no trained mixture weights, so an engine build would validate machinery only. No known blocker beyond the Release's estimate.
  6. Per-cell evidence in the "Upgraded cells" section. Common to all: everything is achieved at conversion time (trace-level rewrites in the edge export drivers); the shipped edge binaries run these models without modification, and zero custom runtime layers were required.
  7. What "sparse" means after this. The ⊘ cells remain correct in the release table's sense: no static graph performs input-dependent sparse compute (all experts are evaluated). However, the NCNN/ONNX exports of the v0_10 trunks now preserve exact sparse ROUTING SEMANTICS: the top-2 selection plus complexity gate is reproduced arithmetically (max/clamp/ceil emulation, full-width weights with zeros at unselected experts), so exported outputs match the sparse eager model, not a dense softmax approximation. One documented deviation: exact score ties break toward the lower expert index (torch.topk tie order is implementation-arbitrary). Compute cost is dense; accuracy semantics are sparse.
  8. MNN gaps for MoA-MoT and routing-preserved MoLoRA are conversion-not-run, not blockers: the required op set is identical to its validated siblings.
  9. Routing-preserved MoLoRA on ncnn is declined by policy: the merged path collapses to the base architecture, runs everywhere, and is this repo's own recommended deployment route.
  10. See "MoLoRA (routing-preserved) x Core ML" above.

Notes

The YOLO-Master-Edge Release v1.1.0 & v1.1.1 binaries do not support the compatibility upgrades above yet. And the updated source code lives under the edge repo's dev/edge-compatibility-upgrade branch. We are currently building the new bundles and will release them together with the extended task support (pose, OBB, etc.) and INT8 support in v1.2.0 later this month.


Linked to: #51, #227

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions