Skip to content

feat(extensions): add managed Laya decisions for Portal - #7439

Merged
Lightheartdevs merged 27 commits into
feat/portal-laya-integrationfrom
feat/portal-laya-service
Oct 8, 2026
Merged

Lightheartdevs merged 27 commits into
feat/portal-laya-integrationfrom
feat/portal-laya-service

Conversation

@gabsprogrammer

@gabsprogrammer gabsprogrammer commented Oct 6, 2026 •

Copy link
Copy Markdown
Member

Adds Laya 0.3.28 as an optional managed Portal extension, using the tools in #7437. It provides classification, ordered scores and yes/no estimates while preserving the selected chat model and its settings. The recipe pins Python and all dependency hashes, authenticates a dedicated loopback connection and verifies actual inference on all three checkpoints before readiness.

CPU is supported on the qualified platforms; NVIDIA hosts can select a CUDA 13.0 image. Automatic selection leaves 2 GiB free headroom and caps Torch allocations at the smallest of 4 GiB, 25% of total VRAM and the available budget. One GPU checkpoint remains resident and unloads after 60 idle seconds. Missing CUDA, insufficient startup budget or startup allocation exhaustion fall back to CPU; explicit cuda requires GPU startup. The allocator cap excludes driver/context memory and cannot reserve memory against other processes. AMD and Apple hosts currently use CPU.

Updates rebuild local sources and use separate CPU/CUDA image identities. This fixes an observed update that built CUDA but retained the running CPU container under a shared image tag. The regression starts with stale CPU bytes under the CUDA tag, applies the shipped NVIDIA overlay through ordinary Compose up, and checks actual Torch/inference without force-recreate.

Validation at ed86de990:

  • Current-head AMD64 and ARM64 container qualification passed. The same CUDA image without GPU access passed real CPU fallback. Local checks passed: 588 curated-recipe tests, five recipe tests, seven runtime tests, strict extension audit, dependency-pin policy, Dashboard staging and build-context materialization.
  • The normal installer exited 0. The managed Dashboard update activated ods/laya:0.3.28-cuda130-v1, image 2214eebdff7792ad226d2f706c63f236858eb37d96e51d0309b01e37ed970024. Actual installed adapter tests passed authentication, full-context handling and all three checkpoints on RTX 5080 CUDA. All ten local services remained healthy; the idle health readback showed zero loaded checkpoints after the 60-second interval.
  • With two CPUs/two threads in both configurations, the fixed 128-news classifier took a warm median 0.391s on GPU vs 72.977s on CPU, both 118/128 correct. Portuguese sentiment took 0.100s vs 11.512s, both 60/128 correct. Initial and three warm trials are retained. These are classifier timings, excluding conversation/file handling and downloads.

The full installed Portal comparison used unchanged Qwen3.5 9B with all 128 complete news texts supplied to both methods and no expected labels in their input. All four predeclared runs are retained:

Method and run End-to-end seconds Correct / 128 Saved rows
Laya 1 13.729 118 128, ordered
Model alone 1 43.946 97 119, incomplete
Model alone 2 39.973 107 128, ordered
Laya 2 14.426 118 128, ordered

The slower Laya run was 2.77x faster than the complete baseline, with +8.59 percentage points accuracy. The incomplete baseline is a retained integrity failure, with missing IDs counted as wrong in the displayed secondary score. This is a small observed workload-specific gain, not general reliability proof. An earlier file-only baseline failed to read all source texts and remains disqualified; supplying full input for this comparison does not fix automatic file pagination.

Limits remain explicit: earlier bilingual and Portuguese sentiment tasks underperformed the main model; semantic review and failure disclosure were inconsistent. Public benchmark training overlap is unknown. ARM64 GPU hardware and physical competing-memory pressure were not tested; low-memory selection and allocation failure have unit coverage. A cold-cache update hit the existing 600-second start limit after slow CUDA dependency downloads; normal rollback restored the previous definition, and a later cached update passed. No timeout or quality assertion was relaxed.

CI follow-up at 09cd72ea5 integrates the validated #7437 base and fixes the extended suite's 45-minute prerequisite stall. Existing jq and ShellCheck are verified and reused; only missing tools need APT. Missing-tool installs have bounded transfer/command deadlines, strict refresh errors and visible logs, with no signature or test checks disabled. Five setup regression tests cover preinstalled, missing, broken and failed-install cases.

GitHub validation: 63/63 checks passed, including x86-64 and ARM64 Laya container qualification. Extended contract suite prepared its tools in 6 seconds and passed all 251 tests. The complete matrix and remaining project checks also passed. This PR remains stacked on #7437; no GitHub PR merge was performed.

Automatic-selection follow-up at 82fb49ef3:

Automatic subtask selection now considers repeated semantic judgments inside larger website/catalog/report tasks, while skipping ordinary replies, obvious single labels, CSS/code changes and deterministic calculations. Only enabled extensions advertise the helper; opt-outs and execution-time revocation remain enforced. No keyword router or extra per-turn inference was added. Questions describe requested semantic outputs, not file formats, identity, order or build success. Invalid question errors identify the offending field so repairs can preserve unrelated dataset paths.

Validation on October 7: 132/132 focused WSL Node tests passed. The new regression verifies invalid-question rejection before inference/writes, useful field diagnostics, fresh admission and unchanged source paths on repair. A reusable installed-Portal selection runner records each declared attempt and distinguishes attempted calls, saved receipts and task termination from actual artifact correctness. The normal installer exited 0; installed/configured plugin hashes match the candidate, with Qwen3.5 9B and its 65,536-token context unchanged.

Natural-request evidence (no Laya keyword in positive prompts): baseline selection passed 7/9 cases, candidate 9/9, then 2/2 held-out wordings. A repeated editorial request after lifecycle recovery skipped Laya despite the enabled hint being present in the actual model prompt, so candidate observations are 11/12, not universal routing success. All six initial negative/opt-out cases avoided Laya. A separate disabled-extension request also made zero Laya calls and contained no enabled hint. Schema-repair attempts remain recorded; one held-out call added redundant double classification. Successful reports independently preserved IDs/order and source hashes. These small, mostly single-run observations are not a latency or accuracy benchmark.

End-to-end generated-site qualification did not pass: the first website had a JavaScript quote error; the held-out website rendered 32 cards and filtered topics, but search raised a variable-initialization error and only 20/32 original texts survived the model's manual transcription exactly. Both baseline and candidate arithmetic fixtures also contained incorrect totals. These are preserved main-model delivery failures, not relabeled Laya-selection successes or fixed by the classifier. Laya's verified saved report does not certify downstream generated code.

Lifecycle qualification found the local catalog still carried an older recipe than the reviewed service branch. The reviewed recipe source was synchronized with a retained backup; the normal authenticated update preserved the disabled state and rollback, then normal enable succeeded. No receipt or authorization control was bypassed. Laya is active again, real installed English-adapter inference passed on CUDA, and all ten services are healthy. Current image: sha256:72843a0ddf3ef8223556a2532533b75edd917d270bc14e7dc268253617daf59e.

Latest-head GitHub validation at 82fb49ef3: 63/63 checks passed, with no pending, failed or skipped checks. The PR has no merge conflicts. Main requires two approving reviews; no GitHub PR merge or force-push was performed. Qualification grading also matched native result digests across 22 retained cases and passed five synthetic evidence checks, including invalid and unbound results.

@gabsprogrammer
gabsprogrammer marked this pull request as ready for review October 6, 2026 21:16
@Lightheartdevs
Lightheartdevs merged commit 251679d into feat/portal-laya-integration Oct 8, 2026
63 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants