Repository navigation
feat(extensions): add managed Laya decisions for Portal - #7439
Merged
Lightheartdevs merged 27 commits intoOct 8, 2026
Merged
Conversation
gabsprogrammer
marked this pull request as ready for review
October 6, 2026 21:16
…to feat/portal-laya-service
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Adds Laya 0.3.28 as an optional managed Portal extension, using the tools in #7437. It provides classification, ordered scores and yes/no estimates while preserving the selected chat model and its settings. The recipe pins Python and all dependency hashes, authenticates a dedicated loopback connection and verifies actual inference on all three checkpoints before readiness.
CPU is supported on the qualified platforms; NVIDIA hosts can select a CUDA 13.0 image. Automatic selection leaves 2 GiB free headroom and caps Torch allocations at the smallest of 4 GiB, 25% of total VRAM and the available budget. One GPU checkpoint remains resident and unloads after 60 idle seconds. Missing CUDA, insufficient startup budget or startup allocation exhaustion fall back to CPU; explicit
cudarequires GPU startup. The allocator cap excludes driver/context memory and cannot reserve memory against other processes. AMD and Apple hosts currently use CPU.Updates rebuild local sources and use separate CPU/CUDA image identities. This fixes an observed update that built CUDA but retained the running CPU container under a shared image tag. The regression starts with stale CPU bytes under the CUDA tag, applies the shipped NVIDIA overlay through ordinary Compose up, and checks actual Torch/inference without force-recreate.
Validation at
ed86de990:ods/laya:0.3.28-cuda130-v1, image2214eebdff7792ad226d2f706c63f236858eb37d96e51d0309b01e37ed970024. Actual installed adapter tests passed authentication, full-context handling and all three checkpoints on RTX 5080 CUDA. All ten local services remained healthy; the idle health readback showed zero loaded checkpoints after the 60-second interval.The full installed Portal comparison used unchanged Qwen3.5 9B with all 128 complete news texts supplied to both methods and no expected labels in their input. All four predeclared runs are retained:
The slower Laya run was 2.77x faster than the complete baseline, with +8.59 percentage points accuracy. The incomplete baseline is a retained integrity failure, with missing IDs counted as wrong in the displayed secondary score. This is a small observed workload-specific gain, not general reliability proof. An earlier file-only baseline failed to read all source texts and remains disqualified; supplying full input for this comparison does not fix automatic file pagination.
Limits remain explicit: earlier bilingual and Portuguese sentiment tasks underperformed the main model; semantic review and failure disclosure were inconsistent. Public benchmark training overlap is unknown. ARM64 GPU hardware and physical competing-memory pressure were not tested; low-memory selection and allocation failure have unit coverage. A cold-cache update hit the existing 600-second start limit after slow CUDA dependency downloads; normal rollback restored the previous definition, and a later cached update passed. No timeout or quality assertion was relaxed.
CI follow-up at
09cd72ea5integrates the validated #7437 base and fixes the extended suite's 45-minute prerequisite stall. Existing jq and ShellCheck are verified and reused; only missing tools need APT. Missing-tool installs have bounded transfer/command deadlines, strict refresh errors and visible logs, with no signature or test checks disabled. Five setup regression tests cover preinstalled, missing, broken and failed-install cases.GitHub validation: 63/63 checks passed, including x86-64 and ARM64 Laya container qualification. Extended contract suite prepared its tools in 6 seconds and passed all 251 tests. The complete matrix and remaining project checks also passed. This PR remains stacked on #7437; no GitHub PR merge was performed.
Automatic-selection follow-up at
82fb49ef3:Automatic subtask selection now considers repeated semantic judgments inside larger website/catalog/report tasks, while skipping ordinary replies, obvious single labels, CSS/code changes and deterministic calculations. Only enabled extensions advertise the helper; opt-outs and execution-time revocation remain enforced. No keyword router or extra per-turn inference was added. Questions describe requested semantic outputs, not file formats, identity, order or build success. Invalid question errors identify the offending field so repairs can preserve unrelated dataset paths.
Validation on October 7: 132/132 focused WSL Node tests passed. The new regression verifies invalid-question rejection before inference/writes, useful field diagnostics, fresh admission and unchanged source paths on repair. A reusable installed-Portal selection runner records each declared attempt and distinguishes attempted calls, saved receipts and task termination from actual artifact correctness. The normal installer exited 0; installed/configured plugin hashes match the candidate, with Qwen3.5 9B and its 65,536-token context unchanged.
Natural-request evidence (no Laya keyword in positive prompts): baseline selection passed 7/9 cases, candidate 9/9, then 2/2 held-out wordings. A repeated editorial request after lifecycle recovery skipped Laya despite the enabled hint being present in the actual model prompt, so candidate observations are 11/12, not universal routing success. All six initial negative/opt-out cases avoided Laya. A separate disabled-extension request also made zero Laya calls and contained no enabled hint. Schema-repair attempts remain recorded; one held-out call added redundant double classification. Successful reports independently preserved IDs/order and source hashes. These small, mostly single-run observations are not a latency or accuracy benchmark.
End-to-end generated-site qualification did not pass: the first website had a JavaScript quote error; the held-out website rendered 32 cards and filtered topics, but search raised a variable-initialization error and only 20/32 original texts survived the model's manual transcription exactly. Both baseline and candidate arithmetic fixtures also contained incorrect totals. These are preserved main-model delivery failures, not relabeled Laya-selection successes or fixed by the classifier. Laya's verified saved report does not certify downstream generated code.
Lifecycle qualification found the local catalog still carried an older recipe than the reviewed service branch. The reviewed recipe source was synchronized with a retained backup; the normal authenticated update preserved the disabled state and rollback, then normal enable succeeded. No receipt or authorization control was bypassed. Laya is active again, real installed English-adapter inference passed on CUDA, and all ten services are healthy. Current image:
sha256:72843a0ddf3ef8223556a2532533b75edd917d270bc14e7dc268253617daf59e.Latest-head GitHub validation at
82fb49ef3: 63/63 checks passed, with no pending, failed or skipped checks. The PR has no merge conflicts. Main requires two approving reviews; no GitHub PR merge or force-push was performed. Qualification grading also matched native result digests across 22 retained cases and passed five synthetic evidence checks, including invalid and unbound results.