In this field report
The first question was whether Sentinel could run the model. The more important question was whether it could return usable code within a finite request budget.
In the September 14 validation, a verified Glimmer Q8 instance ran across Sentinel's two Intel Arc Pro B70 32GiB cards. Transfer integrity passed, process-associated GPU memory confirmed use of both cards, and no host RAM-exhaustion event was observed. Yet all three neutral coding requests failed the delivery gate: each final-response file was empty.
That is a failed preset test, not a verdict against all Intel inference or every Glimmer configuration. The useful result is the boundary it identifies between capacity, inference activity and an executable deliverable.
The exact tested configuration
The evidence cutoff was September 14, 2026, at 04:18 UTC. The audit used an already loaded Q8 instance; it did not run the staged Q5 launcher or change the model's placement and runtime settings.
| Item | Observed configuration |
|---|---|
| Host | Windows 11 Pro; Core Ultra 9 285K |
| Host memory | 64GiB installed; four DDR5 modules configured at 4800 MT/s |
| GPU devices | Two Arc Pro B70 cards, 32GiB each |
| Graphics driver | 32.0.101.8993 |
| LM Studio | 0.4.24+1 |
| Inference runtime | Windows Vulkan llama.cpp AVX2, version 2.37.0 |
| Model artifact | Muse-Glimmer-30B-Q8_0.gguf |
| Configured context | 131,072 tokens |
| Parallel predictions | One |
| Memory-related settings | Flash attention on; GPU KV offload; K/V Q8_0; unified KV; layer split |
Source: Benchmark Command's Sentinel validation handoff, September 14, 04:18 UTC. The artifact's filename and observed configuration identify this experiment; they do not independently establish a general upstream model specification.
Integrity passed before capability was graded
The source and destination model folders matched byte-for-byte and by SHA-256 for all five transferred artifacts. The separately stored files used by the active runtime were also hashed independently. The Q8 file was 29,612,957,984 bytes, and its projector was 1,400,328,928 bytes.
This matters because a successful copy to one folder does not prove that the runtime loaded that same copy. Here, active-file integrity was checked separately from transfer integrity. Neither check, however, predicts coding accuracy.
The observed Q8 process used both cards. Dedicated process-memory peaks were approximately 15.87GiB on one B70 and 14.77GiB on the other. These are separate per-device observations, not a claim that the hardware became one 64GiB GPU. They also do not measure dual-card scaling.
The host retained at least 24,028MiB of available memory across the three test windows. Maximum committed memory was 48.23GiB against a 93.05GiB commit limit. No RAM-exhaustion event was observed during this experiment. Temperature and board/PSU power were unavailable and were not invented.
The delivery gate was stated concretely
Three one-shot tasks tested bug repair, unit-test creation and a self-contained feature. Each request used temperature 0, top-p 1, a deterministic per-test seed and a 2,048-token output cap. The grader required one executable Python fence and a functional result; reasoning was not accepted as a substitute for final code.
Inference remained on Sentinel. Generated-code grading used a separate lab machine's Python 3.13.9 because only a Windows Store Python alias was found on Sentinel. No answer was retried, manually repaired or reconstructed from reasoning to manufacture a pass.
| Task | First reasoning observed | Final content | Client total | Gate result |
|---|---|---|---|---|
| Bug repair | 76.734s | None | 180.060s; timeout | Fail |
| Unit-test creation | 43.720s | None | 227.814s; length stop | Fail |
| Self-contained feature | 257.651s | None | 360.106s; timeout | Fail |
Source: the retained Sentinel validation handoff's objective coding gate and client results. Every final-response file was zero bytes. The practical score was therefore 0/3 deliverables, not three wrong programs: there was no final program to execute.

Reasoning consumed the completed request's budget
The unit-test request supplied the clearest accounting. It reported 2,048 completion tokens, of which 2,045 were reasoning, and ended at the length limit without final content. Approximately 99.85% of the reported completion tokens were classified as reasoning.
That does not mean reasoning is inherently undesirable. It means this observed request budget and preset did not satisfy the task's delivery requirement. The other two requests timed out without usage totals, so their token allocation cannot be inferred from the completed test.
The practical lesson is to score the requested artifact. A visible stream of reasoning proves activity, not completion. An agent that reports progress for several minutes but returns no executable answer has not passed an unattended coding gate.
These are client latencies, not engine decode scores
The completed request yielded client-derived rates of 8.990 completion tokens per second end-to-end and 11.125 after first reasoning. Those are different denominators, and neither is an independently exposed engine decode measurement. LM Studio's metrics endpoint was unsupported in this audit, leaving engine timing unavailable.
Resource collection also had limits: 219 samples arrived at about 4.56-second effective cadence. Incomplete per-engine Windows GPU counters were not promoted to whole-GPU utilization, and inconsistent Vulkan free-memory reports were rejected. A clean methodology keeps missing measurements missing.
What to test next—and what not to claim
The narrow verdict is that this 131,072-context, two-card Q8 preset is not ready for unattended coding under these three one-shot budgets. It does not establish whether lower context, another quantization, different reasoning controls or another runtime would perform better.
A next experiment should preserve the tasks and scoring criteria, changing one variable at a time. A smaller context profile can be tested before a different quantization; a larger response budget or supported reasoning control can be evaluated separately. Record dispatch, first reasoning, first final content, termination, usage and functional grade for every attempt.
The proposed one-card Q5 arrangement remains untested. Coexistence with a second local model was also unproven: the audit found no loaded local Command-R baseline. Available capacity must not be mistaken for verified concurrent service. LM Studio's native API distinguishes catalogue entries from loaded instances. Official model-list documentation.
Sentinel cleared the file-integrity and local-execution gates. It did not clear the useful-code gate. Reporting both is the result: a workstation can successfully host a substantial model while its selected operating preset still fails the job it was bought to do.

