Benchmark Command build case study illustration: Romeo: the dual-R9700 build; not a lab photograph or screenshot.

Romeo: Building a Local AI Workstation Around Two R9700s

ArticleBy Benchmark CommandPublished About 5 min read
In this field report

Romeo is the inference workstation in the Benchmark Command lab: an Intel i7-14700K system with two AMD Radeon AI PRO R9700 32GB cards and 128GB of system memory. Its purpose is not to win a GPU-count contest. It is to keep local agents useful while giving the operator choices about model precision, conversation capacity and concurrent work.

The distinction matters. Two cards provide 64GB of aggregate VRAM across separate devices. That is not one transparent 64GB GPU, and it does not prove twice the speed. This is a build and configuration case study, not a dual-GPU scaling benchmark.

The machine behind the decisions

The owner-supplied hardware inventory records the following build. A September 14 monitoring audit independently corroborated the CPU family and approximately 128GiB of host RAM; it was not a new physical inspection of every component.

Component Recorded configuration
CPU Intel Core i7-14700K
Motherboard ASRock Z790 Taichi Carrara
GPUs Two Radeon AI PRO R9700 cards, 32GB each
Memory Four 32GB TeamGroup T-Create Expert DDR5 modules
Case Fractal Design North XL
Power supply NZXT C1200 Gold, ATX 3.1
Operating systems Windows 11 and Ubuntu 24.04
AI storage Separate recorded 2TB Windows and Linux AI drives

Source: the owner's Romeo hardware inventory and Benchmark Command's September 14 fleet/configuration audit. The memory kit's DDR5-6000 rating is not a measured operating speed. PSU capacity is not a wall-power measurement, and this record does not establish PCIe link widths or an optimal slot arrangement.

The first useful design: separate roles

An earlier operating arrangement gave a Q4 local Qwen instance and a separate 30B coding model their own cards. The operator had tested that arrangement and valued its isolation: one model could handle the main agent conversation while the other handled independent work.

That is a different objective from splitting one model across both cards. Separate placement can preserve room for concurrent roles; model splitting can make room for a larger instance. Neither should be declared faster without measuring the actual workload.

The historical Hermes setup receipt independently verifies a Q4 model selection, a 131,072-token configured context and one bounded provider check returning the expected readiness response in 6.14 seconds. It does not independently establish the two models' complete per-card residency. The one-card-each placement remains the owner's tested operational account.

What changed: Q8 and a larger context profile

By the September 14 configuration audit, Romeo's native model metadata reported one loaded Q8_0 instance with a configured context length of 262,144 tokens. Hermes selected the local identifier qwen3.8-27b rather than the retired Q4-specific identifier.

Those are useful, narrow facts. They establish the selected quantization and loaded-instance configuration at that observation. They do not demonstrate that a 262,144-token conversation was successfully processed, that every layer and cache allocation was GPU-resident, or that long-context answers remained accurate. The local identifier is also not a substitute for verifying the upstream artifact, architecture and supported context from its model card.

LM Studio's native model-list interface explicitly separates available models from loaded_instances and exposes each instance's configured context_length. That is the appropriate source for configuration verification; a catalogue entry alone is insufficient. LM Studio native model-list documentation.

Historical two-model GPU roles contrasted with current Q8 loaded-context metadata; current per-card residency is not asserted.
Historical placement is the owner’s tested account. Current Q8 context is observed metadata, not a verified full-context quality or scaling result.

Context capacity is not an output budget

A larger context profile addresses one particular frustration: a conversation reaching its configured capacity and requiring compression. It does not automatically shorten the wait for an answer, raise the response-token ceiling, or increase the number of simultaneous requests.

In this deployment, the main model's configured context and the auxiliary vision handler's limits are separate settings. The September 14 audit recorded a 120-second auxiliary HTTP timeout and a 1,024-token output bound. The handler could still wait for the provider or spend time processing its input before generating any output. Spare memory does not bypass those deadlines.

Profiles therefore deserve explicit names and contracts. A lower-context daily profile and a larger-context continuity profile are reasonable candidates for the next comparison, but they are not completed experiments in this article. Record the exact artifact, context, cache settings, concurrency and GPU placement for each before calling either a quality or performance mode.

Why the reported speed change needs a clean test

The operator reported earlier generation around 18–19 tokens per second and later runs around 10–11, including Q8 runs. That is a productivity symptom, not a controlled quantization result. Prompt length, prior context, runtime, request contention and the definition of tokens-per-second were not held constant across those observations.

It would be misleading to say Q8 caused the entire slowdown—or that a second card would cure it. The current evidence supports configuration choices and an unresolved performance question, not a universal Q4-versus-Q8 ranking.

A reproducible next comparison

Before changing the workstation again, capture a neutral, fixed test packet. Keep its text, token counts, output cap, reasoning setting and scoring criteria unchanged. Record the LM Studio/runtime versions and loaded configuration, then separately measure dispatch-to-first-output, prompt processing and generated-token timing when the server exposes them.

Run each candidate profile while genuinely idle, repeat the same tasks, and retain both usable answers and failures. Sample each card's process-associated memory rather than adding dashboard percentages. Only then add a simultaneous second request to measure concurrency costs. Estimates can help plan capacity, but LM Studio labels its estimate-only operation as an estimate, not a successful load or inference result. LM Studio model-loading documentation.

Romeo already has a meaningful advantage: enough hardware to test different local operating modes without immediately replacing the workstation. The next win is a documented profile that reliably delivers the required artifact—not simply the largest context slider setting or the highest quantization label.