24
GPTQ-Pro Smoke24 Agentic workload positioning on RTX 3090-class serving

Terminal-Bench 2.0 Smoke24

Real agentic workload signal for local 3090-class model serving.

Why Smoke24?

A compact, fixed regression slice for real agentic workload positioning before spending full benchmark time.

Fixed corpus24 Terminal-Bench 2.0 tasks under the same Terminus-2 harness shape.
Balanced selection12 shortest prior successes and 12 shortest prior failures from the recovery-corrected Qwopus3.6-27B-v2 aggregate.
Intended readA fast local-serving regression and positioning lens, not a replacement for a full Terminal-Bench leaderboard submission.

Quality vs Local Serving Cost

Default view keeps the maximum served context for each local model family when both 131K and 262K rows exist; audit filters keep the raw context-specific rows available.

Serving Mechanics

Recovered from vLLM serving metrics over each Smoke24 run window; local rows only, because external public rows are not on this hardware boundary.

Task Matrix

Each cell is one Smoke24 task for local rows or successes/attempts for external public rows.

Model Table

Sortable detail view. External public rows carry success signal only; local rows carry runtime and token metrics.

Methodology Boundary

The benchmark is coherent because the local rows share the same task corpus and serving/harness shape.