Why Smoke24?
A compact, fixed regression slice for real agentic workload positioning before spending full benchmark time.
Quality vs Local Serving Cost
Default view keeps the maximum served context for each local model family when both 131K and 262K rows exist; audit filters keep the raw context-specific rows available.
Serving Mechanics
Recovered from vLLM serving metrics over each Smoke24 run window; local rows only, because external public rows are not on this hardware boundary.
Task Matrix
Each cell is one Smoke24 task for local rows or successes/attempts for external public rows.
Model Table
Sortable detail view. External public rows carry success signal only; local rows carry runtime and token metrics.
Methodology Boundary
The benchmark is coherent because the local rows share the same task corpus and serving/harness shape.