Skip to content

Commit 7345c7b

Browse files
committed
feat(studio): find-my-model-server panel, and point failures at configure
The Studio half of `slmcode configure`, plus the wiring that offers it at the moment somebody needs it. A user whose endpoint is not answering is looking at a Studio that cannot run anything, and the three fields in Settings — provider, endpoint, model — are exactly the ones they have no way to fill in correctly by guessing. The answer is on their own machine: whatever server is running knows its own address and can list what it serves. GET /api/configure looks and reports; POST writes and rebuilds the orchestrator. Two endpoints rather than one because a single one doing both would mean the only way to SEE what auto-configuration would do is to let it happen — and the panel is split the same way, so "Look around" is free and reversible while "Configure for me" is the deliberate act. A 422 from the apply path is the harness saying it looked and found nothing usable. That is information, not a failure: the body names which of the three problems it is, and the panel shows it rather than throwing the diagnosis away behind a generic toast. POST refuses while a run is active, for the same reason PUT /api/config does: a run reads the model and endpoint it was started with, and rewriting them underneath it leaves half the run on one model and half on another. Two probe remediations now name the command. Connection refused and "model not found" are precisely the two failures `configure` fixes — it finds a server running elsewhere, or picks a model the endpoint actually serves — while a rejected API key still points at `auth set`, because auto-configuration cannot supply a credential and offering it there would send somebody in a circle. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UiLDxnEu1Gf8Q5hWAFhaAd
1 parent ffb53bc commit 7345c7b

13 files changed

Lines changed: 817 additions & 3 deletions

File tree

‎docs/changelog.md‎

Lines changed: 36 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -295,6 +295,42 @@ A defect that comes back is one defect with another attempt, folded by the same
295295
rule the Fixes tab uses — the summary and the panel must never disagree about
296296
how many things went wrong.
297297

298+
### Added — `slmcode configure`, on both surfaces
299+
300+
Every piece of this existed and none of it was joined up. The harness could list
301+
an endpoint's models, measure what a (model, endpoint) pair can do, probe
302+
decoding support and check whether the configured endpoint was answering — but
303+
nothing could answer the question a new user actually has, which is *what do I
304+
put in the config*. They got a default endpoint that may not be running, a
305+
default model that may not be served, and a refusal at their first run.
306+
307+
`slmcode configure` probes the configured endpoint first, then the addresses
308+
local model servers listen on (oMLX, Ollama, LM Studio, vLLM), then any hosted
309+
provider whose API key is already set. Candidates are probed concurrently, so a
310+
machine with nothing running answers in a couple of seconds rather than waiting
311+
out each address in turn. In the Studio the same thing is the **Find my model
312+
server** panel in Settings, split the same way: *Look around* changes nothing,
313+
*Configure for me* writes.
314+
315+
The model is chosen by ruling out what cannot do the job before preferring what
316+
can. A server serves whatever it was given and the list is rarely all chat
317+
models; picking an embedding or speech model by accident produces a failure that
318+
is baffling rather than obvious — the harness runs, the model answers, and
319+
nothing it says is JSON. Among what survives, coder-tuned beats
320+
instruction-tuned beats bigger, `30B-A3B` is read as a 30B model with 3B active,
321+
and matching is on whole name segments so `codestral` is not ruled out for
322+
containing `tts`.
323+
324+
Three things it will not do: replace a working configuration (yours is probed
325+
first and kept if it answers), send your API key to a local port that merely
326+
might be a model server, or second-guess an explicit `--endpoint`.
327+
328+
A failed pass distinguishes the three real problems — nothing listening, a
329+
server with no models loaded, and a server whose models cannot write code —
330+
because they have different fixes. And it is now the remedy named when a run
331+
refuses to start because the endpoint is down or the configured model is not
332+
served, which is where somebody actually needs it.
333+
298334
### Added — correction tickets are legible on the board
299335

300336
A correction ticket looked exactly like planned work: same card, same badges, a

‎docs/cli.md‎

Lines changed: 73 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -119,6 +119,7 @@ path is embedded in a file that may be committed. Older files are migrated forwa
119119

120120
| Command | Purpose |
121121
|---|---|
122+
| `configure` | Find a model server and write a working config — start here |
122123
| `config` | `show` · `get` · `set` · `unset` · `schema` · `path` |
123124
| `stack` | `list` · `show` · `apply` · `edit` · `new` — provider/model presets |
124125
| `agent` | `list` · `show` · `edit` · `clear-llm` — per-agent LLM pins |
@@ -301,6 +302,78 @@ Details worth knowing:
301302
- The same block prints when a run fails, dies at a gate or is interrupted — "did my files
302303
change?" is a more urgent question after a failure than after a success.
303304

305+
## `configure`
306+
307+
The first command to run, and the one to run when nothing works.
308+
309+
```bash
310+
slmcode configure # look around, then ask before writing
311+
slmcode configure --yes # accept the best candidate
312+
slmcode configure --dry-run # show what it would write and stop
313+
slmcode configure --json # machine-readable
314+
slmcode configure --endpoint http://127.0.0.1:1234/v1
315+
slmcode configure --model Qwen3-Coder-30B-A3B-Instruct
316+
slmcode configure --user # write to the user config, not this project
317+
```
318+
319+
It probes the configured endpoint first, then the addresses local model servers
320+
listen on (oMLX, Ollama, LM Studio, vLLM), then any hosted provider whose API
321+
key is already in the environment. Candidates are probed **at once**, so a
322+
machine with nothing running answers in a couple of seconds rather than waiting
323+
out each address in turn.
324+
325+
```text
326+
Looking for a model server…
327+
✗ omlx http://127.0.0.1:8000/v1 — nothing is listening
328+
✗ ollama http://127.0.0.1:11434 — nothing is listening
329+
✓ lmstudio http://127.0.0.1:1234/v1 — 3 model(s) in 3ms (default lmstudio address)
330+
331+
provider lmstudio
332+
endpoint http://127.0.0.1:1234/v1
333+
model Qwen3-Coder-30B-A3B-Instruct-MLX-4bit
334+
tuned for code (coder), instruction-tuned, 30B, 3B active
335+
also available: Qwen2.5-1.5B-Instruct
336+
```
337+
338+
### How the model is chosen
339+
340+
A model server serves whatever it was given, and the list is rarely all chat
341+
models — embedding models, rerankers, speech, vision and safety classifiers sit
342+
next to the one you want. Picking one of those produces a failure that is
343+
baffling rather than obvious: the harness runs, the model answers, and nothing
344+
it says is JSON.
345+
346+
So the ranking rules those **out** first, then prefers coder-tuned over
347+
instruction-tuned over larger. A mixture-of-experts name is read correctly, so
348+
`30B-A3B` is a 30B model with 3B active rather than a 3B one. Matching is on
349+
whole name segments rather than substrings — `codestral` contains `tts` and
350+
`instruct` contains `stt`.
351+
352+
### Three things it will not do
353+
354+
- **Replace a working configuration.** The endpoint you set is probed first and
355+
kept if it answers. Re-detecting around a working setup is how a tool moves
356+
somebody's configuration out from under them.
357+
- **Send your API key to a local port.** A key travels only to the endpoint you
358+
configured, or to a hosted provider offered because that provider's own key is
359+
present — never to `127.0.0.1:1234` because LM Studio *might* be there.
360+
- **Second-guess `--endpoint`.** That flag is an instruction: a dead address you
361+
named fails rather than falling through to whatever else is running.
362+
363+
### When it finds nothing
364+
365+
The three real problems have different fixes, so they are reported differently:
366+
nothing listening, a server with no models loaded, and a server whose models
367+
cannot write code. Collapsing them into "no endpoint found" is what sends
368+
somebody to restart a server that is already running.
369+
370+
`slmcode configure` is also the remedy named when a run refuses to start
371+
because the endpoint is down or the configured model is not served. In the
372+
Studio it is the **Find my model server** panel in Settings, with the same two
373+
steps: *Look around* changes nothing, *Configure for me* writes.
374+
375+
---
376+
304377
## `compose`
305378

306379
```bash

‎docs/config.md‎

Lines changed: 8 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -41,6 +41,14 @@ never persisted for exactly that last reason. Older files are migrated forward o
4141

4242
## Provider & model
4343

44+
!!! tip "Let the harness fill these in"
45+
46+
`slmcode configure` finds the model server running on your machine, asks it
47+
what it serves, and writes `provider`, `endpoint` and `model` for you —
48+
ruling out the embedding, speech and vision models that sit next to the one
49+
you want. In the Studio it is the **Find my model server** panel in
50+
Settings. See [`configure`](cli.md#configure).
51+
4452
| Key | Default | Meaning |
4553
|---|---|---|
4654
| `provider` | `omlx` | `omlx` `ollama` `openai` `lmstudio` `openrouter` `vllm` `litellm` `together` `groq` `deepseek` `mistral` `google` `fireworks` `anthropic` … Any other name is treated as an OpenAI-compatible gateway. |

‎pkg/cli/probe.go‎

Lines changed: 9 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -112,7 +112,12 @@ func Remediation(provider, endpoint, model string, status int, err string) (caus
112112
le := strings.ToLower(err)
113113
switch {
114114
case strings.Contains(le, "connection refused"):
115-
return "connection refused", "nothing is listening on " + endpoint + " — start your " + provider + " server (e.g. `ollama serve`, LM Studio, oMLX) or point --endpoint elsewhere"
115+
// `configure` leads because it is the only remedy here that does not
116+
// require knowing the answer: if a server IS running somewhere else,
117+
// it finds it, and if none is, it says which of the three problems
118+
// this actually is.
119+
return "connection refused", "nothing is listening on " + endpoint +
120+
" — run `slmcode configure` to find your model server, or start one (`ollama serve`, LM Studio, oMLX)"
116121
case strings.Contains(le, "no such host"), strings.Contains(le, "dns"):
117122
return "host not found", "the endpoint hostname does not resolve — check --endpoint / SLMCODE_ENDPOINT for a typo"
118123
case strings.Contains(le, "timeout"), strings.Contains(le, "deadline exceeded"):
@@ -125,7 +130,9 @@ func Remediation(provider, endpoint, model string, status int, err string) (caus
125130
return fmt.Sprintf("HTTP %d unauthorized", status), "the provider rejected the API key — set one with `slmcode auth set <key>` or SLMCODE_API_KEY"
126131
case 404:
127132
if model != "" {
128-
return "HTTP 404 — model not found", fmt.Sprintf("%q is not served by this endpoint — list what is available with `slmcode agent list` or switch with `slmcode config set model <id>`", model)
133+
return "HTTP 404 — model not found", fmt.Sprintf(
134+
"%q is not served by this endpoint — run `slmcode configure` to pick one it does serve, "+
135+
"or set it yourself with `slmcode config set model <id>`", model)
129136
}
130137
return "HTTP 404", "the endpoint path is wrong — most OpenAI-compatible servers need the /v1 suffix"
131138
case 429:

‎pkg/cli/probe_test.go‎

Lines changed: 42 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -16,7 +16,7 @@ func TestRemediationMapsTransportErrors(t *testing.T) {
1616
wantCause string
1717
wantIn string
1818
}{
19-
{err: "dial tcp 127.0.0.1:1234: connect: connection refused", wantCause: "connection refused", wantIn: "start your"},
19+
{err: "dial tcp 127.0.0.1:1234: connect: connection refused", wantCause: "connection refused", wantIn: "slmcode configure"},
2020
{err: `dial tcp: lookup nope.invalid: no such host`, wantCause: "host not found", wantIn: "typo"},
2121
{err: "context deadline exceeded", wantCause: "timed out", wantIn: "still loading"},
2222
{err: "x509: certificate signed by unknown authority", wantCause: "TLS handshake failed", wantIn: "http://"},
@@ -178,3 +178,44 @@ func TestProbeCacheUnknownKey(t *testing.T) {
178178
t.Fatalf("state=%v", got.State)
179179
}
180180
}
181+
182+
// ── Pointing at the command that fixes it ────────────────────────────────
183+
//
184+
// Both of these used to tell the user to go and work out the answer. Since
185+
// `slmcode configure` exists, it IS the answer for exactly these two: a dead
186+
// address when a server may be running elsewhere, and an endpoint that is up
187+
// but does not serve the configured model.
188+
189+
func TestADeadEndpointPointsAtConfigure(t *testing.T) {
190+
_, remedy := Remediation("omlx", "http://127.0.0.1:8000/v1", "some-model", 0,
191+
"dial tcp 127.0.0.1:8000: connect: connection refused")
192+
if !strings.Contains(remedy, "slmcode configure") {
193+
t.Errorf("remedy = %q, want it to name the command that finds the server", remedy)
194+
}
195+
// And it still says how to start one, for the case where none is running.
196+
if !strings.Contains(remedy, "ollama serve") {
197+
t.Errorf("remedy = %q, want the start-a-server hint kept", remedy)
198+
}
199+
}
200+
201+
func TestAnUnservedModelPointsAtConfigure(t *testing.T) {
202+
_, remedy := Remediation("lmstudio", "http://127.0.0.1:1234/v1", "gpt-9-ultra", 404, "")
203+
if !strings.Contains(remedy, "slmcode configure") {
204+
t.Errorf("remedy = %q, want it to name the command that picks a served model", remedy)
205+
}
206+
if !strings.Contains(remedy, "gpt-9-ultra") {
207+
t.Errorf("remedy = %q, want the model that is missing named", remedy)
208+
}
209+
}
210+
211+
// A rejected key is not something auto-configuration can fix, and offering it
212+
// there would send somebody in a circle.
213+
func TestAnAuthFailureStillPointsAtAuth(t *testing.T) {
214+
_, remedy := Remediation("openai", "https://api.openai.com/v1", "gpt-4o", 401, "")
215+
if strings.Contains(remedy, "slmcode configure") {
216+
t.Errorf("remedy = %q — configure cannot supply a key", remedy)
217+
}
218+
if !strings.Contains(remedy, "auth set") {
219+
t.Errorf("remedy = %q, want the auth command", remedy)
220+
}
221+
}

‎pkg/server/configure.go‎

Lines changed: 167 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,167 @@
1+
package server
2+
3+
import (
4+
"context"
5+
"encoding/json"
6+
"net/http"
7+
"os"
8+
"strings"
9+
"time"
10+
11+
"github.com/UnicoLab/slmcode/pkg/autoconfig"
12+
"github.com/UnicoLab/slmcode/pkg/config"
13+
"github.com/UnicoLab/slmcode/pkg/orchestrator"
14+
)
15+
16+
// ── Auto-configuration from the Studio ───────────────────────────────────
17+
//
18+
// The same job as `slmcode configure`, reachable from the browser. It has to
19+
// exist on both surfaces for the same reason the CLI command does: a user whose
20+
// endpoint is not answering is looking at a Studio that cannot run anything,
21+
// and telling them to go and find a terminal is the point at which they stop.
22+
//
23+
// Split in two on purpose. GET looks and reports; POST writes. A single
24+
// endpoint that did both would mean the only way to SEE what auto-configuration
25+
// would do is to let it happen.
26+
27+
// handleConfigureScan probes for a model server and reports what it found,
28+
// changing nothing.
29+
func (s *Server) handleConfigureScan(w http.ResponseWriter, r *http.Request) {
30+
res, cfg := s.discover(r)
31+
writeJSON(w, configureView(res, cfg, false))
32+
}
33+
34+
// handleConfigureApply probes and writes the result.
35+
func (s *Server) handleConfigureApply(w http.ResponseWriter, r *http.Request) {
36+
// A run reads the model and endpoint it was started with. Rewriting them
37+
// underneath it would leave half the run on one model and half on another,
38+
// which is the same reason PUT /api/config refuses.
39+
if s.rejectMutationWhileRunning(w) {
40+
return
41+
}
42+
res, probeCfg := s.discover(r)
43+
if !res.Found {
44+
w.Header().Set("Content-Type", "application/json")
45+
w.WriteHeader(http.StatusUnprocessableEntity)
46+
_ = json.NewEncoder(w).Encode(configureView(res, probeCfg, false))
47+
return
48+
}
49+
50+
choice := res.Choice
51+
// An explicit model must be one the server actually serves: writing a
52+
// config whose first run fails is the situation this exists to prevent.
53+
if want := strings.TrimSpace(requestedModel(r)); want != "" {
54+
ok := false
55+
for _, f := range res.Findings {
56+
if f.Live() && f.Endpoint == choice.Endpoint {
57+
for _, name := range f.Models {
58+
if strings.EqualFold(name, want) {
59+
choice.Model, choice.Why, ok = name, "chosen in the Studio", true
60+
}
61+
}
62+
}
63+
}
64+
if !ok {
65+
http.Error(w, choice.Endpoint+" does not serve "+want, http.StatusUnprocessableEntity)
66+
return
67+
}
68+
}
69+
70+
var saveErr error
71+
s.withConfigWrite(func(c *config.Config) {
72+
c.Provider = choice.Provider
73+
c.Endpoint = choice.Endpoint
74+
c.Model = choice.Model
75+
// A detected provider is not a stack choice; a stale highlight would
76+
// show a stack the config no longer matches.
77+
c.ActiveStack = ""
78+
c.Normalize()
79+
saveErr = c.Save()
80+
})
81+
if saveErr != nil {
82+
http.Error(w, saveErr.Error(), http.StatusInternalServerError)
83+
return
84+
}
85+
// The running orchestrator holds the old endpoint. Rebuild it, exactly as
86+
// PUT /api/config does, or the Studio would report a configuration the
87+
// harness is not using.
88+
orch, err := orchestrator.New(s.cfg())
89+
if err != nil {
90+
http.Error(w, err.Error(), http.StatusInternalServerError)
91+
return
92+
}
93+
s.setOrch(orch)
94+
writeJSON(w, configureView(res, s.cfg(), true))
95+
}
96+
97+
// discover runs one probe pass for a request.
98+
func (s *Server) discover(r *http.Request) (autoconfig.Result, *config.Config) {
99+
cfg := s.cfg()
100+
probeCfg := *cfg
101+
if ep := strings.TrimSpace(r.URL.Query().Get("endpoint")); ep != "" {
102+
probeCfg.Endpoint = ep
103+
}
104+
timeout := autoconfig.DefaultProbeTimeout
105+
ctx, cancel := context.WithTimeout(r.Context(), timeout+7*time.Second)
106+
defer cancel()
107+
res := autoconfig.Discover(ctx, &probeCfg, os.Getenv, autoconfig.HTTPProber(timeout))
108+
if ep := strings.TrimSpace(r.URL.Query().Get("endpoint")); ep != "" {
109+
res = narrowTo(res, ep)
110+
}
111+
return res, cfg
112+
}
113+
114+
// narrowTo keeps only the endpoint the caller named, so an explicit choice
115+
// cannot fall through to whatever else happens to be running.
116+
func narrowTo(res autoconfig.Result, endpoint string) autoconfig.Result {
117+
want := strings.TrimRight(strings.TrimSpace(endpoint), "/")
118+
var kept []autoconfig.Finding
119+
for _, f := range res.Findings {
120+
if strings.TrimRight(f.Endpoint, "/") == want {
121+
kept = append(kept, f)
122+
}
123+
}
124+
out := autoconfig.Result{Findings: kept}
125+
out.Choice, out.Found = autoconfig.Choose(kept)
126+
return out
127+
}
128+
129+
// requestedModel reads an optional model override from the body.
130+
func requestedModel(r *http.Request) string {
131+
if r.Body == nil {
132+
return ""
133+
}
134+
var body struct {
135+
Model string `json:"model"`
136+
}
137+
_ = json.NewDecoder(r.Body).Decode(&body)
138+
return body.Model
139+
}
140+
141+
// configureView is the shape both endpoints return.
142+
func configureView(res autoconfig.Result, cfg *config.Config, applied bool) map[string]interface{} {
143+
tried := make([]map[string]interface{}, 0, len(res.Findings))
144+
for _, f := range res.Findings {
145+
tried = append(tried, map[string]interface{}{
146+
"provider": f.Provider, "endpoint": f.Endpoint, "reason": f.Reason,
147+
"models": f.Models, "live": f.Live(), "error": f.Err,
148+
"latency_ms": f.Latency.Milliseconds(),
149+
})
150+
}
151+
out := map[string]interface{}{
152+
"ok": res.Found,
153+
"tried": tried,
154+
"applied": applied,
155+
}
156+
if res.Found {
157+
out["choice"] = res.Choice
158+
} else {
159+
out["reason"] = res.NothingFound()
160+
}
161+
if cfg != nil {
162+
out["current"] = map[string]interface{}{
163+
"provider": cfg.Provider, "endpoint": cfg.Endpoint, "model": cfg.Model,
164+
}
165+
}
166+
return out
167+
}

0 commit comments

Comments
 (0)