Local Model Source
A local model runner - Docker Model Runner, LM Studio, Ollama - asked what it is serving NOW.
The local counterpart of CredentialLlmServiceFactory and CredentialEmbeddingServiceFactory, and it exists for the same reason they do: something has to be able to answer for a model the platform did not know about when it was built.
WHY THIS IS NOT JUST STARTUP DISCOVERY. Each runner's autoconfiguration enumerates models in a com.embabel.common.ai.autoconfig.ProviderInitialization and registerSingletons one bean per model, which happens exactly once, while the platform is being constructed. A model pulled a minute later is invisible however reachable it is, so docker model pull costs a restart - the restart bring-your-own-key removed for keys and left in place for models on this machine. A source is asked per call, so a model pulled a second ago serves the next request.
Nothing here is built from a secret and no network identity is missing: the endpoint is known and the model is being served. The only thing missing was that nothing asked again.
Implementations must be thread-safe, and must be CHEAP to ask: LocalModelCatalog caches servedModels for a short interval, but that interval is the only thing between this and an HTTP round trip per embedding. They must also be quiet about an unreachable runner - returning an empty set rather than throwing - since a runner that is not running is the ordinary state of a deployment that does not use one.
Properties
Provider name, matching ModelMetadata.provider - "docker", "lmstudio", "ollama".
Functions
An embedding service for model, or null if this source cannot build one.
A chat service for model, or null if this source cannot build one.
What the runner is serving right now, by the names models are asked for under, each with what it is for.