sail.inference provides thin wrappers over Sail’s hosted inference endpoints.
They POST the JSON payload as given and return the raw JSON response as a
dict. When a Voyage is active, the wrappers attach
correlation headers so the model call shows up on the Voyage timeline, scoped
to the active span and agent.
.aio:
responses.create
payload to /v1/responses and returns the raw JSON response dict.
chat.completions.create
payload to /v1/chat/completions and returns the raw JSON response
dict. Same parameters as responses.create.
Voyage correlation
If a current Voyage exists (or you passvoyage=), the wrappers add the
X-Sail-Voyage-Id header plus the active span/agent context so the dashboard
attributes the model call to the right place on the timeline. Pass voyage=
to correlate with a specific Voyage, or call inference with no active Voyage
for ordinary uncorrelated inference.
SAIL_VOYAGE_AUTO_SPANS=0 to disable.
@sail.agent(...) / with voyage.agent(...) when ownership should appear in
the dashboard.
Streaming with the SDK wrapper
The high-levelsail.inference.* wrappers return parsed JSON objects and do
not expose a streaming iterator. Passing stream=True to those wrappers raises
sail.InferenceError before sending the request. Use a raw HTTP client, the
OpenAI SDK, or the Anthropic SDK pointed at Sail when you need API-level
streaming.
Raw HTTP / OpenAI clients
Wrap an OpenAI-style client pointed at Sail’s API once, and every call attributes itself. Headers are computed at call time, so there is no construction-time snapshot that can go stale. Un-spanned calls get the same synthesized auto-spans as thesail.inference wrappers.
wrap_openai wraps responses.create, responses.retrieve, and
chat.completions.create in place (whichever exist), is idempotent, and
resolves the context-local current Voyage on each call with the process-wide
fallback. Pass voyage= to pin one. For any other HTTP client, call the
endpoint directly and attach the attribution headers yourself with
sail.voyage.headers(). The helper carries the full
context (voyage id plus the span/agent active at call time), so compute it
per request, never once at client construction:
Voyage limitations
Voyages is telemetry only. It records and attributes your runs; it is not an agent framework and adds no orchestration or tool abstractions.Errors
Inference wrappers raisesail.InferenceError (e.g. for unsupported wrapper
options such as stream=True, or a
missing API key) and sail.InferenceHTTPError for non-2xx responses. See
Errors.