Skip to main content
sail.inference provides thin wrappers over Sail’s hosted inference endpoints. They POST the JSON payload as given and return the raw JSON response as a dict. When a Voyage is active, the wrappers attach correlation headers so the model call shows up on the Voyage timeline, scoped to the active span and agent.
Async code uses the same methods through .aio:

responses.create

POSTs payload to /v1/responses and returns the raw JSON response dict.

chat.completions.create

POSTs payload to /v1/chat/completions and returns the raw JSON response dict. Same parameters as responses.create.

Voyage correlation

If a current Voyage exists (or you pass voyage=), the wrappers add the X-Sail-Voyage-Id header plus the active span/agent context so the dashboard attributes the model call to the right place on the timeline. Pass voyage= to correlate with a specific Voyage, or call inference with no active Voyage for ordinary uncorrelated inference.
Auto-spans: a wrapper call made with no active span gets a real, timed span created around it automatically (named after the calling function when derivable), so the model call is attributed to a span instead of appearing unscoped. Auto-spans are marked as auto in the dashboard. Explicit spans always win, so synthesis happens only where you declared nothing. Set SAIL_VOYAGE_AUTO_SPANS=0 to disable.
Auto-spans do not infer an agent. Wrap the function or block in @sail.agent(...) / with voyage.agent(...) when ownership should appear in the dashboard.

Streaming with the SDK wrapper

The high-level sail.inference.* wrappers return parsed JSON objects and do not expose a streaming iterator. Passing stream=True to those wrappers raises sail.InferenceError before sending the request. Use a raw HTTP client, the OpenAI SDK, or the Anthropic SDK pointed at Sail when you need API-level streaming.

Raw HTTP / OpenAI clients

Wrap an OpenAI-style client pointed at Sail’s API once, and every call attributes itself. Headers are computed at call time, so there is no construction-time snapshot that can go stale. Un-spanned calls get the same synthesized auto-spans as the sail.inference wrappers.
wrap_openai wraps responses.create, responses.retrieve, and chat.completions.create in place (whichever exist), is idempotent, and resolves the context-local current Voyage on each call with the process-wide fallback. Pass voyage= to pin one. For any other HTTP client, call the endpoint directly and attach the attribution headers yourself with sail.voyage.headers(). The helper carries the full context (voyage id plus the span/agent active at call time), so compute it per request, never once at client construction:

Voyage limitations

Voyages is telemetry only. It records and attributes your runs; it is not an agent framework and adds no orchestration or tool abstractions.

Errors

Inference wrappers raise sail.InferenceError (e.g. for unsupported wrapper options such as stream=True, or a missing API key) and sail.InferenceHTTPError for non-2xx responses. See Errors.