Postman
ai.trivora.tech
Gateway API

OpenAI-compatible inference,
on your contract.

Point any OpenAI SDK at https://ai.trivora.tech and pass a Trivora key. Chat completions, streaming, tool calling and transcription all behave the way the OpenAI client already expects.

Quick start
curl https://ai.trivora.tech/v1/chat/completions \
  -H "Authorization: Bearer $TRIVORA_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"trivora-premium",
       "messages":[{"role":"user","content":"Hello"}]}'

01Authentication

Every request carries a bearer key. There is no other credential, no signing step and no tenant header — the key already identifies the organisation it belongs to.

Header
Authorization: Bearer sk-…

Two kinds of key reach this gateway, and your integration will use one or the other:

Session keys

Short-lived and minted per conversation by your application. They expire in 15 minutes and carry their own spend ceiling, so one leaking out of a browser is bounded rather than catastrophic.

Project keys

Long-lived, issued by Trivora for server-side systems. One per system, not one shared everywhere — so a key can be revoked without taking the rest down.

Either way the key is scoped to an explicit model list. Calling a model outside that list fails; it does not silently fall back to another tier.

To see exactly what a key may call:

GET/v1/models
Treat a key as a secret in transit and at rest, but design as though the end user can read it. Session keys assume exactly that — which is why they are minutes long and budget-capped rather than merely obfuscated.

02Chat completions

POST/v1/chat/completions

The request and response bodies are the OpenAI schema. Anything the OpenAI client sends is accepted; parameters a given model does not support are dropped rather than rejected.

curl https://ai.trivora.tech/v1/chat/completions \
  -H "Authorization: Bearer $TRIVORA_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "trivora-premium",
    "messages": [
      {"role": "system", "content": "You answer from the records provided."},
      {"role": "user",   "content": "Summarise the open pipeline for this account."}
    ],
    "temperature": 0.2
  }'
Apex cannot stream, and its cumulative callout budget is 120 seconds per transaction — async does not lift it. Use the Apex path for latency-tolerant work: batch classification, tagging, summarisation. Interactive chat should call the gateway from the browser so tokens arrive as they are produced.

03Streaming

Set "stream": true and read server-sent events. Chunks are the OpenAI delta format and the stream ends with data: [DONE].

SSE
data: {"choices":[{"delta":{"content":"Open "}}]}
data: {"choices":[{"delta":{"content":"pipeline "}}]}
data: [DONE]
Browser
const res = await fetch("https://ai.trivora.tech/v1/chat/completions", {
  method: "POST",
  headers: {
    "Authorization": `Bearer ${key}`,
    "Content-Type": "application/json",
  },
  body: JSON.stringify({ model: "trivora-premium", messages, stream: true }),
});

const reader = res.body.pipeThrough(new TextDecoderStream()).getReader();
let buffer = "";

while (true) {
  const { value, done } = await reader.read();
  if (done) break;
  buffer += value;

  // SSE frames are separated by a blank line; a chunk can split one in half.
  const frames = buffer.split("\n\n");
  buffer = frames.pop();

  for (const frame of frames) {
    const line = frame.replace(/^data: /, "").trim();
    if (!line || line === "[DONE]") continue;
    const delta = JSON.parse(line).choices[0].delta.content;
    if (delta) render(delta);
  }
}
Keep the trailing partial frame in the buffer, as above. A network chunk boundary lands mid-frame often enough that parsing each read in isolation looks fine in testing and drops tokens in production.

04Models

Model names are stable aliases. The provider behind each one can change without your code changing — that is the point of the alias.

ModelBuilt forNotes
trivora-premium Everything conversational The default, for interactive chat and bulk work alike. Reliable tool calling and strict JSON adherence.
trivora-extract Structured output Query intent, classification, entity extraction, chart specs — anything where the answer is a shape rather than prose. Markedly more accurate at picking the right tool and arguments than the chat tier, and strong across Indian languages. Stateless: no prompt caching, so don't send conversations to it.
trivora-reasoning Deep analysis Multi-step work the other tiers cannot finish. Materially more expensive per call, and never included in a session key.
trivora-transcribe Audio → text Fast and inexpensive. Best for single-speaker audio.
trivora-transcribe-pro Audio → text, word-aligned Word-level timestamps, verbatim output, wider language coverage. See Transcription.
trivora-standard Deprecated alias Resolves to trivora-premium, at the same price. Kept so existing integrations keep working; point new code at premium.

Which of these a key may call depends on your contract and on the key itself. Read the live list from GET /v1/models rather than hardcoding — it is the same list the gateway enforces against.

05Tools and structured output

Tool calling follows the OpenAI schema. It is the mechanism we would recommend for anything involving numbers: have the model emit a structured intent, execute it deterministically in your own code, and let the model narrate the result.

Request
{
  "model": "trivora-premium",
  "messages": [{"role": "user", "content": "How much closed last quarter?"}],
  "tools": [{
    "type": "function",
    "function": {
      "name": "aggregate_opportunities",
      "description": "Sum opportunity amounts over a period.",
      "parameters": {
        "type": "object",
        "properties": {
          "stage": {"type": "string"},
          "from":  {"type": "string", "format": "date"},
          "to":    {"type": "string", "format": "date"}
        },
        "required": ["stage", "from", "to"]
      }
    }
  }]
}
Models are unreliable at arithmetic over hundreds of rows and exact at describing a number you hand them. Aggregate in your own query layer, then pass the totals back. The same rule applies to charts: have the model emit a chart spec and render it yourself.

For a plain JSON answer with no tool involved, pass "response_format": {"type": "json_object"}.

06Transcription

POST/v1/audio/transcriptions

Multipart, matching the OpenAI audio API. Two models, billed per second of audio rather than per token.

curl
curl https://ai.trivora.tech/v1/audio/transcriptions \
  -H "Authorization: Bearer $TRIVORA_KEY" \
  -F file=@call.m4a \
  -F model=trivora-transcribe
trivora-transcribe

English-only conversations. Fast and inexpensive, returns plain text with segment timestamps. Pass a prompt naming the accounts and contacts likely to come up — it markedly improves proper nouns. 35 credits per audio-hour.

trivora-transcribe-pro

Anything not purely English, and any recording where you need to know who said what. Speaker diarization, word-level timestamps, verbatim output, automatic language detection including code-switched speech. Pass keyterms for entity accuracy. 75 credits per audio-hour.

Giving it your vocabulary

Both tiers accept context, and it is the difference between a usable transcript and one that mangles every company name. They take it in different forms, and the forms are not interchangeable.

ModelParameterShape
trivora-transcribeprompt A sentence, not a list. It is a continuation hint, so the model copies its style as well as its vocabulary — a comma-separated list of names strips punctuation from the transcript and doesn't fix the names.
trivora-transcribe-prokeyterms A repeated form field, one term per field, each under 50 characters. Not a JSON array.
Do not send a prompt with non-English audio. On trivora-transcribe an English prompt against non-English speech can put sentences into the transcript that were never spoken. Use trivora-transcribe-pro with keyterms instead — structured terms cannot leak into the output the way prose can.
Choose on the shape of the audio, not on how important it is. A one-speaker voice note gains nothing from diarization, and a two-party call transcribed without it produces a wall of text no downstream step can attribute.

Both are gated by one contract setting. If transcription is not enabled, neither appears in GET /v1/models and both are refused.

07Limits

LimitBehaviour
Tokens per minute Set per key. Exceeding it returns 429 — back off and retry.
Key budget Each key carries a spend ceiling. A session key that reaches its cap stops working; mint a new one rather than retrying with the old.
Organisation budget Your contracted allowance. We are alerted well before it binds, so it should never surface mid-conversation.
Request duration Streaming responses have no practical ceiling — each chunk resets the idle timer. A non-streaming call is cut at 300 seconds.
If a single non-streaming call might run long — a reasoning-tier one-shot, or a large batch item — either stream it or split it. The 300-second cut is enforced at the load balancer and arrives as a dropped connection, not a JSON error.

08Errors

Errors use the OpenAI envelope, so an OpenAI SDK raises them as its own exception types:

Body
{"error": {"message": "…", "type": "auth_error", "param": "None", "code": "401"}}
StatusCauseWhat to do
401Key missing, expired or revoked For a session key, mint a fresh one — do not retry the same key
400Model not permitted for this key, or a malformed body Check the model against GET /v1/models
429Rate limit Retry with exponential backoff
500 / 503Upstream model provider Retryable. Surface as "try again", not as a wrong answer

A budget exhaustion is refused rather than degraded — you will not silently get a cheaper model than the one you asked for.

09Data handling

Prompts and responses are not persisted. What is recorded is token counts and request metadata, which is what billing is computed from. The content of a conversation is not written to our storage at any point.

Inference runs on third-party model providers. Our gateway and usage ledger run in ap-south-1 (Mumbai); the model providers do not, so treat the gateway as in-region and inference as out-of-region when assessing this.

If specific fields must never leave your systems, strip them before the call. The gateway forwards what it receives — it cannot know which of your fields are sensitive.