Control

One API for every model

Point your existing SDK at the gateway and keep your code. The same endpoint reaches OpenAI, Anthropic, Azure AI Foundry, AWS Bedrock, Google Gemini, Mistral and self-hosted models, so you can change model or provider without a rewrite.

  • /v1/chat/completions, /v1/responses, /v1/embeddings and /v1/messages, plus images, speech and transcription
  • Works with the official OpenAI and Anthropic SDKs, plus any HTTP client
  • Switch provider by changing a model name, not your application
  • Self-hosted and private models sit behind the same controls
Control

Virtual keys, teams and budgets

Issue a gateway key to each application, team or person. Each key carries its own allowed models, spend budget and rate limit, and can be revoked in a click without touching a provider account.

  • Gateway virtual keys with expiry and revocation
  • Teams and users, with per-key model permissions
  • Budgets with alerts, and requests-per-minute rate limits
  • Provider keys stay inside the gateway. Applications never hold them
Secure

PII protection, tuned for Australia

Sensitive data is detected before a prompt leaves your control. Choose what happens for each kind of data: watch it, mask it, swap it for a reversible token, or stop the request. Detection is probabilistic, so test it against your own data.

  • Australian identifiers: Medicare number, TFN, ABN, driver licence, passport
  • New Zealand identifiers: IRD number, NHI number and bank account numbers, with a dedicated NZ policy preset
  • Also date of birth and card numbers (including test card numbers), plus names, emails, phone numbers and addresses
  • Four modes: monitor, redact, tokenise (restored on response) or block
  • Per-policy control over which entity types are acted on

Same prompt, four modes (illustrative)

OriginalPatient Jane, Medicare 2123 45670 1, asks about a rebate.
RedactPatient Jane, Medicare [MEDICARE_NUMBER], asks about a rebate.
Tokenise (restored in the reply)Patient Jane, Medicare <MEDICARE_1>, asks about a rebate.
BlockRequest stopped by policy. The application receives a clear error. Monitor mode records it and lets it through.
Secure

Secrets and credential detection

Developers paste logs and config into chat. The gateway recognises common key formats, tokens and private key blocks and can block or mask them before they reach a third party.

  • Cloud, source-control and payment API key patterns
  • Private keys, bearer tokens and connection strings
  • Block or redact, with an audit record of what was caught
  • Applies to prompts and to model responses
Secure

Prompt-injection detection

Requests and retrieved content are checked for attempts to override your instructions or extract hidden prompts. Run in monitor mode first to see what would be caught, then enforce.

  • Instruction-override and system-prompt extraction patterns
  • Monitor first, enforce when you are confident
  • Events logged with the key, team and model involved
  • Layered defence: a control, not a guarantee against every attack
Secure

Response DLP scanning

Model output can contain sensitive data too, from your own context or from a connected tool. Responses are scanned with the same detectors and the same four modes.

  • Same detectors and modes as requests
  • Tokenised values restored on the way back to your application
  • Streaming-aware handling
  • Separate policy for requests and responses
Optimise

Safe prompt optimisation

The gateway normalises whitespace, compacts JSON and removes duplicated context. You see a before and after token count for each change. It never strips punctuation blindly, because that changes what a prompt means.

  • Whitespace, JSON and duplicate-context normalisation
  • Before and after token savings shown per request
  • Code, quoted text and structured data left intact
  • On or off per policy

Optimisation report (illustrative)

Before{ "items": [ { "id": 1, "name": "A" } ] }    (extra whitespace, repeated context block)
After{"items":[{"id":1,"name":"A"}]}   (compact JSON, duplicate block removed)
ReportedInput tokens before and after, per request, and the total saved.
Optimise

Response caching

Identical requests can be answered from cache at no model cost and with lower latency. Semantic caching, which matches similar rather than identical prompts, is off until you switch it on, because it suits some workloads and not others.

  • Exact-match response caching
  • Opt-in semantic caching with a similarity threshold
  • Cache keyed per policy so tenants never share answers
  • Hits and estimated savings reported in the console
Optimise

Smart model routing

Route simple work to faster, lower-cost models and keep demanding work on premium ones. Every routing decision includes a plain-language "why this model" explanation so you can trust it and tune it.

  • Rules by task type, size, team or key
  • "Why this model" explanation on every decision
  • Never routes outside the models a key is allowed to use
  • Easy to turn off for sensitive workloads

Why this model (illustrative)

DecisionRouted to a faster, lower-cost model.
ReasonShort classification task under the size threshold, allowed for this key, within the team budget.
Optimise

Retries and failover

Transient errors are retried and, when a provider or model is unavailable, traffic fails over to the alternatives you have approved.

  • Automatic retries with back-off
  • Provider and model failover chains
  • Failover only uses models the key is already permitted to use
  • Failover events visible in logs
Control

Spend tracking and alerts

Every request is attributed so finance and engineering can see where money is going. Set budgets, get alerts before limits are hit, and see estimated savings from caching, optimisation and routing.

  • Breakdowns by app, team, user, model and provider
  • Budgets, thresholds and alerts
  • Estimated savings clearly labelled as estimates
  • Export for chargeback and reporting
Control

Audit logs and OpenTelemetry

Who called which model, under which policy, and what the gateway did about it. Administrative changes to organisations, keys, teams, members and invitations are written to an append-only, tamper-evident audit log that Tokard staff can review. Send traces and metrics to the observability stack you already run using OpenTelemetry.

  • Append-only, tamper-evident log of administrative changes, visible to staff
  • Content logging is configurable, so you can keep prompts out of logs
  • OpenTelemetry export for traces and metrics
  • Request and policy-action records attributed to key, team and model
Control

Policy presets, every feature a switch

Eight presets give you a sensible starting point, and every individual feature is a simple on/off switch you can override per team or key.

  • Default, Healthcare, Finance, Government, Developer, High Security, AI Agent and New Zealand presets
  • Every feature is an on/off switch
  • Assign a preset to a team, key or environment
  • Preview in monitor mode before enforcing
DefaultHealthcareFinanceGovernmentDeveloperHigh SecurityAI AgentNew Zealand
Control

Sign-in: Entra for staff, invitations for customers

Microsoft Entra single sign-on is used for the platform's own staff and admin console. Customers sign in at the customer login with the account their organisation administrator invites, using email and password.

  • Staff and administrators sign in to the admin console with Microsoft Entra
  • Customers sign in at /customerlogin with an emailed invitation, then email and password
  • Role-based access: customers see only their own organisation
  • Applications authenticate with gateway keys, not logins
Control

More than chat

One OpenAI-compatible gateway handles chat, responses, embeddings, Anthropic-style messages, image generation, text-to-speech and speech-to-text transcription, all under the same keys, budgets and policies.

  • PII guardrails apply to text-to-speech, embedding and image prompts
  • Transcription: the returned transcript is scanned. Healthcare-style policies mask identifiers one way, finance and government-style policies refuse the request
  • The audio itself reaches your chosen provider unredacted. Only the transcript text is scanned
  • Same keys, budgets and model permissions for every endpoint
Control

Organisations and customer login

Set up customer organisations with their own administrators, teams, keys and budgets. Customers sign in at the customer login and see only their own organisation's keys, usage and cost.

  • Organisations with their own admins, teams, keys and budgets
  • Customer login shows only that organisation's keys, usage and cost
  • Administrators invite members by email
  • Provider keys stay inside the gateway
Control

USD and AUD cost display

Costs can be shown in US dollars or Australian dollars at the touch of a switch, using the live ECB exchange rate. API responses carry cost headers in both currencies, so your own systems can record either one.

  • Switch every cost view between USD and AUD
  • Live ECB exchange rate
  • Cost headers in both currencies on API responses
  • Exchange rates move, so converted amounts are indicative

Keep every AI interaction within bounds

Request access and we will help you set up your first policy, key and budget.