SecureProxySecureProxy

Context compression for AI agents

Reduce the cost of your AI agents.

Reduce tokens and costs with compression that considers relevance, format, and repetition. SecureProxy compacts tool outputs in conversation history, keeping instructions and recent messages intact.

See how compression works

Compression and costs

Reduce input tokens, reuse responses, and control consumption.

Native context compression

Resend fewer tokens. Lower input costs.

When your agent reuses conversation history, it sends tool outputs to the model again. SecureProxy combines format-specific compaction, repetition reduction, and relevance selection to reduce that volume before the next call.

Savings start before the model call.

When input is billed per token, sending fewer tokens reduces that portion of the cost. The effect on total spending depends on the content, model, and agent workflow.

How outputs are reduced

Techniques are combined according to the content.

  • Relevance to the task

    For supported formats, uses terms from the request and tool references to prioritize passages related to the task.

  • Techniques by format

    Applies specific techniques to text, JSON, logs, search results, code, diffs, HTML, and configuration.

  • Less repetition

    Compacts repeated content, including blocks that recur across different tool outputs.

What stays intact

Compression only applies to tool outputs outside the recent conversation window. This content stays protected:

  • System and developer instructions
  • User and assistant messages
  • Tool outputs within the recent window

Caching and usage controls

Reuse responses and set limits.

Caching avoids repeating eligible calls. Budgets and usage limits control how much each gateway may consume.

Response caching

Reuse a response already obtained.

Enable exact caching for identical requests or semantic caching for similar ones, with configurable expiry and isolation by organization and gateway.

  • Explicit activation per gateway and request.
  • Configure expiry and similarity for your use case.
  • With Responses, available for stateless calls only.

Budgets and usage limits

Control spending and request volume.

Set daily and monthly USD budgets per gateway, plus limits on requests per second and tokens per minute.

  • Block new calls when recorded spending reaches the cap.
  • Alerts at budget consumption thresholds.
  • Cost visibility by team, model, and provider.

Budgets use recorded spending in daily and monthly UTC windows. Concurrent calls already admitted can slightly exceed the limit.

Data protection

Define what may reach the model and how to handle sensitive content.

Apply data rules before calling the model.

SecureProxy can replace detected data with placeholders or block a call according to policy. In non-streaming responses, reversible masking can restore these placeholders so your application can work with the original data.

Your system sends

Customer John, ID 123.456.789-00

The AI receives

Customer {{name_1}}, ID {{id_1}}

Your system receives

Response with placeholders restored

Reversible masking in a non-streaming callNon-streaming example: a rule masks ana@acme.com before the call reaches the provider. If the placeholder appears in the response, SecureProxy restores the email for the application.Non-streaming exampleApplicationIn your domainAI providerExternal modelSecureProxyProtection in and outRedact dataRestore response(original value restored)(provider received the placeholder)Request → Responseana@acme.com{{email_1}}Hi {{email_1}}Hi ana@acme.comana@acme.com{{email_1}}Hi {{email_1}}Hi ana@acme.com

The example shows a non-streaming response. With streaming in Chat Completions and Anthropic Messages, data is masked before sending, but placeholders are not restored in the response.

Data detectors

Identify personal identifiers, emails, phones, and cards with structured detectors. Add your own patterns when your content requires them.

Contextual rules

Use natural-language rules to evaluate contracts, patents, and information that depends on message context.

Policy actions

Choose to mask detected data, block the call, or record the occurrence according to the risk in your use case.

Routing and APIs

Connect providers, authorize models, and configure alternatives.

Providers

  • OpenAI
  • Anthropic
  • Gemini
  • Azure OpenAI
  • OpenRouter
  • Mistral
  • xAI
  • Groq
  • DeepSeek
  • Ollama
  • OpenAI-compatible APIs

Ollama and internal or custom endpoints are available with on-premises installation. Compatibility depends on the protocol and model you choose.

Protocols and integration

One gateway for applications and agents.

Choose the interface your application uses and check which features are compatible with your provider and model.

OpenAI Chat Completions

/v1/chat/completions

Connect applications that use the Chat Completions format to different authorized providers.

Features such as tools and streaming depend on the model and provider.

OpenAI Responses

/v1/responses

Use the Responses interface for integrations with OpenAI models.

OpenAI only, with no model fallback. Optional caching supports stateless calls only.

Anthropic Messages and Claude Code

/v1/messages

Connect Claude Code and clients that use Messages to compatible chat providers.

Compatibility covers Messages; it does not include every Anthropic resource API.

Three configuration steps

  1. URL: point your client to the gateway address.
  2. Key: use the access credential issued by SecureProxy.
  3. Model: specify provider/model and check protocol compatibility.

Distribution and alternatives defined by your team.

Configure how calls use credentials, retry requests, and reach allowed alternative models.

Multi-provider routing with fault toleranceA call with an explicit list of alternative models reaches SecureProxy. On an eligible failure, it tries the next authorized model in the list. The example uses OpenAI, Anthropic, and Gemini; models cannot change after the first streamed chunk.Routing and fallbackApplicationAI callSecureProxyFallback listGeminiNEXT IN THE LISTOpenAIOpenAIPRIMARY · FAILEDAnthropicResponsedelivered to the appAnthropicACTIVE · FALLBACKResponsedelivered to the appFallback follows the explicit alternatives in the request, subject to model access policies.

Define allowed or blocked providers and models per gateway. These rules also apply to alternative models in the request.

Weighted keys

Distribute calls across multiple keys for the same provider with configurable weights.

Retries

Configure additional attempts for eligible failures before trying another option.

Explicit fallback

Define which models may replace the primary, including models from other compatible providers.

For streaming, fallback can only happen before the first response chunk. The Responses API does not support model fallback.

Teams and access

Define who can sign in, which resources they can use, and what they can change.

Teams and governance

Company identity and team access.

Connect your company identity through OIDC or SAML SSO, including Microsoft Entra ID. A person can have different permissions in each team, depending on the resources they need to view or manage.

Team access

Illustrative example of members, permissions, and resources.

Team access
PersonTeamPermissionAuthorized resource
Ana CostaEngineeringEditorGateway engineering-prod
Ana CostaResearchViewerGateway research-dev
Bruno LimaSupportViewerGateway support-prod
Carla MeloLegalManagerGateway legal-prod

Permissions apply to a person within a team. The same person can edit in one team and only view resources in another.

Permissions per person and team

Viewer

View the resources and information available to the team.

Editor

Work on the configuration and resources the team can edit.

Manager

Edit resources and manage permissions for existing team members.

Credentials and isolation.

Control application access with gateway keys and centralize provider credentials. Each organization has its own resources.

Organization isolation

Organizations have their own resources and records, with isolation reinforced in the database. Within each organization, permissions determine member access.

Key rotation and access expiration

Keep provider credentials out of application code, rotate keys, and configure gateway access expiration. Expiring access does not delete audit history.

Auditing and integrations

Review what happened and deliver events to your operations.

Durable events for auditing and operations.

A durable event flow feeds queryable history and processes integrations and notifications independently. Content, decisions, cost, and latency are available for investigation; deliveries can repeat when retried.

Independently processed audit records, integrations, and alertsAn AI call produces durable events. Auditing saves a queryable history, integrations deliver signed webhooks, and notifications send cost or DLP alerts by email. Each flow processes work and handles failures independently.Events with independent destinationsAI callContent and decisionDurable eventsAsynchronous processingAuditQueryable history in the HubIntegrationsSigned webhooks per organizationEmail alertsCost and data policies

An unavailable webhook does not stop audit persistence. Integrations and notifications have their own retries and dead-letter queues.

Signed webhooks

Forward events to your systems with an HMAC signature. Payloads can contain confidential audit data and should be handled according to your access policy.

Email alerts

Receive cost and DLP alerts with operational information and links to the authenticated Hub. Prompt and response content is not included in these emails.

OpenTelemetry and metrics

Integrate telemetry through an OpenTelemetry Collector and monitor operational metrics with Prometheus.

Audit dashboard

Decisions and costs in view.

Query history by time period, team, model, policy, and status. Track consumption and latency and open records to investigate decisions. The data below illustrates the dashboard experience.

Audit dashboard · illustrative example

Every call with its team, model, policy, and cost

Calls

24,871

12%vs last month

Total cost

USD 487.32

8%vs last month

Blocks

142

18%vs last month

Avg latency

412 ms

4%vs last month

Calls per day

30d

  • Time

    14:32:08

    Team

    Support

    Model

    openai/gpt-4o-mini

    Policy

    Default

    Status

    SENT

    Cost (USD)

    0.004
  • Time

    14:31:47

    Team

    Legal

    Model

    azure/approved-deployment

    Policy

    Critical

    Status

    MASKED

    Cost (USD)

    0.012
  • Time

    14:31:22

    Team

    Engineering

    Model

    anthropic/claude-sonnet-4-6

    Policy

    Engineering

    Status

    BLOCKED

    Cost (USD)

    0.000
  • Time

    14:30:55

    Team

    Research

    Model

    gemini/gemini-2.5-pro

    Policy

    Research

    Status

    SENT

    Cost (USD)

    0.008
  • Time

    14:30:31

    Team

    Support

    Model

    anthropic/claude-haiku-4-5

    Policy

    Default

    Status

    MASKED

    Cost (USD)

    0.002
  • Time

    14:30:04

    Team

    Engineering

    Model

    openai/gpt-4o

    Policy

    Engineering

    Status

    SENT

    Cost (USD)

    0.015

Filter by team, policy, status, or time window. Each row points back to the original call in the audit log.

Deployment

Choose where the gateway, administration, and records reside.

Choose a dedicated environment operated by us or an installation on your infrastructure. Where data goes depends on the providers and endpoints you authorize.

Deployment modes: managed isolated and on your infrastructureIn the managed model, SecureProxy runs in a dedicated tenant we operate. In on-premise mode, gateway, audit, and vault all live inside the customer's perimeter. In both cases, only authorized calls reach external providers.Managed isolatedDedicated tenant · operated by SecureProxyRECOMMENDEDYour organizationClient applicationsIsolated tenant · SecureProxyGatewayAuditPolicies and keysOperated by SecureProxyManaged updates and backupsExternal providersOn your infrastructureEverything inside the perimeter you controlYour perimeterClient applicationsGatewayAuditVault · secrets storageHistory and admin stayin your controlCompatible with Docker and restricted environmentsExternal providers

Managed isolated

We run it. The environment is yours alone.

Recommended
  • ✓Dedicated environment for your organization
  • ✓No execution resources shared with other customers
  • ✓Assisted setup of providers, applications, and policies
  • ✓Data processing agreement and security documentation
  • ✓Managed updates, backups, and operation

On your infrastructure

Inside the network your team controls.

  • ✓Deploy with Docker/Compose or your own servers
  • ✓Ollama and internal or custom endpoints available on-premises
  • ✓Call history and administration inside your perimeter
  • ✓Traefik and automatic TLS in the local installation package
  • ✓Secrets-vault integration with HashiCorp Vault or OpenBao
  • ✓Administration and metrics separated from AI traffic

FAQ

Before you start.

What to consider when validating SecureProxy in your environment.

It depends on the authorized destination. For external providers, policies analyze content and can mask or block detected data before sending. Ollama and internal or custom endpoints are available with on-premises installation.

No tool guarantees compliance on its own. SecureProxy offers technical controls to support your program: data policies, organization isolation, access and key management, and queryable history for auditing.

The impact depends on context size and enabled rules. During the pilot, we measure latency, input tokens, and cost with your actual workflow, including the effect of compression.

We select a real AI workflow with your team and configure providers, access, policies, and budgets. The environment can be managed or on-premises. We measure latency, consumption, and the effect of caching or compression with pilot traffic.

Savings in your workflow

Bring your agent’s context. Assess the token reduction.

We assess compression with tool outputs from your use case and measure its effect on input tokens and cost. The pilot starts with a real application workflow.

Read the FAQ