Verbena is in public beta — the full platform is free during beta.Details →
AI chat · Knowledge · Ticketing

AI support that cuts costs—and keeps your customers' data private.

A full support desk: chat, email, tickets, and human handoff. Personal data is redacted before inference, so the model sees the question, not the customer.

  • Run the model on a paid API, a hosting service, or your own hardware — cheap models for the simple work, your best for the rest.
  • Run Verbena on our cloud, or provision an isolated Verbena instance that's yours alone.
Free during beta · No credit card
console.verbena.ai
Deploy anywhere

One agent. Four deployment types.

Publish once, serve everywhere. Each deployment carries its own appearance, allowed hosts, proactive rules, and rate limits — so staging and production never blur.

Support Agent
Website widgetFloating launcher, themed
MCP serverOther AIs use your agent
Sharable Chat pageBranded page, QR, expiry
REST APIYour UI, our engine
Model Freedom

Any model. Any infrastructure.

Most platforms hide the model behind a toggle and run every customer in one shared pool. Verbena hands you both dials — the whole model catalog, frontier in the cloud or open on GPUs you control, and a say in whether Verbena itself runs on our cloud or on an instance nobody else touches.

  • Local & self-hosted models — vLLM, Ollama, or any OpenAI-compatible host, plus Amazon Bedrock — direct connections with live capability probing.
  • The OpenRouter catalog — hundreds of cloud models on your own OpenRouter key. Verbena doesn't provide one; you bring yours and pay OpenRouter directly.
  • Ordered fallback chains — local first, cloud on standby; an outage never takes your support down.
  • Utility models — route classification, summaries, and knowledge upkeep to cheap, fast models.
Price watch

Price drops find you.

New models undercut old ones every month. Verbena watches the catalog against the model you actually run — benchmarks and token price — and notifies you when a cheaper model clears your bar. You don't hunt for savings; they arrive.

Price dropLower-cost model availablejust now

kwong-4.1 posts better benchmark scores than claude-sonnet at about a third off the token price.

save 34% on token costs
compared against the model you run · your test suite gates the switch
Cost management

You control the cost.

Busywork runs on cheap models

Only customer-facing answers need your best model. Classification, summaries, and knowledge upkeep route automatically to fast, low-cost utility models.

Classificationrouting, tagging, escalation decisions~$0.0003 / message
Summariesticket and conversation digests~$0.001 / ticket
Knowledge syncincremental — only changed pages are re-read$0.00 / unchanged page
Background work, last 30 days: $41.30 — under a quarter of total spend.

Know what every answer costs.

Verbena meters every model call — answering and background — so spend is a dashboard you manage, not a bill you discover.

  • Discover cheaper models — Run test suites to find the cheapest model that works for you.
  • Cost per AI resolution — every call attributed per conversation, per model, per account.
  • Local models cap the curve — send high-volume traffic to your own GPUs and per-token spend drops off the bill.
Model spendLast 30 days
$182.40$0.22 / AI resolution
claude-sonnetOpenRouter · your key$96.20
gpt-miniOpenRouter · your key$44.90
gemini-flashutility · summaries, sync$41.30
llama-70blocal · your GPUs$0.00

Verbena Editions - Our Cloud or Yours

Start hosted, move to an instance of your own whenever you're ready. Same agents, same knowledge, same console.

We host

Verbena Cloud

The whole product, hosted by us, pointed at whichever models you choose.

Runs on
Verbena's cloud
Everything you need to run support
  • Bring your own models - local, hosted or through OpenRouter
  • Simple tiered pricing
  • Cost optimized model usage
  • Personal and sensitive data redacted before inference and storage
Start free
Both editions run on any model — a paid API, a hosting service, or your own GPUs. Model freedom isn't a tier.Compare editions
OpenRouterAmazon BedrockLocal LLMAny OpenAI-compatible host
Privacy

The model sees the question, not the customer.

Email addresses, phone numbers, and card numbers are stripped before anything is sent for inference — and stripped again when a resolved ticket becomes knowledge. Your agent learns what the answer was, not who asked.

  • Before inference — patterns are replaced on the way to the model, not after it has answered.
  • Before knowledge — ticket transcripts are scrubbed unconditionally before the assimilator reads them, so nothing personal reaches an entry.
  • In what we store — the redacted form is what gets written to the chat record, not the raw message.
  • Purged on your schedule — retention windows delete chats and tickets automatically, without anyone remembering to.
What the visitor typed

Hi — I'm john@acme.com and my card 4242 4242 4242 4242 was charged twice. Call me on 555 210 4477.

redacted before inference
What the model receives

Hi — I'm [redacted email] and my card [redacted card] was charged twice. Call me on [redacted phone].

How Verbena works

You build Agents. Agents draw on Knowledge. You publish an Agent to Deployments. Everything else — models, tests, handoff, analytics — hangs off that.

01

Build an Agent

Instructions, models, tools, persona, and tests on one versioned configuration. Every save is an immutable version you can diff and roll back.

Support Agentv14 · published
Models3 + fallbacks
Tests12 / 12 passing
02

Ground it in Knowledge

Upload files, crawl your site, import resolved tickets. Content is synthesized into curated, linkable entries — not chunks — and re-synced on schedule.

Billing24 entries
Returns & refunds11 entries
docs.acme.comsynced 2h ago
03

Publish to Deployments

Bind the agent to a surface — website widget, MCP server, Sharable Chat page, or REST API — with its own appearance, allowed hosts, version pin, and rate limits.

WidgetMCPSharable Chat pageAPI
acme.com · widgetpinned v14
Knowledge synthesis

Raw content in. Curated knowledge out.

Uploading a PDF isn't knowledge. Verbena synthesizes every source — files, crawled sites, resolved tickets — extracting what matters, deduping it against what's already known, and merging it into curated, linked entries your agents can actually cite.

  • Entries, not chunks — synthesis writes human-readable entries you can open, edit, and audit. No opaque vector soup.
  • Dedupe & merge — new content folds into what the agent already knows instead of piling up duplicates.
  • Linked like a wiki — entries reference each other, so answers draw on related topics, not just the closest match.
  • Incremental re-sync — sources refresh on schedule; unchanged pages are skipped and cost nothing.
Deployment cycle

Draft → test → publish agents.

Rollback to prior versions if there's a problem.

Every save snapshots an immutable version of your agent. The test suite gates publish. Deployments serve the version you point them at — latest published, or pinned — so rolling back is simple, not a migration.

Integrations

Plugs into the stack you already run.

Connectors are MCP client connections — OAuth 2.1, tools discovered live from the server, callable by your agent mid-chat. If it speaks MCP, Verbena connects to it. No adapter code, no waiting on us.

Zendesk
Intercom
HubSpot
Shopify
Linear
GitHub
Atlassian
Notion
Asana
GitLab
Stripe
Zapier
Custom MCP serversPoint at any MCP endpoint — OAuth 2.1, API key, or none. Plus 17 curated providers, Zoho, PayPal, Square, Attio, Close and more.
+ REST API · Webhooks · Widget events

Ship your first agent today.

Create an agent, point it at your docs, run the tests, embed the widget. Free during beta — no credit card.