AI support that cuts costs—and keeps your customers' data private.
A full support desk: chat, email, tickets, and human handoff. Personal data is redacted before inference, so the model sees the question, not the customer.
- Run the model on a paid API, a hosting service, or your own hardware — cheap models for the simple work, your best for the rest.
- Run Verbena on our cloud, or provision an isolated Verbena instance that's yours alone.
One agent. Four deployment types.
Publish once, serve everywhere. Each deployment carries its own appearance, allowed hosts, proactive rules, and rate limits — so staging and production never blur.
Any model. Any infrastructure.
Most platforms hide the model behind a toggle and run every customer in one shared pool. Verbena hands you both dials — the whole model catalog, frontier in the cloud or open on GPUs you control, and a say in whether Verbena itself runs on our cloud or on an instance nobody else touches.
- Local & self-hosted models — vLLM, Ollama, or any OpenAI-compatible host, plus Amazon Bedrock — direct connections with live capability probing.
- The OpenRouter catalog — hundreds of cloud models on your own OpenRouter key. Verbena doesn't provide one; you bring yours and pay OpenRouter directly.
- Ordered fallback chains — local first, cloud on standby; an outage never takes your support down.
- Utility models — route classification, summaries, and knowledge upkeep to cheap, fast models.
Price drops find you.
New models undercut old ones every month. Verbena watches the catalog against the model you actually run — benchmarks and token price — and notifies you when a cheaper model clears your bar. You don't hunt for savings; they arrive.
kwong-4.1 posts better benchmark scores than claude-sonnet at about a third off the token price.
You control the cost.
Busywork runs on cheap models
Only customer-facing answers need your best model. Classification, summaries, and knowledge upkeep route automatically to fast, low-cost utility models.
Know what every answer costs.
Verbena meters every model call — answering and background — so spend is a dashboard you manage, not a bill you discover.
- Discover cheaper models — Run test suites to find the cheapest model that works for you.
- Cost per AI resolution — every call attributed per conversation, per model, per account.
- Local models cap the curve — send high-volume traffic to your own GPUs and per-token spend drops off the bill.
Verbena Editions - Our Cloud or Yours
Start hosted, move to an instance of your own whenever you're ready. Same agents, same knowledge, same console.
Verbena Cloud
The whole product, hosted by us, pointed at whichever models you choose.
- Bring your own models - local, hosted or through OpenRouter
- Simple tiered pricing
- Cost optimized model usage
- Personal and sensitive data redacted before inference and storage
Verbena Dedicated
Run Verbena using your own private infrastructure.
- An isolated instance or fully dedicated server — no shared infrastructure
- Pair it with models on your own hardware and nothing leaves your perimeterMaximum Privacy
- Same agents, same knowledge, same console
The model sees the question, not the customer.
Email addresses, phone numbers, and card numbers are stripped before anything is sent for inference — and stripped again when a resolved ticket becomes knowledge. Your agent learns what the answer was, not who asked.
- Before inference — patterns are replaced on the way to the model, not after it has answered.
- Before knowledge — ticket transcripts are scrubbed unconditionally before the assimilator reads them, so nothing personal reaches an entry.
- In what we store — the redacted form is what gets written to the chat record, not the raw message.
- Purged on your schedule — retention windows delete chats and tickets automatically, without anyone remembering to.
Hi — I'm john@acme.com and my card 4242 4242 4242 4242 was charged twice. Call me on 555 210 4477.
Hi — I'm [redacted email] and my card [redacted card] was charged twice. Call me on [redacted phone].
You build Agents. Agents draw on Knowledge. You publish an Agent to Deployments. Everything else — models, tests, handoff, analytics — hangs off that.
Build an Agent
Instructions, models, tools, persona, and tests on one versioned configuration. Every save is an immutable version you can diff and roll back.
Ground it in Knowledge
Upload files, crawl your site, import resolved tickets. Content is synthesized into curated, linkable entries — not chunks — and re-synced on schedule.
Publish to Deployments
Bind the agent to a surface — website widget, MCP server, Sharable Chat page, or REST API — with its own appearance, allowed hosts, version pin, and rate limits.
Raw content in. Curated knowledge out.
Uploading a PDF isn't knowledge. Verbena synthesizes every source — files, crawled sites, resolved tickets — extracting what matters, deduping it against what's already known, and merging it into curated, linked entries your agents can actually cite.
- Entries, not chunks — synthesis writes human-readable entries you can open, edit, and audit. No opaque vector soup.
- Dedupe & merge — new content folds into what the agent already knows instead of piling up duplicates.
- Linked like a wiki — entries reference each other, so answers draw on related topics, not just the closest match.
- Incremental re-sync — sources refresh on schedule; unchanged pages are skipped and cost nothing.
Draft → test → publish agents.
Rollback to prior versions if there's a problem.Every save snapshots an immutable version of your agent. The test suite gates publish. Deployments serve the version you point them at — latest published, or pinned — so rolling back is simple, not a migration.
Plugs into the stack you already run.
Connectors are MCP client connections — OAuth 2.1, tools discovered live from the server, callable by your agent mid-chat. If it speaks MCP, Verbena connects to it. No adapter code, no waiting on us.
Ship your first agent today.
Create an agent, point it at your docs, run the tests, embed the widget. Free during beta — no credit card.