Security and data handling

This page is generated from this deployment's live configuration, not written by hand. If it says a component sends data somewhere, it does.

Where conversations are processed

This deployment sends some AI traffic to a third party. Exactly which, and nothing more:

Chat model
the configured provider for each tenantleaves your network
Embeddings
text-embedding-3-small at https://api.openai.comleaves your network
Reranking
not enabledstays inside
Verification email
mail.digitalcloud.com.np:465leaves your network

To seal this deployment

  • chat: No self-hosted inference endpoint is configured. Set LOCAL_LLM_BASE_URL to an Ollama, vLLM or LM Studio address, and set every tenant provider to "local".
  • embeddings: Embeddings are sent to https://api.openai.com. Every visitor question is embedded, so this is the highest-volume egress in the product. Set EMBEDDING_BASE_URL to a local embedding server.
  • email: Verification email is relayed through mail.digitalcloud.com.np:465, which is outside the network. Point SMTP_HOST at a mail server you run, or unset it -- in which case the process refuses to start in production rather than silently logging codes instead of sending them.

Embeddings are the highest-volume outbound call in the product: every visitor question is embedded before it is answered. It is also the one that fails quietly — a blocked endpoint degrades to full-text search, which looks like a retrieval quality problem rather than a configuration error. That is why no-egress mode is verified when the process starts rather than reported in a log.

What is stored, and how

Everything
PostgreSQL, including the vector index. No external vector store.
Provider credentials
AES-256-GCM at rest. Never logged, never returned by any API, never rendered in the UI.
Passwords
bcrypt, with a per-user salt.
Embed keys
Stored in clear, deliberately — an embed key is published inside a script tag on your own site, so it is not a secret. Its security comes from being unguessable, a per-key domain allowlist, and per-key rate limits.
Analytics
Event names and low-cardinality properties. Keys that look like credentials or personal data are stripped before the write.

Tenant isolation

Every query is scoped to a tenant inside the query, never by a filter applied to results afterwards and never by trusting a caller to pass the right id. A test suite enumerates every route on disk and fails the build if one is neither guarded nor explicitly declared public with a reason.

14 models carry tenant data. A separate test reads the schema and fails if any of them is missing from the erasure path, so a table added later cannot quietly escape deletion.

Answers you can audit

  • Every answer returns the sources it was grounded in, rendered as links in the widget.
  • Retrieved documents are fenced and declared untrusted data in the system prompt. Anyone who can add a page to a knowledge base would otherwise be able to write instructions the model follows.
  • An answer whose supporting material scores below the confidence threshold declines and offers a person, rather than assembling something plausible out of adjacent text.
  • Accuracy and hallucination rate are measured against a fixed question set on every change, not asserted.

Outbound requests

Every fetch built from a URL a tenant supplied passes through one guard: a scheme allowlist, DNS resolution with private and link-local ranges rejected (including the cloud metadata address), re-validation after every redirect, and caps on both time and response size.

The crawler obeys robots.txt, and treats an unreachable robots.txt as a full disallow rather than as permission. One request per second per host. Page and depth limits come from the plan and are enforced where the fetch happens.

Getting your data out, or erased

Export
Every conversation and lead, as JSON, from the dashboard.
Erasure
Deletes the tenant and everything belonging to it, including analytics, which has no cascade of its own and is cleared explicitly.
Retention
Configurable per tenant. Analytics are month-partitioned, so expiry drops a partition rather than deleting rows one at a time.
Audit trail
Who did what, to which tenant, from which address — including bulk reads of conversation data, which is the first thing a compromised admin account does.