# Multi-tenancy for AI systems: stopping cross-customer data leaks at four layers

> When one RAG system serves many customers, a single missing WHERE clause can put one customer's documents into another customer's answer. To prevent that, the boundary has to be built into the database, the vector store and the encryption keys.

Bản gốc: https://fdetimes.net/en/guides/multi-tenant-ai-tenant-isolation-four-layers/

It is Friday, and the internal document assistant is being demoed. A user from customer A asks about the refund policy, and the answer quotes a passage straight from customer B's contract. Nobody attacked the system. One query was simply missing its `tenant_id` condition.

Every FDE building AI solutions will meet this situation sooner or later. When one platform serves many customers at once, each customer has its own documents, encryption keys and configuration.

AWS's whitepaper on tenant isolation strategies states the requirement concisely: even when tenants run in a shared environment, one tenant's resources must not be accessible to another. According to AWS, crossing that boundary just once, by any route, can be an incident a SaaS business does not recover from.

That is why tenant isolation has to be decided at the initial design stage, not left for the final sprint. This article covers the four layers that must be locked down: relational data, vectors, encryption keys and configuration, all tied to one working example.

## Silo, bridge or pool: which model?

AWS says plainly that no single strategy fits every case. The right approach depends on the domain, compliance requirements, deployment model and the services you choose. For RAG on Amazon Bedrock Knowledge Bases, AWS describes three models, and the table below ranks them by degree of isolation.

| | Silo | Bridge | Pool |
|---|---|---|---|
| Dedicated per tenant | Knowledge base, S3 bucket, OpenSearch Serverless collection | Knowledge base only | Nothing; the whole architecture is shared |
| Per-tenant KMS key | Possible, since each tenant has its own bucket | No end-to-end per-tenant encryption | One shared key |
| Filtering at query time | Separated at the infrastructure level | By knowledge base | Relies on a tenant metadata field |
| Suits | Customers with the strictest isolation needs who accept the highest cost | Cases that need a balance | Very many small tenants |

Imagine deploying a platform whose customers include one bank and a few dozen small retailers. The bank will almost certainly demand its own encryption key, which only silo can provide.

Pool makes sense for the small retailers. A single platform can run both models, provided each tenant is clearly assigned to an isolation tier.

## Layer one: let PostgreSQL enforce it

When relational data is shared in one database, AWS guidance states that row-level security (RLS) is required to keep tenants isolated.

RLS moves enforcement from hundreds of queries scattered through application code to a single place: the database. AWS recommends having the application set a tenant context variable at runtime, rather than creating a separate database user for each tenant.

```sql
ALTER TABLE documents ENABLE ROW LEVEL SECURITY;
ALTER TABLE documents FORCE ROW LEVEL SECURITY;

CREATE POLICY tenant_isolation ON documents
USING (tenant_id = current_setting('app.current_tenant')::uuid);
```

On the application side, every query must go through one function:

```python
def query_as_tenant(conn, tenant_id, sql, params):
with conn.transaction():
conn.execute(
"SELECT set_config('app.current_tenant', %s, true)",
(str(tenant_id),),
)
return conn.execute(sql, params).fetchall()
```

The `true` argument makes the value apply only to the current transaction. If it is set at session level in a connection pool, the next request may pick up a connection still carrying the previous request's tenant. That kind of leak is very hard to catch in testing.

Why the extra `FORCE` line? The PostgreSQL documentation states that superusers and roles with the `BYPASSRLS` attribute always bypass RLS. Table owners bypass it too, unless the table has `FORCE ROW LEVEL SECURITY` enabled.

Conversely, if RLS is enabled on a table but no policy exists, PostgreSQL denies by default: no rows are visible and none can be modified. When configuration is incomplete, the system returns nothing rather than the wrong data.

## Layer two: one namespace per tenant

In vector stores, the most common mistake is attaching `tenant_id` as metadata and filtering at query time. Pinecone's documentation recommends one namespace per tenant, because each namespace is stored separately, so tenant data is physically isolated.

With metadata filtering alone, Pinecone notes that a query still scans the entire namespace regardless of the filter. It is both weaker and more expensive.

This may seem to contradict the table above, where Bedrock's pool model relies on exactly that tenant metadata field. In fact they are different choices: pool on Bedrock Knowledge Bases shares the whole architecture, so the metadata filter is the only thing left to separate tenants.

Pinecone gives you an extra option in namespaces. If your vector store has this mechanism, keep pooling for everything else, such as the pipeline and the shared KMS key, but still give each small tenant its own namespace instead of relying on the filter alone.

```python
results = index.query(
vector=embedding,
top_k=5,
namespace=tenant_cfg["vector_namespace"],
)
```

The detail that matters is that `namespace` is not taken from user input. It must come from the tenant config, which is looked up from the authenticated login session. If the namespace arrives as a URL parameter, the physical isolation underneath is worthless.

## Layer three: bind the tenant to the encryption key

Even when you are forced to share one KMS key, you can still extend the tenant boundary into the cryptographic layer with encryption context. These are non-secret key-value pairs that AWS KMS cryptographically binds to the ciphertext.

To decrypt, you must supply the same encryption context. Encryption context can also be used as a condition in key policies and grants to scope permissions per tenant.

```python
dk = kms.generate_data_key(
KeyId=tenant_cfg["kms_key"],
KeySpec="AES_256",
EncryptionContext={"tenant_id": tenant_cfg["tenant_id"]},
)
```

If tenant B's code fetches tenant A's blob and calls decrypt with B's `tenant_id` context, KMS refuses. Encryption context is also recorded in CloudTrail, giving you a ready-made per-tenant key usage log to hand to the customer's audit team.

Precisely because encryption context is logged, however, AWS requires that it contain no sensitive information. Use a meaningless UUID, not the customer's name or tax ID.

## Layer four: configuration is where everything converges

The three layers above only hold when a single source answers the question "what does this tenant use?". Keep one config record per tenant; every request entering the system must look it up first, and everything downstream reads from it. For the bank and retailer example above, the two records might look like this:

```yaml
- tenant_id: 8f1c2e4a-5b7d-4c21-9a3e-0d6f7b2c1a90
isolation_tier: silo
vector_namespace: ns-8f1c2e4a
kms_key: alias/tenant-8f1c2e4a
model: 
system_prompt: prompts/8f1c2e4a.md
limits: { requests_per_minute: 60 }

- tenant_id: 3a9d7c10-2e4f-4b88-b1c5-7e0a9f6d2b34
isolation_tier: pool
vector_namespace: ns-3a9d7c10
kms_key: alias/pool-shared
model: 
system_prompt: prompts/default.md
limits: { requests_per_minute: 20 }
```

Note that both use a UUID as the identifier, with no customer name anywhere, so the value can go straight into the encryption context. The pool tenant shares a key, as the pool model dictates, but still has its own namespace, following layer two.

When the bank's security team asks "where is our data, and which key encrypts it?", you open their record and point to each line instead of digging through code.

The config record also feeds the isolation test suite. The test below logs in as tenant A and deliberately tries to read, search and decrypt B's data:

```python
def test_tenant_a_cannot_reach_tenant_b(conn, index, kms, b_doc_ids, b_blob):
a, b = load_tenant(TENANT_A), load_tenant(TENANT_B)

# Read: RLS must return 0 rows, even when asking directly for B's tenant_id
rows = query_as_tenant(conn, a["tenant_id"],
"SELECT id FROM documents WHERE tenant_id = %s", (b["tenant_id"],))
assert rows == []

# Search: A's namespace must not contain any of B's documents
hits = index.query(vector=PROBE_VECTOR, top_k=50,
namespace=a["vector_namespace"])
assert not {m.id for m in hits.matches} & set(b_doc_ids)

# Decrypt: B's blob with A's context must be rejected by KMS
with pytest.raises(kms.exceptions.InvalidCiphertextException):
kms.decrypt(CiphertextBlob=b_blob,
EncryptionContext={"tenant_id": a["tenant_id"]})
```

Run this test in CI with exactly the role the application uses in production. Run it as the owner role and the read check may pass even though RLS was never actually applied.

**Điểm mấu chốt:** The boundary between customers must live in the infrastructure, not only in the code.

## Mistakes that defeat isolation

The most common mistake is letting the application connect as the role that owns the tables, or as a role with `BYPASSRLS`. The policy is still there but does not apply to that role, so every test passes. Keep the migration role and the application role separate.

The second is treating a vector store metadata filter as an isolation mechanism. The filter is just a condition in a query, and one piece of code that forgets to pass it is enough to leak data.

The third is choosing bridge for convenience, then discovering the customer requires end-to-end encryption with its own key. Bridge cannot meet that requirement, so ask about encryption keys in the very first discovery session.

The last is testing only the happy path. The isolation suite in layer four is only worth having if it deliberately does the wrong thing: reading, searching and decrypting another tenant's data, then asserting that all three operations fail.

## Turning this skill into an interview advantage

When reading job descriptions for FDE or solutions engineer roles, watch for keywords such as "multi-tenant", "data isolation", "customer-managed keys" or "compliance". When you see them, have a concrete story ready about how you separated data, keys and configuration between customers.

On your CV, do not just write "experienced with multi-tenancy". Be specific: RLS with a tenant context variable set per transaction, one vector store namespace per tenant, per-tenant KMS encryption context, and automated tests that attempt cross-tenant reads. A line like that shows an interviewer you understand where systems break, not just the terminology.

The next time someone says "let's separate tenants later", you will know the price of that "later": a single contract passage surfacing in the wrong demo is enough.

**Thử ngay tuần này:**

- Set up a Postgres table with two tenants, enable ENABLE and FORCE ROW LEVEL SECURITY, then write a test proving tenant A reads 0 rows of B's data, even when running as the owner role.
- Review the vector store in your current project: if tenants are separated by a metadata filter, estimate the effort to move to one namespace per tenant.
- Add a line to your CV describing the specific tenant isolation mechanism you built and how you tested it.

## Nguồn

- [SaaS Tenant Isolation Strategies: Isolating Resources in a Multi-Tenant Environment (AWS Whitepaper)](https://docs.aws.amazon.com/whitepapers/latest/saas-tenant-isolation-strategies/saas-tenant-isolation-strategies.html)

- [Multi-tenant RAG with Amazon Bedrock Knowledge Bases](https://aws.amazon.com/blogs/machine-learning/multi-tenant-rag-with-amazon-bedrock-knowledge-bases/)

- [Implement multitenancy (Pinecone docs)](https://docs.pinecone.io/guides/index-data/implement-multitenancy)

- [Row-level security recommendations - AWS Prescriptive Guidance](https://docs.aws.amazon.com/prescriptive-guidance/latest/saas-multitenant-managed-postgresql/rls.html)

- [PostgreSQL Documentation: Row Security Policies](https://www.postgresql.org/docs/current/ddl-rowsecurity.html)

- [Encryption context - AWS Key Management Service Developer Guide](https://docs.aws.amazon.com/kms/latest/developerguide/encrypt_context.html)
