Hostzero Logo
Back to Articles

Multi-tenant Kubernetes on one Ceph cluster: isolating RBD and CephFS per tenant with ceph-csi

One Ceph cluster, many Kubernetes tenants: the CephX cap set, RADOS namespaces and mgr grants that stop cross-tenant CephFS access — and the ISO 27001 evidence.

Sven Völlmecke
September 2026

If you provide Kubernetes clusters to several customers and back them with one Ceph cluster, ceph-csi is the obvious driver. It is also, with the capabilities from its own documentation, four cross-tenant paths in one key: anyone holding a cluster’s csi-cephfs-secret can mount the CephFS volumes of every other cluster read-only, list every object in the shared data pool, and delete the CSI bookkeeping of other clusters. In a multi-tenant setup that key sits in a Kubernetes Secret any tenant cluster-admin can read.

This article describes a configuration that closes that, verified on Ceph Squid with ceph-csi v3.17 via ceph-csi-operator 1.0.5 (September 2026; the cap grammar is unchanged in Tentacle), and how to migrate existing volumes into it. Nothing here needs per-tenant filesystems, MDS or mgr daemons.

The setup

  • One Ceph cluster, one CephFS filesystem (kubernetes_cephfs), one RBD pool per tenant.
  • One Kubernetes cluster per tenant, each with its own ceph-csi deployment and two CephX users: k8s-<t>-rbd and k8s-<t>-ceph.
  • CephFS volumes of a tenant live in their own subvolume group /volumes/<t>-csi, set through subVolumeGroup in the ceph-csi ClientProfile.

RBD is the easy half. Pool-scoped caps (osd allow rwx pool=<tenant>, mon profile rbd) are a hard boundary, and the RBD journal ceph-csi keeps lives inside that pool. Everything below is about CephFS, where the pool is shared.

Where the upstream ceph-csi template leaks

The ceph-csi documentation gives this user for CephFS:

mon  allow r fsname=kubernetes_cephfs
mgr  allow rw
osd  allow rw tag cephfs metadata=kubernetes_cephfs, allow rw tag cephfs data=kubernetes_cephfs
mds  allow r fsname=kubernetes_cephfs path=/volumes, allow rws fsname=kubernetes_cephfs path=/volumes/<group>

It assumes one tenant per filesystem. With several tenants it opens four paths:

Cap

What a tenant key can do with it

mds allow r path=/volumes

Mount /volumes read-only and read every other tenant’s files.

osd allow rw tag cephfs data=…

rados ls / rados get on the whole data pool: all tenants’ file contents, no names, but the data.

osd allow rw tag cephfs metadata=…

Read and write raw MDS metadata objects, and the CSI journal of all clusters (RADOS namespace csi), including deleting other clusters’ PV bookkeeping.

mgr allow rw

Any mgr command: list, create, delete subvolumes in any group, run ceph orch.

ceph fs subvolume create --namespace-isolated does not help; it creates one namespace per subvolume and assumes one key per subvolume. ceph-csi has one key per cluster and, as of v3.17, never sets the flag — an open pull request (ceph-csi #6358) would add a namespaceIsolated StorageClass parameter, but even merged it isolates per subvolume, not per tenant, and does not touch the mgr or journal caps.

The cap set that holds

mon  allow r fsname=kubernetes_cephfs
mds  allow rwps fsname=kubernetes_cephfs path=/volumes/<t>-csi
osd  allow rw pool=customers_cephfs_metadata namespace=<t>,
     allow rw namespace=<t> tag cephfs data=kubernetes_cephfs
mgr  allow command "fs volume ls",
     allow command "<cmd>" with vol_name=kubernetes_cephfs group_name=<t>-csi     # for each <cmd> below
     allow command "fs subvolume snapshot clone" with vol_name=kubernetes_cephfs group_name=<t>-csi target_group_name=<t>-csi

The mgr commands ceph-csi v3.17 issues, and therefore the 21 <cmd> grants:

fs subvolume create        fs subvolume snapshot create      fs subvolume snapshot metadata set
fs subvolume rm            fs subvolume snapshot rm          fs subvolume snapshot metadata rm
fs subvolume resize        fs subvolume snapshot info        fs subvolume snapshot metadata ls
fs subvolume getpath       fs subvolume snapshot ls          fs subvolume snapshot protect
fs subvolume info          fs subvolume snapshot getpath     fs subvolume snapshot unprotect
fs subvolume exist         fs clone status                   fs clone cancel
fs subvolume metadata set  fs subvolume metadata rm          fs subvolume metadata ls

Plus fs volume ls (no group argument, used at startup) and fs subvolume snapshot clone with the additional target_group_name constraint, so a clone cannot be written into another group. 23 grants. The list changes between releases; for another ceph-csi version, the command strings are the "prefix" values in go-ceph’s cephfs/admin package (vendored in ceph-csi), and internal/cephfs/core shows which of its methods ceph-csi calls. Or run one PVC lifecycle against a key with these grants and read the driver logs for EACCES.

Four layers, each closing one row of the table above.

One CephFS filesystem, many tenants: upstream ceph-csi template versus the narrowed cap set, by layer — MDS path, mgr commands, OSD data pool namespace, OSD metadata pool namespace

MDS: path cap on the group only. rwps on /volumes/<t>-csi and nothing on /volumes. ceph-csi never mounts above the group; the kernel client mounts the subvolume path directly. p is needed for quotas (ceph.quota.max_bytes) which ceph-csi sets on resize, s for snapshots.

mgr: one grant per command, with argument constraints. ceph-csi v3.17 uses 23 mgr commands. allow command "<cmd>" with vol_name=… group_name=… makes the mgr reject the call unless those arguments are present and equal, so fs subvolume rm against another group fails with EACCES. fs subvolume ls is deliberately absent; ceph-csi does not use it and it would list other groups. Note that allow module volumes with group_name=… does not do this; module grants do not validate arguments of individual commands.

OSD data pool: one RADOS namespace per tenant. CephFS lets a directory carry a layout with pool_namespace; files created below inherit it, and the OSD cap grammar accepts namespace=<t> tag cephfs data=…. Set once per tenant, as admin, on a mount of the filesystem:

setfattr -n ceph.dir.layout.pool_namespace -v <t> /volumes/<t>-csi

New subvolumes inherit it (the mgr volumes module sets only layout.pool; the MDS fills the rest from the nearest ancestor layout), clones and snapshot restores keep it. From then on the tenant’s data objects sit in namespace <t>, and the key sees nothing else in the pool. This must happen before the first PVC; the isolation depends on the order.

OSD metadata pool: journal namespace. ceph-csi stores its PV↔subvolume mapping as omap objects (csi.volumes.default, csi.volume.<uuid>, csi.snaps.default, …) in the metadata pool, RADOS namespace csi by default. ClientProfile.spec.cephFs.radosNamespace: <t> moves it, and the cap narrows to that namespace. The field is immutable once set. The ceph-csi-drivers Helm chart 1.0.5 does not render it yet (ceph-csi-operator #629, fix in #630, both September 2026); until the fix ships, a Flux postRenderers patch on the ClientProfile does it:

postRenderers:
  - kustomize:
      patches:
        - target: {kind: ClientProfile, name: <clusterID>}
          patch: |
            - op: add
              path: /spec/cephFs/radosNamespace
              value: <t>

Raw MDS metadata objects (inode backtraces, directory fragments) are in the default namespace of the metadata pool, so the tenant key cannot touch them either.

Onboarding a new tenant

The order matters, because the data-pool isolation only applies to files created after the xattr is set:

  1. ceph fs subvolumegroup create kubernetes_cephfs <t>-csi
  2. setfattr -n ceph.dir.layout.pool_namespace -v <t> /volumes/<t>-csi on an admin mount
  3. ceph auth get-or-create client.k8s-<t>-ceph with the cap set above; client.k8s-<t>-rbd with profile rbd pool=<t>
  4. ClientProfile with subVolumeGroup: <t>-csi and cephFs.radosNamespace: <t>, secrets in the driver’s namespace
  5. Acceptance test (below), then hand the cluster over

Migrating existing tenants

Clusters already running on the upstream template need two migrations. Both work online, except for files that are written in place.

Data-pool namespace. A file’s layout is fixed at creation, so files written before the xattr stay in the default namespace until rewritten. Per tenant: set the xattr on the group directory and on every existing <subvolume>/<uuid> directory (those carry their own layout), then rewrite every regular file with cp -a f f.tmp && mv -f f.tmp f, skipping hard links and files modified in the last few minutes, and run a second pass. mv is atomic, so readers never see a partial file, and a file the application creates meanwhile is already in the new namespace. Safe for write-once data; for files written in place (SQLite, MySQL, Grafana) stop the pod, rewrite, start. As a sizing reference, a single admin host with a kernel mount and 12 parallel workers does about 150 MiB/s or 500 files/s, whichever binds: a 40 GiB volume with 375,000 files takes around 12 minutes, a small database restart 30 to 120 seconds. Before narrowing the cap: getfattr -n ceph.file.layout.pool_namespace on every file must return <t>. After: a full read of every file through the tenant key, and rados -p <data> ls with that key must be denied.

Two things will surprise you. The rewrite gives every file a new inode, so the next Velero/kopia run rescans everything (dedup keeps the upload small; the scan is the cost). Files in existing CephFS snapshots keep their old layout until the snapshot expires. And the data pool keeps one zero-byte object per file in its default namespace afterwards: MDS backtraces, written there regardless of file layout, needed by cephfs-data-scan and scrub. They hold no data and the tenant cap denies them. Leave them.

Journal namespace. Copy the tenant’s csi.volume.<uuid> objects and their csi.volume.<pv> keys from namespace csi to <t> (a 60-line librados script; which entries belong to the tenant follows from the cluster’s PV list), change the cap, restart the ceph-csi controller and node plugins, switch the ClientProfile, run a PVC lifecycle test, delete the copied entries from csi. The restart is not optional: librados sessions keep the caps they were opened with, and a long-running plugin will fail every journal access with rados: ret=-1, Operation not permitted until it reconnects. Kernel mounts are unaffected by the plugin restart.

Acceptance test

For every tenant, after every cap change, from the tenant’s cluster with the tenant’s secret: create a PVC and write to it, snapshot it and an existing PVC, remount an existing PVC, restore and clone from the snapshot, expand, delete everything; then grep the ceph-csi logs for EACCES and permission denied. From an admin host with the tenant key: mount another tenant’s group (must fail), rados ls on the data pool default namespace and on another tenant’s namespace (must fail), fs subvolume ls on another group (must fail). Around 3 minutes per cluster; script it.

What this has to do with ISO 27001

Tenant separation on shared storage is exactly the kind of control an ISO 27001 auditor asks about and that usually lives in someone’s head. Annex A 5.15 (access control) and 8.3 (information access restriction) of ISO 27001:2022 want the rule “a customer’s access ends at their namespace boundary” to be enforced technically and demonstrably; the cap set above is the enforcement, the acceptance test is the evidence. Two more points belong in the ISMS alongside it:

  • The order of onboarding steps is a control. Group, namespace xattr, CephX user, ClientProfile, first PVC. Set the xattr after the first PVC and the tenant’s first volumes are unprotected. That belongs in a written procedure with a checklist, not in a chat log.
  • Caps limit what a leaked key can do; they do not prevent the leak. A tenant with cluster-admin on their cluster reads csi-cephfs-secret any time. Namespaced admin RoleBindings instead of cluster-admin, and CSI secrets in the driver’s namespace rather than default, are the other half. So is keeping shared material out of tenant secrets: an RBD encryptionPassphrase copied from one cluster’s secret to the next is a shared key for every tenant the day someone sets encrypted: "true" on a StorageClass.

When a shared cluster is the wrong answer

If a customer’s contract or regulator demands physically separate storage — separate hardware, separate failure domain, separate key hierarchy — no cap set substitutes for that, and a dedicated Ceph cluster is the honest offer. The same goes for two or three tenants with very different performance profiles: a shared filesystem shares its MDS, and one tenant’s metadata storm is everyone’s latency. The configuration above is for the common case in between: dozens of tenants, ordinary confidentiality requirements, one storage team.

What we do at Hostzero

We run Ceph and Kubernetes for EU companies from our data center in Frankfurt: shared Ceph clusters behind many customer Kubernetes clusters, operated under our ISO 27001 ISMS. The configuration above is what we apply to every tenant — on the same clusters we described in Ceph as a MinIO alternative. If you are planning a Ceph cluster or want a second look at the one you have, see our Ceph storage cluster planning and services and our Kubernetes clusters in Germany.

FAQ

Do I need a separate CephFS filesystem or MDS per tenant?

No. MDS path caps, mgr command grants and RADOS namespaces give per-tenant isolation on one filesystem with one set of daemons. Separate filesystems cost an MDS pair each and do not scale to dozens of tenants.

Does --namespace-isolated on fs subvolume create solve this?

No. It creates one namespace per subvolume and expects one key per subvolume. ceph-csi uses one key per cluster and, as of v3.17, never passes the flag. The per-tenant namespace on the group directory achieves the same at the right granularity.

Why not allow module volumes for the mgr cap?

Module-level grants accept a with clause but do not validate the arguments of individual commands, so group_name= constraints are ignored. Per-command grants are validated (MgrCapGrant::validate_arguments requires the constrained keys to be present and equal).

Can I change radosNamespace on an existing ClientProfile?

Setting it for the first time works; changing it afterwards is rejected (CEL rule self == oldSelf). Copy the journal entries before you set it.

How long is a tenant down during the migration?

Zero for write-once data (uploads, media, static assets). For a database on CephFS, the rewrite of its files plus a pod restart, typically 30 to 120 seconds. Database volumes are usually better off on RBD anyway; the migration window is a good moment to move them.

What does an ISO 27001 auditor want to see for tenant separation on shared Ceph storage?

Three things: the rule (a written access-control policy that says where a tenant’s access ends), the enforcement (the CephX cap set per tenant, exported with ceph auth get), and the evidence that it works (the acceptance test results, dated, per tenant, after every cap change). The onboarding procedure with its fixed step order is the fourth document; it shows the control is applied consistently, not once.

Have questions about this topic?

Our experts are happy to advise you on your individual strategy.

Schedule a consultation