Mount point limits and behavior
Capacity, quotas, and scale
EFS and S3 Files are elastic. There is no provisioned size. The storage value on the volume is required by Kubernetes but ignored, so do not promise users a per-Mount-Point size quota.
| Limit | Value |
|---|---|
| Access points per EFS file system | About 10,000. A default that AWS Support can increase. |
| Access points per S3 Files file system | 25,000. Cannot be increased. |
| Mount targets per Availability Zone | 1 per Availability Zone. Hard limit. |
| Mount Points created per file system, per region | 1,000. Enforced by SecurSpaces and not configurable. |
Every Create New Mount Point consumes one access point, and the ceiling is shared across all projects on that file system. An access point you create by hand occupies a slot from the moment you create it, so count pre-created access points into the same budget.
When the 1,000 limit is reached, new Create New requests are rejected with a message that the file system storage limit has been reached. The gate runs only for Create New. An Attach Existing Mount Point is neither counted against the 1,000 nor blocked by it, because it consumes no new access point slot.
The limit is compiled into the product. It is not a settings, Helm, or environment value. It is also approximate, and there are three ways a file system goes past it:
- A burst of simultaneous creates can push a file system slightly over.
- During a datastore outage the count cannot be read, and the create is allowed through unchecked.
- When the StorageClass list cannot be read, the file system id cannot be resolved, so the gate is skipped entirely and the create proceeds unguarded. This is the same missing permission that switches off the Create New encryption checks, described in Configure SecurSpaces.
Treat the limit as a leading indicator and the AWS ceiling as the real backstop. Both of the last two cases are counted, so you can alert on them. See Metrics.
Prefer a file system dedicated to SecurSpaces Mount Points. On a shared file system SecurSpaces counts only its own access points, so the real total can approach the AWS ceiling while SecurSpaces still believes there is room.
Isolation between Mount Points is directory and POSIX identity on a shared file system. Access points on one file system share one encryption key and one throughput pool. Put a tenant that needs its own key, throughput pool, and failure domain on a separate file system.
Note
A separate file system is not by itself an access boundary. Who may mount is decided by IAM and the network. To make the split real, scope the node role’s client permissions to named file systems and keep the file system’s mount targets closed to other clusters’ nodes.
Deletion and data protection
What a delete removes depends on how the Mount Point was created:
| Mode | What SecurSpaces removes | Backing data |
|---|---|---|
| Create New | The claim. The StorageClass reclaim policy then removes the volume. | Destroyed. The access point is removed. |
| Attach Existing, SecurSpaces-built volume | The claim and the volume. | Retained. |
| Attach Existing, volume you authored | Only the claim. | Retained. |
Deleting a Create New Mount Point removes the access point but leaves its directory and files on the file system. Plan periodic cleanup of these orphaned directories, and never attach a file-system-root volume over such a file system.
A Mount Point cannot be deleted while a live workspace or workspace template uses it. It also cannot be deleted while a deleted workspace in the recycling bin still uses it, because that workspace can be restored. Detaching does not help in that case. Either permanently delete the named workspaces from Deleted Workspaces, wait for the retention period to expire, or restore the workspace and remove the Mount Point from it.
Updates are metadata-only. Only the name and mount path can be changed. The file system id, access point, and subpath are fixed when the Mount Point is created, AWS Mount Points cannot be resized, and the read-write or read-only setting cannot be changed after creation.
Warning
A Mount Point is collaborative read-write storage with no undo. Every workspace that mounts it can delete any file, and the access point squashes all activity to a single POSIX identity, so file ownership gives no per-user protection. SecurSpaces does not create, schedule, verify, or monitor backups. Backup and recovery of Mount Point data are entirely the responsibility of the AWS account owner.
For Amazon EFS, the only recovery path is an AWS Backup recovery point that existed before the loss. Backups are taken per file system, not per Mount Point. A full restore returns every project’s directory on that file system together, and whoever can run it can read all of it. An item-level restore can target individual files and directories, so recovering one deleted folder does not mean bringing back everything. Prefer an item-level restore when only one folder is being recovered.
For Amazon S3 Files, bucket versioning is the recovery path: a file deleted through the file system becomes a non-current object version. Pair it with a lifecycle rule that keeps non-current versions long enough to be useful.
Warning
Never let live data move into an archive tier. Objects in Glacier Flexible Retrieval or Deep Archive, and Intelligent-Tiering objects that have fallen into an Archive Access tier, cannot be read through the file system at all. They must be restored through the S3 API before the file system can read them again, and nothing in SecurSpaces or in the mount warns you: the files simply become unreadable to every workspace. Apply transition rules only to non-current versions, and keep every current object in a directly readable class: Standard, Standard-IA, One Zone-IA, or Intelligent-Tiering with the archive tiers left off. A transition rule is only one of the two ways in. The archive tiers of Intelligent-Tiering are a bucket setting, not a lifecycle rule.
A restore never puts files back where they were. AWS Backup writes into a new directory off the file system root,
named aws-backup-restore_<datetime>, and every attempt creates another one. Because every Mount Point is rooted
inside an access point directory, that directory sits above the mount root and is invisible from every workspace.
Budget for a manual copy-back as part of every restore:
- Run the restore. Restore into the source file system whenever it still exists. That keeps the file system id and every access point valid, so no Mount Point has to change. Restore into a new file system only if the original is gone.
- From an administrative host in the same VPC, not a developer workspace, mount the file system without an
access point:
sudo mount -t efs -o tls <file-system-id>:/ /mnt/recovery. - Copy the wanted files from
/mnt/recovery/aws-backup-restore_<datetime>/into the access point directory the Mount Point uses, preserving ownership and permissions withcp -aorrsync -a. - Delete the recovery directory once you have confirmed the copy-back, because it keeps billing as file system
storage. Then unmount
/mnt/recovery.
Warning
That mount has no access point, so it exposes every project’s data on the file system. Run it under change control and unmount as soon as the copy-back is finished. Keep the
tlsoption: the SecurSpaces requirement that every mount carriestlsapplies to the volumes the product creates, not to a mount you run by hand, so without it the copy-back crosses the VPC unencrypted on port 2049.
If the original file system is gone, a restore recovers bytes, not Mount Points. A new file system has a new file system id and no access points, and a Mount Point’s file system id, access point, and subpath are fixed at creation. Every affected Mount Point must be deleted and recreated against the restored file system, with its data copied into the newly carved directory.
Rehearse the whole procedure, not just the restore job. A restore that reports success proves nothing about whether developers get their files back. Confirm from inside a workspace that the files and their ownership appear at the mount path.
Note
Keep your StorageClass and volume manifests in source control. If you lose the cluster, the volumes, claims, and StorageClasses are gone, while the Mount Point records survive and still report as ready.
Amazon S3 Files behavior to plan around
-
No hardlinks. One file is one S3 object key, and an attempt returns
Too many links. Tools that deduplicate using hardlinks fail or fall back. Symlinks work. - Export is asynchronous, taking roughly 60 to 72 seconds. A file written in a workspace appears as an object after a sync, not instantly. Within the file system, reads, writes, and locking are immediately consistent.
-
Conflicting writes discard the workspace’s version. If the same file changes both in a workspace and
directly in S3, the bucket wins. No error is returned to whoever wrote the file, and SecurSpaces does not
surface the conflict either, so the only detection channel is AWS. Alarm on
LostAndFoundFiles, the per-file-system count Amazon S3 Files publishes in theAWS/S3/Filesnamespace under theFileSystemIddimension. It counts the files currently in the lost-and-found directory, so alarm on any increase. Partition the bucket by writer, or keep external processes read-only. -
Displaced files land where workspaces cannot see them, in a directory named
.s3files-lost+found-<file-system-id>in the file system root. The name is dot-prefixed, so a plainlsdoes not show it, and it sits above every mount root, so only an operator can recover the files. Those copies are never exported, so no object version ever exists for them: bucket versioning does not cover them, and recovery from that directory is the only path. AWS keeps and bills for them indefinitely. -
A path longer than the 1,024-byte S3 key limit never reaches the bucket. Export fails terminally, so no
object version exists to restore and bucket versioning does not cover it. Alarm on the CloudWatch
ExportFailuresmetric in theAWS/S3/Filesnamespace.
Use Amazon EFS for latency-sensitive work such as git trees and build caches. Use Amazon S3 Files for large, mostly immutable datasets, where it is substantially cheaper.
Files are owned by the access point identity rather than the workspace user, so ownership-sensitive tools such as
git’s safe.directory check may warn. This is normal access point behavior.
Recover a displaced file
Reaching the lost-and-found directory means mounting the file system with no access point, which exposes every project’s data on that file system. Treat the whole procedure as a change-controlled operation on an administrative host in the same VPC, never something run from a developer workspace. If you prefer not to mount without an access point, create a temporary access point with no root directory restriction and a fixed POSIX user, allow it only from the recovery host, and delete it in the same change window. An access point rooted at the lost-and-found directory is not a workaround: files cannot be moved or renamed inside it, and the copy-back writes to the file’s original path, which lies outside it.
Mount the file system and list the directory:
sudo mount -t s3files <file-system-id>:/ /mnt/recovery
ls -la /mnt/recovery/.s3files-lost+found-<file-system-id>
<!--NeedCopy-->
The names there are not the original ones. A hexadecimal id is prepended so that repeated conflicts on one file
stay distinct, names longer than 100 characters are truncated, and the original directory path is not kept. The
path survives only in an extended attribute. Read it with getfattr. The timestamp in the attribute name is
required, because it forces a fresh status read:
getfattr -n "user.s3files.status;$(date -u +%s)" \
/mnt/recovery/.s3files-lost+found-<file-system-id>/<hex-id>_<name> --only-values
<!--NeedCopy-->
S3Key is the object key, and is empty if the object was deleted in the bucket. FilePath is the path the file
had before the conflict.
To keep the workspace’s version, copy the file back to FilePath under the mount. S3 Files exports it again as a
new object version. To keep the bucket’s version, delete the displaced copy instead. You can copy out of this
directory and delete inside it, but you cannot move or rename within it, and you cannot delete the directory
itself.
Finish in the same change window: unmount, and delete any temporary access point you created. Anything left in the directory keeps billing as file system storage.
Multi-region
Mount Points are not supported in multi-region configurations. The create dialog offers no region choice, so a Mount Point created there always belongs to the primary region, and the workspace Resources step prevents adding a Mount Point to a workspace in any other region.
Warning
That restriction is enforced in the create dialog, not on the REST create endpoint. On a deployment with a secondary region configured, an API caller that sets
region_idto that region creates a Mount Point there. The record is not filtered out afterwards, so it stays in the project’s Mount Points list and can be attached to a primary-region workspace, where the volume does not exist. Leaveregion_idunset when you create a Mount Point through the API.Workspace templates are not covered by that block. The template editor allows a Mount Point to be attached regardless of region, so a template can be saved carrying storage that workspaces outside the primary region cannot use. Do not treat this as a supported way to use Mount Points in a secondary region.
Timeouts
Provisioning is subject to two fixed timeouts. Neither is configurable.
| Timeout | Value |
|---|---|
| Volume bind | 3 minutes |
| Overall provisioning call | 5 minutes |
Exceeding either moves the Mount Point to an error state, from which it can be deleted and re-created.
Metrics
SecurSpaces publishes Prometheus metrics for Mount Points on two endpoints, and only one of them is reachable as shipped:
-
The control plane serves
/metricsover HTTPS, on port 8080, alongside the rest of its API. -
The workspace service serves
/metricsover plain HTTP, on port 2112. The chart does not put that port in a Kubernetes Service and does not declare it as a container port, so nothing can scrape it until you expose it yourself.
The chart also ships no ServiceMonitor, PodMonitor, PrometheusRule, scrape annotations, or dashboards. Building
your own alerting starts with building the scrape configuration. Because the control-plane endpoint is HTTPS, a job
pointed at http:// fails the TLS handshake, with nothing in the product to explain it:
scrape_configs:
- job_name: securspaces-central
scheme: https
metrics_path: /metrics
tls_config:
# The CA that signed the control plane's certificate.
ca_file: /etc/prometheus/certs/ca.crt
static_configs:
- targets: ["<release>-central-service:8080"]
<!--NeedCopy-->
To reach the workspace service metrics, add port 2112 to the workspace service container ports and to a Service in
front of it, then add a second job for that Service with scheme: http.
Published metrics
| Metric | Labels | What it records |
|---|---|---|
sds_mountpoint_create_total |
type, result, reason
|
Terminal outcome of a mount point create |
sds_mountpoint_create_rejected_total |
reason |
Creates rejected before provisioning starts, excluding the soft-quota rejection counted separately |
sds_mountpoint_quota_gate_total |
outcome: allowed, rejected, failed_open, skipped_no_fsid
|
Decisions of the per-file-system soft quota gate on Create New |
sds_mountpoint_attach_admit_total |
sub_mode, decision, reason
|
Attach Existing admission decisions |
sds_mountpoint_readback_total |
provider, mount_point_type, result
|
Outcome of the volume handle read-back after bind |
sds_mountpoint_state_write_total |
outcome: persisted, write_error, lost_race. state: MOUNT_POINT_PROVISIONING, MOUNT_POINT_READY, MOUNT_POINT_ERROR
|
Every Mount Point state transition except the two claim writes that begin a delete and an update |
sds_mountpoint_stranded_recovered_total |
state: MOUNT_POINT_CREATING, MOUNT_POINT_PROVISIONING, MOUNT_POINT_DELETING
|
Rows re-admitted after being stranded in a transient state |
sds_mountpoint_enable_flag_failopen_total |
cause |
Times the enable-flag check failed open, allowing a create despite an unreadable or absent configuration |
sds_mountpoint_picker_classes_hidden |
region |
StorageClasses currently hidden from the Create New picker for missing the tls mount option |
Every metric in this table is published by the control plane except sds_mountpoint_readback_total, which is
published by the workspace service. An alert on the read-back counter stays permanently empty until port 2112 is
exposed and scraped.
Query the label values exactly as they are written above. The state label carries the protocol buffer enum
names, in upper case, while every other label value is a lower-case token, so a query for state="ready" matches
nothing and reads as healthy. Note also that a counter’s series does not exist until it is first incremented, so
you cannot discover these values by scraping a deployment that has not yet hit the condition.
Warning
Do not rename these metrics to match the product name. They kept their original spelling deliberately, and renaming one silently breaks every dashboard and alert built on it.
Alerts worth building
No alerts ship with the product. These are the ones worth building first.
-
sds_mountpoint_quota_gate_totalwith anoutcomeoffailed_openorskipped_no_fsid. Either value means a Create New was allowed through with the per-file-system count unchecked. Alert on any non-zero value of either. -
sds_mountpoint_attach_admit_total{decision="deny"}. A denied attach, including a cross-project attempt. These denials exist as control-plane log lines and this counter only. They are not written as audit events, so they do not appear on the Audit page and do not reach a CEF or SIEM export. If your security team expects attach denials in the SIEM, this counter is the only channel. -
sds_mountpoint_readback_total{mount_point_type="create_new",result="no_access_point"}. Scope the alert tocreate_new. A Create New whose volume handle resolves without an access point is refused: SecurSpaces does not mark it ready and the record goes to the error state, so no such Mount Point is ever usable. A non-zero count means a Create New was accepted although its StorageClass should have been refused. It does not mean data has been exposed. Do not alert on everyresultother thanresolved: an Attach Existing that deliberately names no access point reads backno_access_point, and that is a supported configuration. This counter is the one published by the workspace service, so the alert needs the port 2112 scrape described above. -
sds_mountpoint_state_write_totalwith anoutcomeofwrite_error, andlost_racewithstate="MOUNT_POINT_READY". A lost READY write can leave a record that looks correct in the product but never truly reached ready. -
sds_mountpoint_enable_flag_failopen_total. Any non-zero value means a create was allowed while the configuration could not be read. -
sds_mountpoint_picker_classes_hiddenabove zero. That means at least one StorageClass is hidden from the Create New list. Zero does not mean there is no misconfiguration: a class you deliberately keep visible is not counted, and a probe that fails also reports zero, including when it fails because the StorageClass read permission is missing. Scope the alert to the regions still in service, because the per-region series is never deleted and a decommissioned region keeps reporting its last value.
The one failure with no counter
If the workspace service cannot read StorageClasses, the Create New encryption checks switch off and creates proceed without them. That failure increments no counter. It appears only as a log line, and only in the workspace service logs, not the control plane’s.
Build a log alert on the literal text RBAC forbidden in the workspace service logs. Match that string rather
than a longer phrase: the permission failure is reported from several different checks, and a narrower match
catches only one of them.
Configure SecurSpaces covers
the same permission, the one-time kubectl auth can-i check, and the two Helm settings that remove the grant.