Networking

Containers and Process Isolation

Understand how Linux namespaces, cgroups, capabilities, and filesystems combine to run containers.

On this page
  1. A container is assembled from kernel features
  2. Namespaces change what a process can see
  3. Cgroups control resource consumption
  4. The filesystem view shapes the process
  5. Capabilities narrow privileged operations
  6. The runtime defines the effective boundary
  7. Inspect a container as a process
  8. A useful review checklist

A container is assembled from kernel features

A container is often described as a lightweight virtual machine, but that analogy can hide important differences. A conventional Linux container shares the host kernel. The runtime creates a process with selected views and limits, then prepares a filesystem and starts the requested command. There is no separate guest kernel unless a virtual machine or another isolation layer is involved.

This design makes containers quick to start and efficient to pack, but it also means isolation depends on the kernel and runtime configuration. A container is not automatically a complete security boundary. Understanding the mechanisms helps explain both what the runtime provides and what still needs attention.

Namespaces change what a process can see

Linux namespaces give processes separate views of selected system resources. A process ID namespace changes the visible process tree. A mount namespace gives a process its own mount table. A network namespace separates interfaces, routes, and network settings. User namespaces map identities so a process can have elevated IDs inside a namespace without necessarily having the same authority on the host.

Namespaces are inherited by child processes unless the process enters another namespace. This means a process launched inside a container typically sees the container's process and network views, while the host retains its own. Some namespaces can be shared intentionally, which is why a container's actual isolation depends on the runtime flags and the surrounding configuration.

Cgroups control resource consumption

Control groups organize processes and apply resource controls. CPU quotas can limit how much processing time a group receives. Memory limits constrain resident memory and can lead to an out-of-memory kill when exceeded. I/O controls can influence access to block devices. Cgroup accounting also provides measurements for operators.

Limits need to match the workload. A memory limit that is too low can cause repeated restarts; a limit that is too high can allow one service to crowd out others. Measure ordinary and peak behavior, then test how the service handles the boundary. Resource controls do not guarantee fair performance on their own, but they make consumption more visible and bounded.

The filesystem view shapes the process

Container images commonly provide a read-only set of filesystem layers, with a writable layer added for the running container. A runtime selects a root filesystem and may mount configuration, secrets, temporary storage, or persistent data into it. The process sees paths inside that filesystem view, while the host manages the underlying files.

Read-only roots reduce accidental mutation. Temporary files can go to a bounded temporary mount. Persistent application state should use an explicit volume or external service rather than depending on the lifetime of the writable container layer. This separation makes replacement easier and clarifies which data must survive a redeployment.

Filesystem permissions remain important. A bind mount exposes host paths to the container, so the mounted directory and its ownership matter. Mount only the paths needed by the workload and use read-only mounts where possible. Avoid broad host mounts that make a compromised process more consequential.

Capabilities narrow privileged operations

Linux capabilities split many traditional root privileges into smaller permissions. A container process may run as root within its user namespace yet still lack capabilities that would let it modify host networking or load kernel modules. Dropping unused capabilities reduces the set of operations available after a process compromise.

A safer default is to grant only what the application requires. The same principle applies to seccomp filters, Linux security modules, and device access. These controls are complementary: namespaces change views, cgroups limit resources, capabilities constrain privileged actions, and system-call filters restrict interfaces. None should be treated as a substitute for updating the host kernel or isolating untrusted workloads appropriately.

The runtime defines the effective boundary

The runtime coordinates these primitives: it prepares mounts, configures namespaces, sets cgroups, applies security settings, and starts the process. Orchestrators add scheduling and service discovery, but they do not make every workload configuration safe automatically. Privileged mode, host networking, host PID visibility, and broad device or filesystem access can remove important isolation properties.

Review the configuration that is actually deployed, not only the image declaration. A compact container may still have a large host footprint if it mounts sensitive paths or runs with broad privileges. Conversely, a carefully configured container with a narrow process role, explicit resources, and minimal access can be operationally useful without claiming to be a virtual machine.

Inspect a container as a process

When debugging, inspect the running process from both perspectives. Check its command line, mounts, namespace identifiers, cgroup membership, and network interfaces. Compare the host view with the container view. This usually explains why a path is missing, a port cannot be reached, or a process sees a different PID tree.

Record the runtime, kernel version, image digest, and security settings when reproducing a problem. Container behavior can change when any of these change. A reproducible test should include resource limits and mounts, not just the image tag.

A useful review checklist

Before deploying a container, identify which namespaces it shares, what resource limits apply, which capabilities remain, what filesystem paths it can write, and how it gets network access. Verify how logs and persistent data leave the container. Test startup, shutdown, and resource exhaustion rather than checking only that the process launches.

Containers are a practical way to package and isolate processes, but the word container does not describe a single level of protection. Their properties come from a set of kernel features and runtime choices. Naming those choices makes the design easier to operate and makes the remaining risks visible.