Foundations
Cloud native primer
Before you read: Linux foundations (namespaces, cgroups v2, file descriptors) and IaaS and cloud native (what the engine is responsible for).
The engine is a thin, careful layer over Linux kernel features and a handful of open specifications. This page gives you just enough of each concept to read the code, and says where it lives in this repository. After it you can take any mechanism — a namespace, a cgroup limit, an image layer, a firewall chain, a VM boot, a manifest apply — and name the file and symbol that implements it. For depth, follow the official links — they are better than any summary here.
Concepts are taught once in this handbook. The kernel primitives (namespaces, user namespaces,
cgroups v2) are taught hands-on in Linux foundations, so sections 4.1 and
4.2 only recap them and map them to the code. What each open standard requires, and how far the
engine conforms, is in Cloud native standards, later in the course.
Paths name crates (crates/<layer>/<crate>); the layers are explained in
Architecture, and for now a path simply says where the code is.
Each section has three parts: the concept, In Delonix (files and symbols you can grep),
and Read more. Paths are relative to the repository root.
For orientation in the wider ecosystem, the CNCF Landscape and the CNCF Glossary are good maps. Comparisons with runc, crun, containerd or Podman appear only where they help explain a design choice.
4.1 Linux namespaces and rootless operation#
Recap. A container is a process started in a fresh set of namespaces (mount, PID, network,
IPC, UTS, cgroup, user). The user namespace is what makes it rootless: inside, the process is
uid 0 over resources that namespace owns; outside, it is an ordinary user. An unprivileged user can
map only its own uid; a range needs newuidmap/newgidmap and /etc/subuid. Both are taught,
with commands to type, in Linux foundations — Namespaces and
User namespaces and uid mapping. Some
distributions additionally restrict unprivileged user namespaces through AppArmor — see
Environment
for the practical consequences.
In Delonix
fn spawnincrates/adapters/delonix-linux/src/lib.rsbuilds theCloneFlags(CLONE_NEWNS,CLONE_NEWPID,CLONE_NEWNET,CLONE_NEWUSER, …) and callsnix::sched::clone. IPC/UTS sharing between pod members is handled bysetnsincontainer_init.write_userns_mapsin the same file writes the maps from the parent: a single-uid map (0 <euid> 1) rootless, or a subuid range throughnewuidmap/newgidmapwhenhave_subid_helpers()says they are available.USERNS_UID_BASE/USERNS_RANGEdefine the range used when running as root.setup_rootfsmounts the container's root and callspivot_root;container_initis the code that runs inside the new namespaces beforeexecvp.- Rootless operations on files owned by mapped subuids re-execute the binary inside a mapped user
namespace:
reexec_mapped,reexec_mapped_hold,remove_tree_mapped.
Read more: namespaces(7),
user_namespaces(7),
newuidmap(1),
subuid(5),
pivot_root(2).
4.2 cgroups v2 and delegation#
Recap. cgroup v2 is one tree at /sys/fs/cgroup whose files (memory.max, cpu.max,
pids.max, memory.events) limit and account for a group of processes. An unprivileged user can
only write a subtree systemd has delegated to them, and a shell in an SSH session scope sits
outside it — so limits can silently not apply there. The tree, the "no internal processes" rule and
delegation are taught hands-on in Linux foundations — cgroups v2;
the delegation contract as a standard, and the engine's conformance, are in
Cloud native standards — 13.15.
The engine is v2-only.
In Delonix
- Root mode places containers under
delonix_compute::DELONIX_SLICE(/sys/fs/cgroup/delonix.slice). - Rootless mode finds the user's service cgroup and creates leaves under
<user@uid.service>/dlx-containers— seeuser_service_baseandtry_delegated_baseincrates/adapters/delonix-linux/src/lib.rs.cgroup_limits_applyanswers "will limits apply on this host?" without starting a container. - The design decision about the intermediate level is
ADR-0015; how the CRI follows the kubelet's cgroup
hierarchy is ADR-0038, with the kubelet's
parent validated by
KubeCgroupParent::parseincrates/contexts/delonix-compute/src/record.rs.
Read more: Linux kernel — Control Group v2,
systemd — Control Group APIs and Delegation,
cgroups(7).
4.3 Capabilities, seccomp, AppArmor, masked paths#
Root's power is split into capabilities (CAP_NET_ADMIN, CAP_SYS_ADMIN, …). A container
keeps a small default set and drops the rest. seccomp installs a BPF filter that allows or
denies system calls, reducing kernel attack surface. AppArmor (and SELinux on other
distributions) are Linux Security Modules that confine a process by profile. Finally, runtimes
mask sensitive /proc and /sys paths (bind something empty over them) and make others
read-only, because those files leak host information or allow host control.
A subtle point the code documents: clone3 passes its flags through a pointer that a seccomp
filter cannot inspect, so a filter that blocks clone(CLONE_NEWUSER) must also make clone3
fail with ENOSYS to force libc back to the filterable clone.
In Delonix
- Capabilities:
KEPT_CAPSandresolve_cap_keepincrates/adapters/delonix-linux/src/capabilities.rs;drop_capabilitiesinlib.rs. - seccomp:
apply_seccompincrates/adapters/delonix-linux/src/lib.rs(built with theseccompilercrate, including theclone3→ENOSYSpre-filter); custom JSON profiles are parsed and compiled inseccomp_profile.rs(parse,compile). - AppArmor:
apply_apparmorinlib.rs. - Masked and read-only paths:
DEFAULT_MASKED_PATHS,DEFAULT_READONLY_PATHS,apply_masked_paths,apply_readonly_paths,mask_proc_pathsinlib.rs. - The node-level security decisions (policy, admission, events, score) are a separate pure crate:
evaluateincrates/contexts/delonix-security-runtime/src/admission.rs(ADR-0026).
Read more: capabilities(7),
kernel — Seccomp BPF,
AppArmor documentation,
OCI runtime spec — Linux config
(the maskedPaths/readonlyPaths/seccomp fields other runtimes consume).
4.4 OCI images, content-addressed storage and overlayfs#
The Open Container Initiative publishes three specifications:
- the image spec — an image is a manifest (JSON) that points to a config and an ordered list of layers (tarballs), and possibly an index that points to one manifest per platform;
- the distribution spec — the HTTP API registries serve (
/v2/<name>/manifests/<ref>,/v2/<name>/blobs/<digest>, token auth); - the runtime spec — how a runtime such as runc or crun is told to run a filesystem bundle.
Everything is content-addressed: a blob is named by the SHA-256 digest of its bytes, so a
client verifies what it downloaded by hashing it. Pulling by digest (name@sha256:…) is only a
guarantee if the manifest is checked against that digest as well as each blob against the
manifest.
At run time, layers are stacked with overlayfs: read-only lowerdirs, a writable upperdir
where changes are copied up, and a workdir. Many containers can share the same lower layers.
In Delonix
- Registry client (distribution spec):
crates/adapters/delonix-oci/src/registry.rs—pull_from_registry*functions, theACCEPT_MANIFESTmedia types, andverify_manifest_digest. Types come from theoci-speccrate. - Content-addressed blob store:
Casincrates/adapters/delonix-oci/src/cas.rs. - Writing an OCI image layout archive:
write_oci_archiveinsave.rs. - Overlay preparation:
ImageStore::prepare_overlayinoverlay.rswrites anoverlay-lowersmarker (LOWERS_FILE); the mount itself happens inside the container's user and mount namespaces inmount_overlay_if_marked(crates/adapters/delonix-linux/src/lib.rs), using the new mount API — see ADR-0016 and ADR-0037. - The engine runs containers itself rather than handing an OCI runtime bundle to runc/crun.
Read more: OCI image spec, OCI distribution spec, OCI runtime spec, kernel — Overlay Filesystem.
4.5 Container networking#
Linux networking building blocks:
- a network namespace has its own interfaces, routes and firewall;
- a veth pair is a virtual cable with one end in each namespace;
- a bridge is a virtual switch joining many veth ends;
- nftables is the kernel packet filter and NAT engine; DNAT rewrites a destination (how a
published port reaches a container), and conntrack tracks flows so reply traffic of an
allowed connection passes (
ct state established,related); - slirp4netns gives an unprivileged network namespace outbound connectivity by emulating a TCP/IP stack in user space, and forwards host ports into it;
- VXLAN carries L2 frames over UDP between hosts, and WireGuard encrypts a tunnel.
CNI (Container Network Interface) is a spec where a runtime executes plugin binaries
(bridge, host-local, portmap, …) with ADD/DEL commands and a JSON config from
/etc/cni/net.d. Kubernetes runtimes use it for pod networking.
In Delonix
- Rootless networking cannot create interfaces on the host, so the engine keeps a long-lived
holder network namespace: a minimal pin process owns the namespaces, and a restartable
control process serves a Unix socket. See
start_pin,start_controlandensure_upincrates/adapters/delonix-sdn/src/infra.rs. - Attaching a workload:
attach_container(IPAM + control command) anddo_attach(veth into the bridge, inside the holder). Bridge names come frombridge_namein the dependency-freecrates/foundation/delonix-net-rules/src/lib.rs. - Outbound and port forwarding:
slirp_attachandslirp_add_hostfwdincrates/adapters/delonix-sdn/src/lib.rs(which spawnslirp4netns); in-holder publishing inpublish_port/do_publish(infra.rs). - Firewall:
table ip dlxingwith base chainsfwguard,fwdeny,fwcontand thefwmapverdict map (FWMAP), generated ininfra.rs(do_firewall,apply_firewall_all,ns_set_joinfor namespace isolation sets). - Internal DNS (standard name
<name>.<namespace>.svc.delonix.internal, the older<name>.<namespace>.delonix.internalstill answers;service_fqdn,parse_internal_name):dns_server_main,handle_dns,dns_resolve_for,dns_resolve_multi_forininfra.rs. - Overlay networks:
set_vxlan(infra.rs) and the WireGuard helpers incrates/adapters/delonix-sdn/src/wg.rs, orchestrated byrealize_overlayinbins/delonix-runtime-bin/src/cmd/network.rs. - CNI:
crates/adapters/delonix-sdn/src/cni.rs—add,del,readiness,attach_named_netns. Rootless use is opt-in (enabled_confchecksDELONIX_CNI=1); the root CRI path uses the node's CNI chain (root_cni_readinessincrates/interfaces/delonix-cri/src/runtime_svc.rs). - Topology decisions: ADR-0013, ADR-0014.
Read more: network_namespaces(7),
veth(4),
nftables wiki,
slirp4netns,
kernel — VXLAN,
WireGuard,
CNI and its specification.
4.6 Kubernetes: CRI, kubelet, kubeadm and kind#
The kubelet is the Kubernetes node agent. It does not run containers itself; it talks to a
container runtime over the Container Runtime Interface, a gRPC API (RuntimeService,
ImageService) served on a local Unix socket. The kubelet creates a pod sandbox
(RunPodSandbox) and then containers inside it. It also has a cgroup driver setting
(systemd or cgroupfs) that must match how the runtime manages cgroups, or pod cgroups and
container cgroups diverge.
kubeadm bootstraps a cluster on existing machines (kubeadm init, kubeadm join).
kind runs Kubernetes nodes as containers built from the kindest/node image.
In Delonix
crates/interfaces/delonix-criis a CRIruntime.v1server. The protobuf isproto/api.protoinside that crate, compiled bybuild.rswithtonic-build. The binary issrc/bin/delonix-cri.rs.- The cgroup driver reported to the kubelet:
engine_cgroup_driverinruntime_svc.rs; its doc comment records why the answer is what it is and what would have to change for the other one. - A round-trip over real gRPC is tested in
crates/interfaces/delonix-cri/tests/grpc_status.rs. - Cluster bootstrap commands: kubeadm over SSH in
bins/delonix-runtime-bin/src/cmd/cluster.rs(withkubeadm_config.rs,etcd.rs,lb.rs), and kind-style local clusters inkindmode.rs.
Read more: Kubernetes — Container Runtime Interface, cri-api repository, Configuring a cgroup driver, kubeadm, kind.
4.7 Virtualization: KVM, virtio, Cloud Hypervisor, libvirt, cloud-init#
KVM is the kernel's hypervisor, exposed as /dev/kvm; a user-space VMM (QEMU, Cloud
Hypervisor) uses it to run guests. virtio is the family of paravirtual devices (disk, network,
9p filesystem sharing) guests use for efficient I/O. Cloud Hypervisor is a Rust VMM focused on
cloud workloads that can run unprivileged with access to /dev/kvm; it boots guests through a
firmware (an EDK2 UEFI build or rust-hypervisor-firmware) or directly from a kernel image.
libvirt manages QEMU/KVM domains described in XML, via virsh and libvirtd.
Cloud images are generic; per-instance configuration (hostname, SSH keys, users, network) comes
from cloud-init, which reads a datasource. The NoCloud datasource is a small ISO labelled
cidata holding user-data, meta-data and optionally network-config.
In Delonix
- The port is
VmBackendincrates/adapters/delonix-vm/src/lib.rs, implemented byCloudHypervisorBackendandLibvirtBackendthere and byProxmoxBackendincrates/providers/delonix-proxmox(ADR-0008). - Cloud Hypervisor firmware search order:
DEFAULT_CH_FIRMWARES(EDK2CLOUDHV.fdbeforehypervisor-fw); the VMM command line is built inboot_ch. - libvirt domain XML:
libvirt_domain_xml. - NoCloud seed generation:
generate_seed_isoinbins/delonix-runtime-bin/src/cmd/vm.rs. - VMs on the rootless SDN get a DHCP lease derived from their MAC:
dhcp_lease_ipincrates/adapters/delonix-sdn/src/infra.rs. - Practical setup and known host pitfalls: MicroVM setup.
Read more: kernel — KVM, virtio specification (OASIS), Cloud Hypervisor and its documentation, libvirt, cloud-init NoCloud.
4.8 Declarative reconciliation#
Kubernetes popularised a model where users submit desired state as typed objects
(apiVersion, kind, metadata, spec), and controllers repeatedly compare it with actual
state and act to converge. kubectl apply adds a three-way diff: it stores the last
applied configuration on the object, so it can tell "you removed this field from your file"
(revert it) from "someone set this field by hand" (leave it alone).
The principle, and what it asks of you when you add a Kind, are in IaaS and cloud native — Declarative and convergent.
In Delonix
- The engine has its own Kinds in API groups (
delonix api-resourceslists them). Facts about each Kind (domain, whether it converges, teardown, namespacing) live in one table:KindFactsincrates/contexts/delonix-stack/src/kinds.rs. - The planner is pure:
plan(desired, actual, stack)incrates/contexts/delonix-stack/src/reconcile.rs. The module doc has the three-way truth table, and the last-applied spec is stored on the resource itself under theLAST_APPLIEDannotation (delonix.io/last-applied) — there is no separate state file. - Ownership is a label on the resource; revision history is in
revision.rs(ADR-0019). - Manifests are parsed in
bins/delonix-runtime-bin/src/cmd/manifest.rs;stack plan/applylive incmd/stack.rs. - There is no controller loop running in the background: reconciliation happens when a command runs (daemonless). The proposed pull reconciler keeps that property by being a systemd timer that invokes the same apply, not a resident process (ADR-0021, status Proposed).
Read more: Kubernetes — Objects,
Controllers,
Declarative management with kubectl apply.
4.9 Observability and the MCP interface#
OpenTelemetry is a CNCF standard for traces, metrics and logs, exported over OTLP to a
collector. Prometheus scrapes metrics from an HTTP /metrics endpoint in a text exposition
format. The Model Context Protocol is an open protocol that lets AI clients discover and call
tools exposed by a server, commonly over stdio.
In Delonix
- Structured logging and optional OTLP spans:
initincrates/adapters/delonix-telemetry/src/telemetry.rs(spans are exported only whenDELONIX_OTLP_ENDPOINTis set; the exporter runs on its own thread so the synchronous CLI needs no async runtime). - Prometheus registry (prefix
delonix) and text encoding:crates/adapters/delonix-telemetry/src/metrics.rs(encode). The local management API serves/metricsincrates/interfaces/delonix-mgmt/src/lib.rs. - MCP:
crates/interfaces/delonix-mcp(built onrmcp, stdio transport), with the binary inbins/delonix-mcp-bin. Scope and limits: ADR-0025.
Read more: OpenTelemetry documentation, OTLP specification, Prometheus — exposition formats, Model Context Protocol.
4.10 Daemonless, in one paragraph#
containerd and the Docker Engine keep a resident daemon that owns container state; Podman showed
that a runtime can instead be a command that exits, with per-container helper processes and
systemd for anything that must persist. Delonix follows the second model: the CLI does the work
and exits, state is files under the state root guarded by flock (see
Rust primer §3.8), a supervisor process
exists per detached container, the network holder exists only while something needs it, and boot
persistence is systemd units (bins/delonix-runtime-bin/src/cmd/boot.rs). The consequences —
good and bad — are discussed in Architecture and
System design interview. The principle itself, and the rule that a new
daemon needs an ADR, are in IaaS and cloud native — Daemonless.
Next: Rust primer for this codebase — the Rust this codebase is written in — workspace, errors, traits as ports, unsafe, serde, clap, concurrency and tests — pinned to real files.