DDelonix RuntimeContributor handbook

Foundations

Cloud native primer

Before you read: Linux foundations (namespaces, cgroups v2, file descriptors) and IaaS and cloud native (what the engine is responsible for).

The engine is a thin, careful layer over Linux kernel features and a handful of open specifications. This page gives you just enough of each concept to read the code, and says where it lives in this repository. After it you can take any mechanism — a namespace, a cgroup limit, an image layer, a firewall chain, a VM boot, a manifest apply — and name the file and symbol that implements it. For depth, follow the official links — they are better than any summary here.

Concepts are taught once in this handbook. The kernel primitives (namespaces, user namespaces, cgroups v2) are taught hands-on in Linux foundations, so sections 4.1 and 4.2 only recap them and map them to the code. What each open standard requires, and how far the engine conforms, is in Cloud native standards, later in the course. Paths name crates (crates/<layer>/<crate>); the layers are explained in Architecture, and for now a path simply says where the code is.

Each section has three parts: the concept, In Delonix (files and symbols you can grep), and Read more. Paths are relative to the repository root.

For orientation in the wider ecosystem, the CNCF Landscape and the CNCF Glossary are good maps. Comparisons with runc, crun, containerd or Podman appear only where they help explain a design choice.


4.1 Linux namespaces and rootless operation#

Recap. A container is a process started in a fresh set of namespaces (mount, PID, network, IPC, UTS, cgroup, user). The user namespace is what makes it rootless: inside, the process is uid 0 over resources that namespace owns; outside, it is an ordinary user. An unprivileged user can map only its own uid; a range needs newuidmap/newgidmap and /etc/subuid. Both are taught, with commands to type, in Linux foundations — Namespaces and User namespaces and uid mapping. Some distributions additionally restrict unprivileged user namespaces through AppArmor — see Environment for the practical consequences.

In Delonix

  • fn spawn in crates/adapters/delonix-linux/src/lib.rs builds the CloneFlags (CLONE_NEWNS, CLONE_NEWPID, CLONE_NEWNET, CLONE_NEWUSER, …) and calls nix::sched::clone. IPC/UTS sharing between pod members is handled by setns in container_init.
  • write_userns_maps in the same file writes the maps from the parent: a single-uid map (0 <euid> 1) rootless, or a subuid range through newuidmap/newgidmap when have_subid_helpers() says they are available. USERNS_UID_BASE/USERNS_RANGE define the range used when running as root.
  • setup_rootfs mounts the container's root and calls pivot_root; container_init is the code that runs inside the new namespaces before execvp.
  • Rootless operations on files owned by mapped subuids re-execute the binary inside a mapped user namespace: reexec_mapped, reexec_mapped_hold, remove_tree_mapped.

Read more: namespaces(7), user_namespaces(7), newuidmap(1), subuid(5), pivot_root(2).


4.2 cgroups v2 and delegation#

Recap. cgroup v2 is one tree at /sys/fs/cgroup whose files (memory.max, cpu.max, pids.max, memory.events) limit and account for a group of processes. An unprivileged user can only write a subtree systemd has delegated to them, and a shell in an SSH session scope sits outside it — so limits can silently not apply there. The tree, the "no internal processes" rule and delegation are taught hands-on in Linux foundations — cgroups v2; the delegation contract as a standard, and the engine's conformance, are in Cloud native standards — 13.15. The engine is v2-only.

In Delonix

  • Root mode places containers under delonix_compute::DELONIX_SLICE (/sys/fs/cgroup/delonix.slice).
  • Rootless mode finds the user's service cgroup and creates leaves under <user@uid.service>/dlx-containers — see user_service_base and try_delegated_base in crates/adapters/delonix-linux/src/lib.rs. cgroup_limits_apply answers "will limits apply on this host?" without starting a container.
  • The design decision about the intermediate level is ADR-0015; how the CRI follows the kubelet's cgroup hierarchy is ADR-0038, with the kubelet's parent validated by KubeCgroupParent::parse in crates/contexts/delonix-compute/src/record.rs.

Read more: Linux kernel — Control Group v2, systemd — Control Group APIs and Delegation, cgroups(7).


4.3 Capabilities, seccomp, AppArmor, masked paths#

Root's power is split into capabilities (CAP_NET_ADMIN, CAP_SYS_ADMIN, …). A container keeps a small default set and drops the rest. seccomp installs a BPF filter that allows or denies system calls, reducing kernel attack surface. AppArmor (and SELinux on other distributions) are Linux Security Modules that confine a process by profile. Finally, runtimes mask sensitive /proc and /sys paths (bind something empty over them) and make others read-only, because those files leak host information or allow host control.

A subtle point the code documents: clone3 passes its flags through a pointer that a seccomp filter cannot inspect, so a filter that blocks clone(CLONE_NEWUSER) must also make clone3 fail with ENOSYS to force libc back to the filterable clone.

In Delonix

Read more: capabilities(7), kernel — Seccomp BPF, AppArmor documentation, OCI runtime spec — Linux config (the maskedPaths/readonlyPaths/seccomp fields other runtimes consume).


4.4 OCI images, content-addressed storage and overlayfs#

The Open Container Initiative publishes three specifications:

  • the image spec — an image is a manifest (JSON) that points to a config and an ordered list of layers (tarballs), and possibly an index that points to one manifest per platform;
  • the distribution spec — the HTTP API registries serve (/v2/<name>/manifests/<ref>, /v2/<name>/blobs/<digest>, token auth);
  • the runtime spec — how a runtime such as runc or crun is told to run a filesystem bundle.

Everything is content-addressed: a blob is named by the SHA-256 digest of its bytes, so a client verifies what it downloaded by hashing it. Pulling by digest (name@sha256:…) is only a guarantee if the manifest is checked against that digest as well as each blob against the manifest.

At run time, layers are stacked with overlayfs: read-only lowerdirs, a writable upperdir where changes are copied up, and a workdir. Many containers can share the same lower layers.

In Delonix

  • Registry client (distribution spec): crates/adapters/delonix-oci/src/registry.rs — pull_from_registry* functions, the ACCEPT_MANIFEST media types, and verify_manifest_digest. Types come from the oci-spec crate.
  • Content-addressed blob store: Cas in crates/adapters/delonix-oci/src/cas.rs.
  • Writing an OCI image layout archive: write_oci_archive in save.rs.
  • Overlay preparation: ImageStore::prepare_overlay in overlay.rs writes an overlay-lowers marker (LOWERS_FILE); the mount itself happens inside the container's user and mount namespaces in mount_overlay_if_marked (crates/adapters/delonix-linux/src/lib.rs), using the new mount API — see ADR-0016 and ADR-0037.
  • The engine runs containers itself rather than handing an OCI runtime bundle to runc/crun.

Read more: OCI image spec, OCI distribution spec, OCI runtime spec, kernel — Overlay Filesystem.


4.5 Container networking#

Linux networking building blocks:

  • a network namespace has its own interfaces, routes and firewall;
  • a veth pair is a virtual cable with one end in each namespace;
  • a bridge is a virtual switch joining many veth ends;
  • nftables is the kernel packet filter and NAT engine; DNAT rewrites a destination (how a published port reaches a container), and conntrack tracks flows so reply traffic of an allowed connection passes (ct state established,related);
  • slirp4netns gives an unprivileged network namespace outbound connectivity by emulating a TCP/IP stack in user space, and forwards host ports into it;
  • VXLAN carries L2 frames over UDP between hosts, and WireGuard encrypts a tunnel.

CNI (Container Network Interface) is a spec where a runtime executes plugin binaries (bridge, host-local, portmap, …) with ADD/DEL commands and a JSON config from /etc/cni/net.d. Kubernetes runtimes use it for pod networking.

In Delonix

  • Rootless networking cannot create interfaces on the host, so the engine keeps a long-lived holder network namespace: a minimal pin process owns the namespaces, and a restartable control process serves a Unix socket. See start_pin, start_control and ensure_up in crates/adapters/delonix-sdn/src/infra.rs.
  • Attaching a workload: attach_container (IPAM + control command) and do_attach (veth into the bridge, inside the holder). Bridge names come from bridge_name in the dependency-free crates/foundation/delonix-net-rules/src/lib.rs.
  • Outbound and port forwarding: slirp_attach and slirp_add_hostfwd in crates/adapters/delonix-sdn/src/lib.rs (which spawn slirp4netns); in-holder publishing in publish_port/do_publish (infra.rs).
  • Firewall: table ip dlxing with base chains fwguard, fwdeny, fwcont and the fwmap verdict map (FWMAP), generated in infra.rs (do_firewall, apply_firewall_all, ns_set_join for namespace isolation sets).
  • Internal DNS (standard name <name>.<namespace>.svc.delonix.internal, the older <name>.<namespace>.delonix.internal still answers; service_fqdn, parse_internal_name): dns_server_main, handle_dns, dns_resolve_for, dns_resolve_multi_for in infra.rs.
  • Overlay networks: set_vxlan (infra.rs) and the WireGuard helpers in crates/adapters/delonix-sdn/src/wg.rs, orchestrated by realize_overlay in bins/delonix-runtime-bin/src/cmd/network.rs.
  • CNI: crates/adapters/delonix-sdn/src/cni.rs — add, del, readiness, attach_named_netns. Rootless use is opt-in (enabled_conf checks DELONIX_CNI=1); the root CRI path uses the node's CNI chain (root_cni_readiness in crates/interfaces/delonix-cri/src/runtime_svc.rs).
  • Topology decisions: ADR-0013, ADR-0014.

Read more: network_namespaces(7), veth(4), nftables wiki, slirp4netns, kernel — VXLAN, WireGuard, CNI and its specification.


4.6 Kubernetes: CRI, kubelet, kubeadm and kind#

The kubelet is the Kubernetes node agent. It does not run containers itself; it talks to a container runtime over the Container Runtime Interface, a gRPC API (RuntimeService, ImageService) served on a local Unix socket. The kubelet creates a pod sandbox (RunPodSandbox) and then containers inside it. It also has a cgroup driver setting (systemd or cgroupfs) that must match how the runtime manages cgroups, or pod cgroups and container cgroups diverge.

kubeadm bootstraps a cluster on existing machines (kubeadm init, kubeadm join). kind runs Kubernetes nodes as containers built from the kindest/node image.

In Delonix

  • crates/interfaces/delonix-cri is a CRI runtime.v1 server. The protobuf is proto/api.proto inside that crate, compiled by build.rs with tonic-build. The binary is src/bin/delonix-cri.rs.
  • The cgroup driver reported to the kubelet: engine_cgroup_driver in runtime_svc.rs; its doc comment records why the answer is what it is and what would have to change for the other one.
  • A round-trip over real gRPC is tested in crates/interfaces/delonix-cri/tests/grpc_status.rs.
  • Cluster bootstrap commands: kubeadm over SSH in bins/delonix-runtime-bin/src/cmd/cluster.rs (with kubeadm_config.rs, etcd.rs, lb.rs), and kind-style local clusters in kindmode.rs.

Read more: Kubernetes — Container Runtime Interface, cri-api repository, Configuring a cgroup driver, kubeadm, kind.


4.7 Virtualization: KVM, virtio, Cloud Hypervisor, libvirt, cloud-init#

KVM is the kernel's hypervisor, exposed as /dev/kvm; a user-space VMM (QEMU, Cloud Hypervisor) uses it to run guests. virtio is the family of paravirtual devices (disk, network, 9p filesystem sharing) guests use for efficient I/O. Cloud Hypervisor is a Rust VMM focused on cloud workloads that can run unprivileged with access to /dev/kvm; it boots guests through a firmware (an EDK2 UEFI build or rust-hypervisor-firmware) or directly from a kernel image. libvirt manages QEMU/KVM domains described in XML, via virsh and libvirtd.

Cloud images are generic; per-instance configuration (hostname, SSH keys, users, network) comes from cloud-init, which reads a datasource. The NoCloud datasource is a small ISO labelled cidata holding user-data, meta-data and optionally network-config.

In Delonix

Read more: kernel — KVM, virtio specification (OASIS), Cloud Hypervisor and its documentation, libvirt, cloud-init NoCloud.


4.8 Declarative reconciliation#

Kubernetes popularised a model where users submit desired state as typed objects (apiVersion, kind, metadata, spec), and controllers repeatedly compare it with actual state and act to converge. kubectl apply adds a three-way diff: it stores the last applied configuration on the object, so it can tell "you removed this field from your file" (revert it) from "someone set this field by hand" (leave it alone).

The principle, and what it asks of you when you add a Kind, are in IaaS and cloud native — Declarative and convergent.

In Delonix

  • The engine has its own Kinds in API groups (delonix api-resources lists them). Facts about each Kind (domain, whether it converges, teardown, namespacing) live in one table: KindFacts in crates/contexts/delonix-stack/src/kinds.rs.
  • The planner is pure: plan(desired, actual, stack) in crates/contexts/delonix-stack/src/reconcile.rs. The module doc has the three-way truth table, and the last-applied spec is stored on the resource itself under the LAST_APPLIED annotation (delonix.io/last-applied) — there is no separate state file.
  • Ownership is a label on the resource; revision history is in revision.rs (ADR-0019).
  • Manifests are parsed in bins/delonix-runtime-bin/src/cmd/manifest.rs; stack plan/apply live in cmd/stack.rs.
  • There is no controller loop running in the background: reconciliation happens when a command runs (daemonless). The proposed pull reconciler keeps that property by being a systemd timer that invokes the same apply, not a resident process (ADR-0021, status Proposed).

Read more: Kubernetes — Objects, Controllers, Declarative management with kubectl apply.


4.9 Observability and the MCP interface#

OpenTelemetry is a CNCF standard for traces, metrics and logs, exported over OTLP to a collector. Prometheus scrapes metrics from an HTTP /metrics endpoint in a text exposition format. The Model Context Protocol is an open protocol that lets AI clients discover and call tools exposed by a server, commonly over stdio.

In Delonix

Read more: OpenTelemetry documentation, OTLP specification, Prometheus — exposition formats, Model Context Protocol.


4.10 Daemonless, in one paragraph#

containerd and the Docker Engine keep a resident daemon that owns container state; Podman showed that a runtime can instead be a command that exits, with per-container helper processes and systemd for anything that must persist. Delonix follows the second model: the CLI does the work and exits, state is files under the state root guarded by flock (see Rust primer §3.8), a supervisor process exists per detached container, the network holder exists only while something needs it, and boot persistence is systemd units (bins/delonix-runtime-bin/src/cmd/boot.rs). The consequences — good and bad — are discussed in Architecture and System design interview. The principle itself, and the rule that a new daemon needs an ADR, are in IaaS and cloud native — Daemonless.


Next: Rust primer for this codebase — the Rust this codebase is written in — workspace, errors, traits as ports, unsafe, serde, clap, concurrency and tests — pinned to real files.