Practical knowledge. Stronger foundations.
YOUR PREPARATION TOOLKIT

Practical labs

Build something you can explain.

These are exercise specifications for your own isolated environment. Lab completion is your self-assessment; the site does not run or inspect your infrastructure. The Python reference uses only the standard library and makes no network or device changes.

Use disposable VMs or an isolated lab network, compatible authorized images, and synthetic data. Keep experimental DHCP, routing, load generation, and certificate trust away from everyday and production networks.

01 · Python planner

Time: 60–90 minutes. Open the coding exercise below, implement before looking at the solution, and run the tests. All code is standard-library, local, and read-only with respect to networks/devices. The reference suite has 22 tests.

Deliver: your implementation, test output, two intentionally broken safety checks caught by tests, and a two-minute explanation of the boundary between local deterministic planning and distributed execution. Interview target: advanced Python engineering rather than merely scripting a sequence of SSH commands. References: PY-IPADDRESS, PY-MOCK, HTTP-SEM.

Coding exercise, inputs & reference files

Download the student workspace, implement planner.py, and run python3 -m unittest -v test_planner.py. The starter intentionally fails until you implement it. Keep the reference solution closed on your first attempt.

Time box: 45 minutes implementation, 15 minutes tests, 10 minutes explanation. Work without AI for the first attempt. Python 3.10+ and standard library only. Do not read reference_planner.py first.

Implement a pure function that receives desired and observed JSON snapshots, an expected observed revision, and an explicit permission to remove prefixes. It returns an ordered plan of additions/removals per tenant/device. It never connects to a device or writes configuration.

Input contract

Use the two example JSON files. Required top-level fields: complete, revision, devices. Device fields: tenant, device_id, allowed_prefixes. Completeness must be the actual boolean true; a string or integer is invalid. Revisions and identities must be explicit nonempty strings. Prefixes require explicit CIDR notation and must represent network addresses, with no host bits. IPv4 and IPv6 are supported.

The same private prefix may exist in different tenants. Duplicate tenant/device identities, duplicate normalized prefixes, invalid input, stale observation revision, and incomplete inventories are errors. Desired/observed managed device scope must match: this exercise does not create or decommission devices. Prefix removals require explicit permission; that permission must not waive any other validation.

Treat allowed prefixes as an order-insensitive set of exact network entries. This is NOT an ordered first-match firewall rule engine. Do not merge or collapse overlapping prefixes as a hidden optimization. Sort deterministically. Do not mutate the inputs. Include desired/observed revision and normalized fingerprints in the plan. The same effective inputs must produce the same output.

Required demonstrations

Show an addition, rejected removal, approved removal, stale revision, incomplete page result, duplicate tenant/device, invalid prefix, equivalent IPv6 normalization, overlapping tenant addresses, reordered data producing the same plan, and a second run after simulated convergence producing no changes. Show your tests failing when you intentionally remove a safety check.

Run the reference only after your attempt

python3 -m unittest -v
python3 reference_planner.py desired.json observed.json --expected-observed-revision obs-001

On Windows, py -3 may replace python3.

Discussion that separates the levels

A basic solution computes sets. A stronger solution validates completeness, scope, and revision. A senior explanation identifies the limits: a flag does not prove a remote API was genuinely complete; a fingerprint is not a distributed lock; the world may change after planning; an executor needs current-state preconditions, authorization, idempotency/reconciliation, and per-device recovery. A local unit test cannot prove vendor API compatibility or on-wire correctness.

The reference has no network, database, credentials, or device writes. It is a teaching implementation, not a ready-made production network controller. Its CLI errors return status 2; a successful plan returns status 0. Documentation-only addresses are intentional. No external packages are needed.

Extension — 60 minutes

Build a fake in-memory executor with these events: success; known rejection; write accepted then acknowledgment lost; observer unavailable; concurrent revision change. Persist an operation record to a local temporary JSON file. Do not add real credentials or device APIs. Prove that unknown is distinct from failure and that restart/reconciliation does not blindly replay a potentially successful write. Explain the gap between this simulation and a distributed transactional system.

References: PY-IPADDRESS, PY-TUTORIAL, PY-MOCK, HTTP-SEM, NAUTOBOT-API. See the source catalog.

After your attempt: reference solution

Compare the implementation and test coverage after writing your own planner.

Sign in to save private evidence notes and your completion checklist.

02 · Alma service incident — isolate three different causes

Time: 60–90 minutes. Use one disposable Alma VM and a trivial nonprivileged local HTTP service. Record OS/kernel, service definition, user, port, resource controls, and expected response. Select a currently compatible package path; do not disable SELinux to make setup easier.

Create one fault at a time: a wrong working directory or file permission; a service bound to loopback when the test expects a remote path; and a constrained cgroup memory limit for a controlled memory-using process. Keep any resource-pressure experiment strictly inside the VM and set a bounded size/time. Have a console before changing the service or network.

For each fault: preserve the symptom, identify a discriminating command/result, state two alternatives the result rejects, fix the actual cause, and demonstrate the same request succeeds under the intended user and constraints. Do not count a root-shell success as service recovery. Explain why “more RAM,” “open every port,” and “chmod 777” are not interchangeable remedies.

Deliver: three small incident notes, exact service context, commands, outputs, and before/after tests. Pass criterion: three distinct causal explanations, not one generic restart procedure. References: LINUX-CGROUP, SYSTEMD, RHEL-SELINUX, LINUX-NETNS.

Sign in to save private evidence notes and your completion checklist.

03 · Puppet/Hiera — controlled convergence

Time: 90–120 minutes if a compatible lab already exists. Otherwise spend 45 minutes on a catalog/data-flow design rather than forcing an installation under interview pressure.

Use a supported Puppet/agent version combination and one disposable node. Current Puppet and Foreman ecosystem versions differ; establish the version matrix first. Build a small profile that owns a package, a configuration file with explicit mode/owner, and a service. Use a real supported service validation mechanism before accepting new configuration. Put an environment-specific value in Hiera and show how the lookup selects it.

Run once to converge, then again without changes. Change one Hiera value and predict the exact resource effects. Inject an invalid configuration and demonstrate that validation prevents a harmful replacement or service action. Introduce one intentional manual change and explain how the next run reconciles it. Explain what happens when an earlier resource succeeds and a later dependent resource fails; do not claim general automatic rollback.

Deliver: minimal manifest/profile, Hiera example, dependency graph, first/second-run evidence, and a 90-second comparison with your own experience. Redact secrets; do not reuse any actual prior employer module. Pass criterion: explain ordering versus notification, which resource owns the file, and why the second run is quiet. References: PUPPET-ARCH, PUPPET-REL, PUPPET-HIERA, PUPPET-EXEC.

Sign in to save private evidence notes and your completion checklist.

04 · Foreman provisioning — stage-boundary troubleshooting

Time: 60-minute design drill or a longer authorized lab. Draw host intent → DHCP/boot → install content → first boot → trust enrollment → configuration convergence → readiness. Identify every DNS, DHCP, template, proxy, certificate, content, and network dependency.

Given this synthetic exhibit: “DHCP lease acquired; UEFI bootloader fetched; installer starts; template URL returns HTTP 403.” State the last successful stage, whether you would begin by replacing DHCP, and what request identity/path and authorization evidence you need next. Then change the exhibit to “no DHCP reply reaches the client” and give a different investigation.

A live lab must use an isolated DHCP segment and a supported host/proxy/content combination. Foreman does not imply Katello is installed. Confirm actual host-group inheritance and proxy/subnet associations. A provisioning success flag must not substitute for application readiness.

Deliver: stage diagram, one failure timeline, effective host parameters, and a definition of “ready for traffic.” Pass criterion: every proposed check has a result that changes the next step. Reference: FOREMAN.

Sign in to save private evidence notes and your completion checklist.

05 · Routing and MTU — one path, two directions

Time: 90 minutes in an existing isolated routing lab. Use three routers or router VMs plus two endpoint VMs. FRR or Linux routing can teach protocol behavior; it does not prove Junos/QFX, NSX, or firewall-vendor operational experience. Do not import full public Internet tables or peer a lab with a production provider.

Advertise only documentation/test prefixes. Demonstrate an intended route, then a route that is received but cannot forward because the next hop or route installation is wrong. Record control-plane and forwarding evidence separately. Add two paths and explain the outbound policy decision; write a negative test proving that a transit-learned prefix is not exported to another simulated provider.

Next, lower one lab path MTU and compare a small transaction with a larger one. Establish what ICMP/PMTUD behavior actually occurs; do not assume a packet disappears because one capture lacks it. Capture at two points and compare direction, size, and counters. Restore the lab configuration and repeat the test.

Deliver: topology, expected routes, actual RIB/FIB evidence, bounded captures, a safe rollback, and one test for each injected fault. Pass criterion: explain why BGP Established is insufficient and why an MSS workaround is not a universal UDP fix. References: RFC-BGP, RFC-BGP-OPS, RFC-PMTUD, TCPDUMP.

Sign in to save private evidence notes and your completion checklist.

06 · TLS trust and proxy reasoning

Time: 60-minute whiteboard plus an optional isolated TLS lab. Draw client → inspection proxy → origin. Place each certificate, private key, trust store, server name, ALPN choice, and validation step. Do not import a lab interception CA into the everyday browser/system trust store; use an isolated client/profile/VM and remove the lab environment afterward.

Create or reason through four cases: valid server-authenticated sessions; wrong origin hostname; client certificate authentication; and a client with a pinned server key. Explain which leg fails and why adding a broadly trusted CA is not a universal fix. Explain why a CONNECT relay is not automatically a TLS terminator and why a capture plus an origin signing key does not recover ephemeral TLS 1.3 session secrets.

Deliver: two-leg diagram, expected failure at each boundary, a trust-validation test, and an approved-exception record template including owner, scope, risk, visibility, and expiry. F5 SSLO hands-on work requires legitimate compatible lab access; the public training link does not provide a license. References: RFC-TLS, HTTP-SEM, F5-SSLO.

Sign in to save private evidence notes and your completion checklist.

07 · Duplicate jobs and ambiguous writes — paper first, then simulation

Time: 60 minutes. You have a source-of-truth database, message queue, worker, and device API. Walk these crash points: before DB commit; after DB commit but before publish; after publish but before consumer acknowledgment; after device write but before response; after recording success but before queue acknowledgment.

For each, state what persists, who can observe it, whether retry can duplicate a side effect, and what operation ID or invariant is needed. Explain publisher confirmation versus consumer acknowledgment. Show how a local transactional outbox addresses one boundary while not creating global exactly-once execution.

Optional implementation: extend the local Python planner with a fake in-memory device and a local temporary operation journal. Inject acceptance-with-lost-acknowledgment and process restart. No live APIs or credentials. Pass criterion: unknown outcomes remain explicit and recovery does not blindly overwrite newer operator changes. References: RABBIT-ACK, POSTGRES-ISO, HTTP-SEM.

Sign in to save private evidence notes and your completion checklist.

08 · Network-performance experiment design

Time: 45 minutes, no load generator required. Synthetic observation: “32 cores, average CPU 15%; one receive core 100%; RX queue 3 drops climb; 64-byte-heavy workload; large-flow benchmark is healthy.” Produce at least three hypotheses and the next evidence for each.

Expected distinctions: aggregate versus per-core capacity; queue/flow distribution; hardware drops versus software processing pressure; burst buffering versus sustained service rate; IRQ/app memory locality; offload effects on captures. Do not assert that one setting fixes all cases. Define a controlled experiment with unchanged workload, rollback, p99 latency, loss, CPU, and inspection-correctness gates.

Deliver: a before/after measurement table with units and the decision rule. Pass criterion: a throughput gain is not accepted automatically when latency or security correctness regresses. References: LINUX-SCALING, LINUX-CONNTRACK, WIRESHARK-OFFLOAD, SRE-SLO.

Sign in to save private evidence notes and your completion checklist.

09 · Cross-team incident tabletop

Time: 20 minutes. Synthetic observation: network sees traffic reaching the host; infrastructure sees a healthy process; one tenant's requests still fail. Use a single timestamp and flow, identify the application success criterion, and trace socket, TLS, proxy, dependency, and response boundaries. Assign one evidence-gathering task to each team without assuming ownership of the fault.

Deliver: a six-line incident update separating known facts, leading hypotheses, mitigation, residual risk, next experiment, and owner. Pass criterion: you can state what would disprove your current hypothesis and what evidence justifies a permanent correction. Reference: HTTP-SEM.

Sign in to save private evidence notes and your completion checklist.