Practical knowledge. Stronger foundations.
YOUR PREPARATION TOOLKIT

Mock panel & story practice

Five truthful engineering stories

Prepare five examples from your actual experience: a migration or recovery; an automation tool; a configuration rollout; a difficult fault investigation; and a cross-team or architectural tradeoff. For each, write the environment, constraint, your ownership, system or packet path, difficult decision, verification, measured outcome, and lesson.

Use a two-minute core narrative with a technical appendix you can expand under questions. Separate your own implementation from work you inherited. Distinguish a design target from a tested recovery result, and a lab from production experience. Keep confidential identifiers and customer data out of your notes.

Introduction template: “I work at the intersection of [your actual strengths]. In [sanitized example], I owned [specific work] under [constraint]. I verified it through [evidence] and learned [lesson]. I am interested in this role because [connection to the work].” Use your own voice and examples.

A 75-minute rehearsal

Minutes 0–5: introduction and role interest. Minutes 5–15: one production story, interrupted for technical depth. Minutes 15–30: Linux, configuration, or provisioning incident. Minutes 30–45: network, TLS, or stateful incident. Minutes 45–60: Python planner, error handling, and code critique. Minutes 60–70: cross-team architecture and partial failure. Minutes 70–75: your questions and close.

Rehearse interruptions: Why? What would you observe next? What result would disprove that? What if the write succeeded but the reply was lost? How does rollback work if management access is gone? Which part did you personally build?

For practice, score correctness, concrete evidence, causal troubleshooting, safe change design, maintainability/testing, and communication from 0 to 3: absent; vocabulary only; sound applied explanation; tradeoffs plus verification. This is a self-assessment exercise, not an employer rubric.

Ask the panel about the first 90 days, team ownership boundaries, deployed versions, testing and release workflow, source-of-truth ownership, service verification, and on-call expectations. Choose the questions that help you understand the work.

Four architecture whiteboards

These are practice scenarios, not predictions of exact panel questions. Use 12–15 minutes each and reserve the last two minutes for failure handling.

Scenario A — provision 200 private Linux service nodes

Prompt: provision a fleet consistently across multiple facilities without public-cloud tools. Some nodes are bare metal. Rollouts must not interrupt every redundant service member together.

Start by asking for service placement, existing Foreman/Puppet versions, trust/PKI, image/package sources, boot networks, number of sites, capacity margin, out-of-band access, and readiness criteria. Propose an inventory/source-of-truth boundary, approved versioned configuration, site-local provisioning dependencies where needed, authenticated enrollment, Puppet convergence, and traffic admission after verification. Explain what happens when DNS works but DHCP fails, when installation completes but the catalog fails, and when one site loses the central controller.

Strong answer: limited privilege, tested templates, compatible packages, identity lifecycle, clear stage transitions, complete per-host outcomes, and an independent recovery route. The fleet does not enter service just because the installer returned success. A central dependency has a declared availability/recovery strategy. Avoid introducing a dozen unfamiliar systems when existing tools could meet the requirement.

Panel follow-ups: How do rebuilt hosts avoid inheriting stale credentials? How do you stop a bad default from reaching every site? What is the blast radius of a wrong Hiera value? How do you prove a restored controller has consistent database, content, and certificate state? FOREMANPUPPET-ARCHPUPPET-HIERAPOSTGRES-BACKUP

Scenario B — automate a tenant network change

Prompt: create a tenant service spanning a source of truth, two switches, a firewall, and an inspection service. The second switch times out after accepting its write.

Model desired state and validation first: scoped addresses, VLAN/VNI mappings, allowed route export, policy order, and supported platform versions. Produce a deterministic plan and establish a coherent input revision. Explain a per-target execution state machine: planned, authorized, in progress, succeeded, failed, unknown, and recovered. Use local device transactional features where available, but do not claim a global ACID transaction across unrelated devices.

Strong answer: bounded concurrency, a canary gate, a durable change identifier, unknown-result reconciliation, protection against concurrent writers, and service-level validation. A rollback to an old full configuration must not erase unrelated legitimate changes. A failed inventory read must not become an empty inventory. An accepted API request is not operational completion.

Panel follow-ups: What if the orchestrator dies halfway through? What if the same event is delivered twice? What if a device's running configuration changes outside the platform? What if the database commits but the queue is unavailable? HTTP-SEMNETCONFNAUTOBOT-APIRABBIT-ACKPOSTGRES-ISO

Scenario C — TLS application outage after security-path change

Prompt: new flows from most clients work; one application fails; existing sessions also reset during failover. Routing adjacencies remain healthy.

Separate the symptoms rather than forcing one cause. For the failing application, identify transport, TLS version, ALPN/SNI, mutual authentication or pinning, proxy behavior, and upstream certificate validation. For failover resets, examine state synchronization, NAT identity, return path, and connection draining. Select exact timestamps and both original/translated tuples before collecting data.

Strong answer: a packet/path model, independent verification on both TLS legs, scoped captures, and explicit inspection-exception ownership. Health checks must test the relevant service, not merely the management address. “Just bypass the security stack” is not a complete decision; specify authorized traffic class, risk, observability, and expiration.

Panel follow-ups: Why does having the certificate private key not guarantee passive TLS 1.3 decryption? What happens to mTLS client identity at the proxy? Why might an HTTP/3 client follow a different handling path? RFC-TLSRFC-QUICF5-SSLOLINUX-CONNTRACK

Scenario D — intermittent loss on a high-throughput Linux node

Prompt: average utilization is modest, one core is hot, drops increase during small-packet traffic, and a buffer increase improves throughput but worsens tail latency.

Define the load profile and compare per-queue counters, IRQ/softirq CPU, NUMA placement, ring/backlog behavior, flow distribution, and application work. Establish whether drops occur before capture, in host processing, in policy, or downstream. Separate a single flow from aggregate flows. Make one controlled change under repeatable load, with correctness and latency gates.

Strong answer: an evidence table with each hypothesis, distinguishing measurement, low-risk experiment, and acceptance criterion. Do not assert RSS, RPS, offloads, huge buffers, or faster BFD is a universal fix. Keep security visibility and stateful correctness in the benchmark.

Panel follow-ups: Why does idle aggregate CPU not disprove a CPU bottleneck? What does an offloaded packet look like in a host capture? How can a queue reduce loss while making users less happy? LINUX-SCALINGWIRESHARK-OFFLOADSRE-SLO