Practical knowledge. Stronger foundations.
YOUR PREPARATION TOOLKIT

Study guide & prep plans

A repeatable structure for technical answers

Start with the symptom and scope. State the relevant invariant or packet/resource path. Name the next discriminating observation and what result would change your hypothesis. Propose a bounded correction. Define service-level verification. Finish with recovery, residual risk, and ownership.

A tool list alone is not a diagnosis. Explain why each observation distinguishes one cause from another. State assumptions about versions, scale, trust, and the failure domain before choosing a product or a command.

Choose your preparation window

Suggested allocations. Adapt them to the target role, your baseline, and the time you have.

48-hour sprint

Budget roughly 12 focused hours. Spend hour 1 reading the target role requirements and outlining five truthful engineering stories. Use hours 2–5 for the two diagnostics. In hour 6, identify your three highest-risk domains using wrong answers, confidence, and ability to explain the mechanism.

Use hours 7–8 for the Python planner and tests, hour 9 for Puppet/Hiera and provisioning, hour 10 for a Linux incident and a packet walk, and hour 11 for TLS, stateful failover, and routing policy. Finish with a mock panel and evidence notes in hour 12.

If time is shorter, use ten-question domain practices as a partial diagnostic. Avoid spending the final night installing a large lab platform or memorizing vendor commands you cannot explain.

Seven-day plan

Day 1: role requirements, infrastructure diagnostic, and evidence map. Day 2: network diagnostic, error journal, and a Linux incident. Day 3: Python planner, tests, and ambiguous-write recovery. Day 4: Puppet/Hiera and provisioning-stage troubleshooting.

Day 5: routing policy, stateful security, and a packet/MTU lab. Day 6: TLS, service chains, monitoring, and an architecture whiteboard. Day 7: full mock panel, repair the two weakest explanations, and a light final review. Budget 2.5–3.5 focused hours per day and adjust after the diagnostics.

Fourteen-day plan

Use the seven-day plan for week one. In week two, complete one isolated Linux/Puppet build and a deliberately broken provisioning stage. Extend the Python planner with a fake executor and crash-recovery tests. Run a routing/MTU lab and rehearse two additional architecture scenarios.

Add exact-vendor hands-on time where a compatible, authorized lab is available. Prepare a small demonstration or annotated design that another engineer can inspect quickly. Retest weak areas with different symptoms and no answer choices.

The technical field guide

12 chapters · Applied reasoning
01Linux: diagnose the resource, not the adjective

Questions: I001–I020, N061–N070.

A “slow host” is not a diagnosis. Start with a timed service symptom and separate CPU execution, scheduling, memory pressure, storage waits, and networking. High load with idle CPU can involve uninterruptible I/O waits; high CPU use can be useful work or wasted contention. A container can hit a cgroup memory or CPU limit even when the host has spare resources. Linux cache is not automatically a leak; available memory, reclaim, faults, pressure, and the application working set provide context. LINUX-MANLINUX-CGROUPLINUX-PSI

Read-only investigation sequence, adjusted for installed tools and access:

uptime
vmstat 1 5
free -h
cat /proc/pressure/cpu /proc/pressure/memory /proc/pressure/io
ps -eo pid,ppid,stat,wchan:24,comm
journalctl -k --since "15 minutes ago"
systemctl status example.service
systemctl show example.service -p User -p WorkingDirectory -p MemoryMax -p CPUQuotaPerSecUSec

Do not dump secret-bearing environment variables or logs indiscriminately. A tool not being installed is not proof that the resource is healthy. /proc, service logs, and cgroup files often provide a lower-level path. A futex wait may be normal synchronization; one stack does not prove deadlock. A zombie is an exited process waiting to be reaped, not a CPU-consuming workload to kill repeatedly. LINUX-MANSYSTEMDLINUX-CGROUP

For “disk full after deleting a large log,” compare allocated space, inode exhaustion, mount/namespace, and open deleted files. df -h, df -i, and a scoped lsof +L1 can discriminate causes. Space for an unlinked but open file is not necessarily reclaimed until the reference is released. Do not truncate arbitrary descriptors or reboot first without understanding the application and evidence retention.

For “works interactively, fails under systemd,” compare effective user, groups, working directory, paths, capabilities, environment sourcing, filesystem access, sandbox restrictions, and SELinux context. A successful root shell test is not equivalent to a constrained service test. LINUX-MANSYSTEMD

Alma/RHEL-specific refresh

Know how to inspect package provenance and enabled repositories, not just install a package. Reproducible hosts need a known package set and compatible repositories; “latest on every run” can create unreviewed change. Treat Alma major-version compatibility, Puppet/Foreman versions, plugins, and Python dependencies as an explicit matrix. RHEL documentation is useful for shared administration concepts, not a promise of identical vendor support on Alma. RHEL-PACKAGESFOREMAN

SELinux troubleshooting begins with the denial and intended access. Check contexts, expected labels, allowed ports, booleans, and whether a supported policy exists. A persistent, scoped label correction is different from an ad hoc chcon that a relabel may undo. Avoid “disable SELinux” and blindly generating broad allow rules from all denials. Prove the original operation succeeds while confinement remains active. RHEL-SELINUX

Exit drill: a process is OOM-killed with host memory available; explain the cgroup evidence. Then diagnose a service failure that is caused by permissions rather than memory. You must not apply the same “increase limits” answer to both.

02Puppet and Hiera: explain convergence and ownership

Questions: I021–I030.

Know the agent/server path: facts and identity inform catalog compilation; the agent applies declared resources and reports outcomes. Compilation and application failures are different. The catalog is a dependency graph, not a top-to-bottom shell script. Distinguish ordering from notification: an ordering edge controls sequence; a notification relationship can also trigger a supported refresh when a resource changes. A Puppet run is not a general atomic transaction that automatically rolls back every earlier successful resource after a later failure. PUPPET-ARCHPUPPET-REL

A role expresses what a node is; profiles compose policy for a function; component modules implement reusable resources. Data belongs in a deliberate Hiera hierarchy with explained precedence and merge behavior. For an unexpected value, trace the lookup with the relevant facts, environment, hierarchy, and lookup options. Do not “fix” five nodes manually before identifying the higher-priority data entry that will reassert the old value. PUPPET-ROLESPUPPET-HIERA

A senior-level service rollout answer should include package/version intent, validated configuration, ownership/mode, service lifecycle, notification behavior, testing, promotion, and health verification. Re-running an unchanged catalog should not repeat destructive work. Prefer native resource semantics; guard necessary execs with an actual state predicate. A creates marker is only meaningful when it accurately represents the operation's resulting state. A false marker can hide incomplete work. No-op is valuable but does not prove every external side effect or application health outcome. PUPPET-EXECPUPPET-REL

Treat CA identity and secrets as engineering concerns. Blind certificate signing is not safe enrollment. Rebuilt nodes require a deliberate identity/certificate lifecycle. Puppet's Sensitive type reduces accidental display; it is not encryption of the value throughout all storage and transport. Explain the actual secret-management boundary rather than claiming a wrapper makes a secret universally safe. PUPPET-ARCHPUPPET-SENSITIVE

Prepare your real story: What did your configuration code own? Which facts and Hiera values selected the resulting configuration? How was it tested and promoted? What happened when a run failed? How did you establish that drift was corrected without an unnecessary service restart? Which part did you build versus inherit? Do not invent missing details; retrieve or reconstruct what you genuinely remember.

Exit drill: draw the graph for package → config → service. Explain whether a changed config restarts or reloads the service, whether that operation is supported, and how application validation prevents deploying invalid content.

03Foreman: provision a machine, then prove a service

Questions: I031–I040.

Think in stages: inventory and intended host identity; network boot and content delivery; OS installation; enrollment/configuration; readiness and handoff. Foreman coordinates host lifecycle and integrations. Smart Proxies can provide local integration points for services such as DHCP, DNS, TFTP, and Puppet-related functions. A proxy existing in the UI does not prove a DHCP request reaches the correct scope or that a host receives compatible boot artifacts. FOREMAN

For provisioning failure, locate the last successful transition. No DHCP lease suggests a different investigation than successful bootloader download followed by a template HTTP error. BIOS and UEFI boot paths can differ. Verify subnet/proxy association, boot method, MAC/identity, reservation, next-server/boot filename, routing, and content URL. Do not keep debugging DHCP after the host has clearly moved on to fetching its install template.

Host groups, parameters, templates, and inherited settings can make builds repeatable or spread a bad default widely. Record the effective configuration, not merely what one screen displays locally. Katello adds content-lifecycle concerns; understand repository synchronization and promoted content versions without assuming every Foreman environment includes Katello. The current documentation is a reference, not proof of the organization’s deployed version. FOREMAN

Define “done” explicitly. An OS reaching its first boot is not the same as a trusted, monitored, configured application node ready to receive traffic. Readiness should require the intended identity, expected package/config revision, successful configuration report, service tests, and controlled enrollment into traffic. Failed bootstrap must not accidentally add an unhealthy node to a live pool.

Isolated lab: management host/proxy plus one disposable Alma VM on an isolated network. Never attach a test DHCP service to your normal LAN, customer network, or production VLAN. Select a supported software combination first; do not spend your final prep day forcing incompatible versions to install. A traced provisioning design and one controlled failure are more valuable than claiming an unverified “fully automated build.”

Exit drill: the VM downloads its bootloader but fails retrieving the installation template. Identify the next three observations, why each matters, and what would make you return to investigating DHCP.

04Python: from script to production application

Questions: I041–I070, N071–N080. Complete the supplied coding lab before studying its reference implementation.

Separate domain logic from I/O. A pure planner accepts validated desired and observed state and returns an explicit plan. Adapters fetch APIs, read inventory, and write devices. This makes invariants testable without a real router. Represent success, known failure, and unknown outcome distinctly. An exception followed by return is dangerous when an empty result means delete everything.

Design data contracts: explicit required fields, types, tenant/namespace scope, canonical addresses, coherent revision, schema version, and inventory completeness. Do not confuse transport success with valid business data. A GraphQL response can contain both data and errors; paginated APIs require complete traversal and validation before using absence as authoritative evidence. PY-IPADDRESSNAUTOBOT-APIGRAPHQL

Use deadlines, not just one vague timeout. Decide connect/read/request limits, an overall operation budget, maximum attempts, and cancellation behavior. Retry only potentially transient failures whose side effects can be reconciled. A timed-out POST may have succeeded. Reuse an idempotency key when the server supports that contract; otherwise reconcile against a durable operation identity. A deterministic authorization or validation error usually needs correction, not ten retries. Honor server throttling and use bounded backoff with jitter. HTTP-SEM

Concurrency is a capacity decision. For I/O-bound work, bounded threads or async tasks can overlap waits. Calling blocking work directly inside an async coroutine can block the event loop. Under a GIL-enabled CPython build, CPU-heavy pure Python is not generally accelerated merely by adding threads; processes or other suitable execution strategies may be appropriate. Free-threaded builds and native extensions change the analysis, so state your assumption. Limits must protect device APIs, database pools, memory, and rate budgets together. PY-ASYNCPY-FUTURESPY-THREAD

Testing must cover wrong and ambiguous outcomes: malformed JSON, missing fields, page two failure, stale revision, duplicate event, timeout after write, partial fleet success, cancellation, invalid certificate, and an idempotent second run. Patch the name where the tested code looks it up, not an unrelated original import location. Avoid mocks that make every operation succeed; include fixtures representing real supported vendor schemas. PY-MOCK

Know common language hazards: mutable default arguments retain shared state across calls; context managers handle cleanup; broad exception swallowing erases error meaning; a compound read-modify-write is not automatically safe because of the GIL. Pass arguments to subprocesses without shell interpolation when possible, validate inputs, and handle nonzero exit status and timeout. Log structured outcome/context, not passwords, bearer tokens, or private keys. PY-TUTORIALPY-EXCEPTIONSPY-THREADPY-SUBPROCESSOWASP-LOGGING

A strong design answer: “The planner is deterministic and has no device credentials. It rejects incomplete inventories and mismatched revisions. The executor uses bounded concurrency, per-target outcomes, time budgets, and a canary gate. A lost acknowledgment becomes unknown, not success or a reason for blind replay. Verification reads the realized state and tests service behavior. The change record joins approval, source revision, intended diff, execution, and recovery.” These are design recommendations; adapt them to the actual system.

Exit drill: in 45 minutes, implement a safe allow-list diff, reject invalid data and unauthorized removals, and write tests. Then explain how your local deterministic function differs from a distributed transaction.

05APIs, databases, caches, and queues: own the failure semantics

Questions: I061–I080.

HTTP 200 does not make every body trustworthy; HTTP 202 is acceptance of an asynchronous operation, not proof of completion. Version or ETag preconditions can prevent overwriting an intervening change. Authentication proves a caller identity; authorization decides whether that caller may act on a particular device, tenant, or operation. Private networking does not eliminate either requirement. Trusted proxy headers need a clear boundary so a client cannot forge its identity by supplying its own forwarding header. HTTP-SEMOWASP-API

For database-backed IP allocation, a Python “check then insert” is not enough under concurrency. Define the correct scoped uniqueness constraint and use appropriate transactions with conflict handling. Do not hold a database transaction open while waiting minutes for a switch to respond. Separate durable intent from slow external work. An outbox pattern can atomically record a local state change plus work to dispatch; it does not by itself make the remote side effect exactly once. POSTGRES-CONSTRAINTPOSTGRES-ISO

Queues separate delivery from processing. Publisher confirmation and consumer acknowledgment are different promises. Acknowledge after the required durable effect, but recognize the crash window after committing the effect and before acknowledging: delivery can repeat. Design idempotent or deduplicated processing using a durable operation identity. Poison messages need bounded attempts, failure classification, a dead-letter path, and a human-visible owner rather than an infinite requeue loop. RABBIT-ACK

For PostgreSQL, know transactions, constraints, indexes, locks, query plans, backups, and restore validation. EXPLAIN ANALYZE executes the statement: treat writes and expensive production queries accordingly. An index helps only if it fits the workload and the planner chooses it; it also has write/storage costs. A backup job's success status does not prove the data can be restored within the recovery target. POSTGRES-EXPLAINPOSTGRES-BACKUP

Redis may be a cache or a deliberately durable component depending on configuration; do not assume its role or persistence. A cache miss is not evidence that the authoritative object was deleted. Elasticsearch search visibility can lag an acknowledged write because search is near-real-time. MongoDB's flexible documents still need intentional schemas, validation, indexes, and operational discipline. Choose stores based on invariants and access patterns, not one universal “SQL is slow” or “NoSQL scales” slogan. REDIS-PERSISTELASTIC-NRTMONGO-VALIDATION

Exit drill: a worker modifies a device, times out, and crashes before recording completion. Explain what can be known, how the next worker finds the operation, how it avoids destructive duplication, and what remains impossible to guarantee without cooperation from the remote system.

06Containers and Windows integration

Questions: I081–I090.

A container is a process isolation and packaging mechanism, not an automatic security boundary equivalent to a separate machine. Exposing the Docker socket to a workload gives it powerful control over the daemon and potentially the host. Explain user privilege, filesystem access, secrets, network exposure, image provenance, and resource limits. Record immutable image identity and configuration revision; a mutable tag is not a reproducible release identifier. DOCKER-SECDOCKER-STORAGE

Durable data needs a deliberate persistence and backup path beyond an ephemeral writable layer. Separate process liveness from readiness to serve work. A dependency failure should not necessarily cause every process to restart in a loop. Docker health information must be integrated with the orchestrator or service-management behavior actually in use; do not assume Kubernetes probe semantics apply automatically to plain Docker. DOCKER-HEALTHDOCKER-STORAGE

For Linux and AD, distinguish identity lookup, authentication, authorization, and federation. DNS/service discovery and time synchronization are important in Kerberos-backed environments. A cached SSSD login succeeding during an outage is not proof the domain is healthy. A valid user identity may still be denied by access policy. AD FS introduces token/claim and trust concerns: validate signatures, intended issuer/audience, lifetime, and protocol-specific conditions rather than trusting an unverified token body. NPS is the Windows RADIUS policy/authentication role, not a generic DNS service. RHEL-ADMS-KERBEROSMS-ADFSMS-NPS

In PowerShell automation, understand terminating versus non-terminating errors, explicit error handling, scoped remoting identity, and structured serialization. Returning text containing “failed” with an apparently successful process status is not a reliable integration contract. In Bash, quote expansions, propagate failures intentionally, avoid parsing human output when structured interfaces exist, and be able to explain pipelines and exit status. set -e is not a comprehensive substitute for error design. MS-POWERSHELLPY-SUBPROCESS

Exit drill: a Linux service account resolves in AD but cannot log in. Split DNS/time/trust, authentication, local access policy, and service-specific authorization before resetting its password or disabling access controls.

07Routing and packet analysis: prove the path

Questions: N001–N040.

Start with the packet, not the protocol name. For an off-subnet IPv4 destination, the host selects a route and usually resolves the next-hop MAC; the Ethernet destination is not generally the remote host's MAC. For IPv6, Neighbor Discovery and other required ICMPv6 functions matter. Blanket ICMPv6 blocking can break neighbor discovery and path MTU behavior. LLDP is discovery evidence, not an infallible substitute for inventory or an enforcement protocol. LACP member health includes actor/partner and collecting/distributing state; an electrically up port can still fail to participate. A single hashed flow does not automatically use every member's bandwidth. RFC-ARPRFC-NDRFC-PMTUDLINUX-BOND

Be able to trace TCP handshake, sequence/acknowledgment progress, retransmissions, receive window, and close/reset behavior. A SYN with no reply at one capture point does not locate the drop. A zero receive window points toward receiver capacity/backpressure; retransmissions can reflect loss, reordering, capture gaps, or other conditions. Correlate both directions, clock accuracy, offloads, and the capture location. SPAN oversubscription and local capture drops can fabricate an apparent loss story. RFC-TCPTCPDUMPWIRESHARKWIRESHARK-OFFLOAD

Useful read-only probes, chosen for a specific hypothesis:

ip -br address
ip route get 192.0.2.10
ip -6 route get 2001:db8::10
ip rule show
ip neigh show
ss -tin
ip -s link show dev eth0
ethtool -S eth0

An interface name, namespace, source address, routing table, and mark can change the answer. Use documentation/test addresses only in an isolated lab. Do not run flood tests or unbounded captures against production. Packet captures may contain credentials and customer content; use authorized scope, bounded size/time, access controls, and a retention decision. LINUX-NETNSLINUX-SYSCTLTCPDUMP

BGP: from adjacency to forwarding

The investigation chain is session state → correct address family → received route → import policy → eligibility/next-hop resolution → best-path selection → export policy → FIB → actual traffic. An Established session or a route in a BGP table proves only one stage. For usual eligible paths, local preference is considered before AS-path length; changing local preference primarily influences your outbound choice, not a remote network's inbound selection. Inbound policy is constrained by what upstreams honor. RFC-BGPRFC-BGP-OPS

Know iBGP propagation and why full meshes or route reflectors are used. Inspect next-hop reachability and intended reflection behavior rather than assuming all established peers receive every route. Know explicit export allow-lists, prefix limits, policy validation, and route-leak negative tests. RPKI origin validation is not full AS-path validation. A route's length, origin authorization, business relationship, and operational reachability are distinct considerations. RFC-RRRFC-BGP-OPSRFC-RPKI

OSPF adjacency states need context. Two DROther routers on a broadcast segment may legitimately remain 2-Way while being Full with the DR/BDR. A stuck ExStart/Exchange state warrants checks such as MTU and neighbor parameters, not indiscriminately restarting the routing process. BFD accelerates failure detection but aggressive timers can flap under CPU or scheduling pressure; a fast false failure is not resilience. RFC-OSPFRFC-BFD

For multi-ISP designs, distinguish loss of link, peer, next-hop, regional connectivity, and the actual service. BGP can remain healthy while a downstream path is broken. RPM/active probes need representative targets, hysteresis, and bounded event behavior. Filter-based forwarding can send traffic to a different routing instance; verify both the selected table and stateful return path. Full tables create memory, RIB/FIB, convergence, and churn requirements; do not memorize a supposedly permanent prefix count. JUNIPER-FBFJUNIPER-RPMJUNIPER-EVENTRFC-BGP-OPS

Exit drill: the BGP session is up, an imported prefix is visible, but customers cannot reach it. Name the next observable that separates an unresolved next hop, missing FIB programming, security drop, and remote return-path failure.

08Stateful security, IPsec, TLS, and application proxies

Questions: N041–N060.

A stateful device evaluates more than an isolated permit rule. Draw the original and translated five-tuples, zones, route lookup, policy decision, session ownership, and return direction. NAT and policy evaluation order varies by platform and rule type; verify the actual semantics. Configuration replication does not necessarily replicate active sessions. A firewall pair may restore new-flow service while breaking existing connections, especially when NAT identity or state moves differently. LINUX-CONNTRACKPAN-NAT

For IKEv2, distinguish the IKE SA from Child SAs and actual protected traffic. A tunnel can negotiate while routes, selectors, NAT exemptions, security policy, or return paths prevent application traffic. Inspect encrypt/decrypt and drop counters on both peers. NAT traversal commonly uses UDP-encapsulated ESP on 4500; native ESP is IP protocol 50, not TCP/UDP port 50. Rekey, lifetimes, overlap, and peer compatibility matter when the tunnel fails only during transitions. RFC-IKERFC-NATT

MTU problems often appear as small requests working while large transfers stall. Added encapsulation reduces payload headroom; inspect inner and outer packet sizes, fragmentation/PMTUD, and error signaling. TCP MSS clamping can help a TCP-specific path but does not solve every UDP or encapsulation problem. Calculate overhead rather than saying “set jumbo frames everywhere.” RFC-PMTUDRFC-VXLANRFC-IKE

TLS inspection requires a correct trust model. In ordinary authorized interception, the proxy terminates one TLS connection and creates another, taking on validation responsibilities at both boundaries. A server certificate private key does not by itself decrypt captured TLS 1.3 ECDHE sessions. Mutual TLS and pinning create additional constraints; certificate trust does not magically preserve the original client's private-key proof through a generic proxy. SNI, ALPN, certificate identity, and HTTP routing are separate concepts. RFC-TLSHTTP-SEM

F5 SSL Orchestrator adds inspection-service-chain concerns beyond generic load balancing: traffic steering, decrypt/inspect/re-encrypt paths, service health, routing/translation, and explicitly governed bypass or drop behavior. Use the official training material to learn terminology, but a public lab guide does not grant a licensed lab environment or establish your production experience. Do not claim that every encrypted application or transport is transparently inspectable. QUIC/HTTP3 requires an explicit supported handling policy. F5-SSLORFC-QUIC

Application protocol transitions matter. SMTP STARTTLS changes the session's encryption state; a byte relay and an application-aware proxy are not interchangeable. TFTP uses UDP transfer identifiers beyond the initial well-known destination port. HTTP CONNECT establishes a tunnel; it does not necessarily terminate the TLS inside. HTTP/2 multiplexing makes one connection a failure domain for multiple streams, and retry safety still depends on request semantics. RFC-SMTP-TLSRFC-TFTPHTTP-SEMRFC-HTTP2

Exit drill: a pinned mTLS application fails only when inspected. Explain why adding another trusted CA may not solve it, what evidence you would obtain, and how an approved exception differs from casually disabling verification.

09Linux network tuning: measure packets, queues, and tail latency

Questions: N061–N070, I011–I020.

Do not start by copying sysctl values. Define the workload: packets per second, bytes per second, flow count/churn, packet-size distribution, encryption/inspection work, connections per second, and latency distribution. One large flow, many short flows, and a small-packet burst can stress different limits despite similar Gbps. Establish hardware, driver, kernel, topology, and offload baseline before a controlled change. LINUX-SCALINGLINUX-SYSCTLSRE-SLO

RSS distributes receive flows across hardware queues using a hash; it does not necessarily split one flow across all cores. RPS/RFS/XPS affect software steering and placement in different ways. Look at per-CPU interrupts/softirq, queue counters, NIC locality, application placement, and memory locality. Thirty idle cores do not help a receive queue saturated on the thirty-first if work is not distributed appropriately. A NUMA-local design can reduce remote access but must be balanced against contention. LINUX-SCALING

Buffering absorbs bursts, not an infinite sustained overload. A larger ring/backlog may reduce drops while increasing queuing latency. Conntrack capacity and timeouts affect memory and valid session behavior; bypassing tracking is a functional/security change, not a harmless speed optimization. Check policy routing, namespace, reverse-path filtering, and MTU before blaming NIC performance. LINUX-CONNTRACKLINUX-SYSCTL

Compute bandwidth-delay product: 1 Gbit/s × 0.080 s = 80 Mbit = 10 MB of data in flight, using decimal units. Achieving that rate in one TCP flow requires suitable effective window/congestion behavior and a path that supports it; the calculation alone does not guarantee throughput. Compare single/parallel streams, both directions, loss, CPU placement, and loaded latency. Host captures may show offload artifacts that are absent on the wire. RFC-TCPLINUX-SCALINGWIRESHARK-OFFLOAD

Exit drill: an adjustment improves aggregate throughput but doubles p99 latency and reduces inspection visibility. Explain the evidence needed to approve or reject it against explicit service objectives. The right answer is not automatically “more Gbps wins.”

10EVPN and named vendor platforms: a bounded ramp plan

Questions: N081–N090. These are preferred-platform preparation, not a reason to neglect core Linux/Python/routing.

EVPN is the reachability control plane in an EVPN-VXLAN design; VXLAN carries encapsulated tenant traffic over the IP underlay. Validate underlay VTEP reachability independently of EVPN route exchange. Route distinguishers make otherwise overlapping routes unique; route targets control import/export membership. Type 2 describes MAC/IP reachability, Type 5 IP prefixes. Tenant-to-VNI mapping, route import, MAC mobility, and multihoming behavior are separate failure points. RFC-EVPNRFC-EVPN-IPRFC-VXLANJUNIPER-EVPN

For a 1,500-byte inner IP packet with untagged inner Ethernet and ordinary IPv4 VXLAN encapsulation, outer IP size is 1500 + 14 + 8 + 8 + 20 = 1550 bytes. That excludes outer Ethernet/FCS and additional tags/tunnels. Device “MTU” and “maximum frame size” settings may refer to different layers. Explain precisely what you counted. For multihoming, do not simplify designated-forwarder election to “only one device can forward every flow.” RFC-VXLANRFC-EVPN

On QFX, prioritize identifying the actual model, Junos release, forwarding resources, and supported features before asserting a design works universally. For NSX-T, distinguish logical intent, realized state, workload routing, north-south connectivity, and edge/stateful services; learn the deployed version's Tier-0/Tier-1 and distributed-firewall model from an authorized current environment or official documentation. Specific NSX documentation retrieval was unavailable during this research; the single exam item is a conceptual topology check, not a release-specific operating procedure. Do not memorize guessed CLI commands. JUNIPER-EVPN

For Panorama, distinguish device groups for policy/objects from templates/template stacks for device/network settings. Check inheritance, local overrides, central commit, intended push scope, and the actual per-device result. A management-plane success does not prove that all target firewalls now enforce the policy. For F5, concentrate on the two TLS legs, service chains, and tested health/failure behavior covered in the previous chapter. PAN-GROUPSPAN-TEMPLATESPAN-COMMITF5-SSLO

Honest bridge answer: “My direct production experience is with [actual platforms]. I understand the stateful routing, TLS, and failure-domain problem. I have not operated [specific product] at your scale. I would verify [two concrete implementation details] before a change, and I can show how I handled the equivalent invariants in [real project].” Replace the brackets; do not read them aloud.

11Source of truth, monitoring, and safe automation

Questions: I091–I100, N071–N080, N091–N100.

A source of truth needs field ownership. Intended VLANs, discovered neighbors, approved prefixes, current health, and observed software versions are not all the same kind of data. Decide who is authoritative for each, how stale data is identified, and how conflicts are reconciled. A discovery process should not silently overwrite approved intent because the current network happens to be misconfigured. A complete, versioned inventory is a prerequisite for destructive reconciliation. NAUTOBOT-MODELNAUTOBOT-API

Explain telemetry layers: Telegraf can collect/process/send metrics; InfluxDB stores time-series data; Grafana queries data sources and visualizes results. ELK components have their own collection/storage/search/visualization roles. Knowing the dashboard is not the same as operating ingestion, retention, authentication, and recovery. Watch freshness: a green panel showing the last value from yesterday is not current proof of health. Label cardinality matters; avoid one metric label value per request ID or other unbounded identity. TELEGRAFGRAFANAPROM-METRICS

Use service-level indicators that correspond to customer outcomes: successful protected transactions, latency, drops/errors, inspection/bypass state, and availability within a defined scope. Infrastructure metrics explain causes and capacity but do not replace these outcomes. A successful configuration write, healthy BGP session, and running process can coexist with a broken service. SRE-SLO

A change record should join source revision, approved diff, target inventory, executor identity, per-target status, observed before/after state, service verification, and recovery decision. “Rollback” requires a compatible old configuration/application/data state and a recovery path that survives management loss. Stage across failure domains; do not update every redundant member together and call the result highly available. These are proposed controls for your design answers, not a claim about any particular organization’s internal procedures.

Exit drill: a scanner misses half a tenant's devices. Explain how the reconciliation process avoids deleting them, how operators see the incomplete read, and who can authorize any real decommissioning.

12Four architecture whiteboard exercises

These are practice scenarios, not predictions of exact panel questions. Use 12–15 minutes each and reserve the last two minutes for failure handling.

Scenario A — provision 200 private Linux service nodes

Prompt: provision a fleet consistently across multiple facilities without public-cloud tools. Some nodes are bare metal. Rollouts must not interrupt every redundant service member together.

Start by asking for service placement, existing Foreman/Puppet versions, trust/PKI, image/package sources, boot networks, number of sites, capacity margin, out-of-band access, and readiness criteria. Propose an inventory/source-of-truth boundary, approved versioned configuration, site-local provisioning dependencies where needed, authenticated enrollment, Puppet convergence, and traffic admission after verification. Explain what happens when DNS works but DHCP fails, when installation completes but the catalog fails, and when one site loses the central controller.

Strong answer: limited privilege, tested templates, compatible packages, identity lifecycle, clear stage transitions, complete per-host outcomes, and an independent recovery route. The fleet does not enter service just because the installer returned success. A central dependency has a declared availability/recovery strategy. Avoid introducing a dozen unfamiliar systems when existing tools could meet the requirement.

Panel follow-ups: How do rebuilt hosts avoid inheriting stale credentials? How do you stop a bad default from reaching every site? What is the blast radius of a wrong Hiera value? How do you prove a restored controller has consistent database, content, and certificate state? FOREMANPUPPET-ARCHPUPPET-HIERAPOSTGRES-BACKUP

Scenario B — automate a tenant network change

Prompt: create a tenant service spanning a source of truth, two switches, a firewall, and an inspection service. The second switch times out after accepting its write.

Model desired state and validation first: scoped addresses, VLAN/VNI mappings, allowed route export, policy order, and supported platform versions. Produce a deterministic plan and establish a coherent input revision. Explain a per-target execution state machine: planned, authorized, in progress, succeeded, failed, unknown, and recovered. Use local device transactional features where available, but do not claim a global ACID transaction across unrelated devices.

Strong answer: bounded concurrency, a canary gate, a durable change identifier, unknown-result reconciliation, protection against concurrent writers, and service-level validation. A rollback to an old full configuration must not erase unrelated legitimate changes. A failed inventory read must not become an empty inventory. An accepted API request is not operational completion.

Panel follow-ups: What if the orchestrator dies halfway through? What if the same event is delivered twice? What if a device's running configuration changes outside the platform? What if the database commits but the queue is unavailable? HTTP-SEMNETCONFNAUTOBOT-APIRABBIT-ACKPOSTGRES-ISO

Scenario C — TLS application outage after security-path change

Prompt: new flows from most clients work; one application fails; existing sessions also reset during failover. Routing adjacencies remain healthy.

Separate the symptoms rather than forcing one cause. For the failing application, identify transport, TLS version, ALPN/SNI, mutual authentication or pinning, proxy behavior, and upstream certificate validation. For failover resets, examine state synchronization, NAT identity, return path, and connection draining. Select exact timestamps and both original/translated tuples before collecting data.

Strong answer: a packet/path model, independent verification on both TLS legs, scoped captures, and explicit inspection-exception ownership. Health checks must test the relevant service, not merely the management address. “Just bypass the security stack” is not a complete decision; specify authorized traffic class, risk, observability, and expiration.

Panel follow-ups: Why does having the certificate private key not guarantee passive TLS 1.3 decryption? What happens to mTLS client identity at the proxy? Why might an HTTP/3 client follow a different handling path? RFC-TLSRFC-QUICF5-SSLOLINUX-CONNTRACK

Scenario D — intermittent loss on a high-throughput Linux node

Prompt: average utilization is modest, one core is hot, drops increase during small-packet traffic, and a buffer increase improves throughput but worsens tail latency.

Define the load profile and compare per-queue counters, IRQ/softirq CPU, NUMA placement, ring/backlog behavior, flow distribution, and application work. Establish whether drops occur before capture, in host processing, in policy, or downstream. Separate a single flow from aggregate flows. Make one controlled change under repeatable load, with correctness and latency gates.

Strong answer: an evidence table with each hypothesis, distinguishing measurement, low-risk experiment, and acceptance criterion. Do not assert RSS, RPS, offloads, huge buffers, or faster BFD is a universal fix. Keep security visibility and stateful correctness in the benchmark.

Panel follow-ups: Why does idle aggregate CPU not disprove a CPU bottleneck? What does an offloaded packet look like in a host capture? How can a queue reduce loss while making users less happy? LINUX-SCALINGWIRESHARK-OFFLOADSRE-SLO

Your readiness gate

Can you implement and test a safe planner, explain a real configuration workflow, draw a provisioning path, diagnose a Linux resource problem, walk a failed request in both directions, defend routing policy, explain stateful failover, and draw both TLS legs of a proxy? Can you name the evidence that supports your correction and the recovery path if it fails?

Use your diagnostic map and lab evidence to answer these questions. A high multiple-choice score is useful feedback, but the goal is reasoning you can explain without prompts.

Take a diagnostic