Skip to content

Infrastructure Troubleshooting

Every component reports healthy and the service still fails.

Faults that cross network, identity, certificate and application boundaries resist single-team diagnosis, because each team can prove its own part is fine. We follow the evidence end to end until the cause is found and fixed.

  • Microsoft Partner
  • AZ-104 and AZ-305 certified engineers
  • Senior engineers only
  • Brisbane-based, working across Australia
Fibre inspection scope, power meter and patch leads on a workbench.

How we investigate

  1. Reproduce and bound the fault: exactly what fails, for whom, from where, how often, and what changed around the first occurrence.
  2. Trace the path: client, DNS, routing, firewall and NAT state, load balancer, TLS handshake, authentication, application, backend. Capture at multiple points to find where the conversation stops matching.
  3. Form hypotheses and eliminate them with evidence, not opinion. Every conclusion is tied to a capture, a log line or a configuration difference.
  4. Prove the cause by reproducing it on demand, then fix it and prove the fix under the same conditions.
Five components in a chain, client, DNS, network, identity and app, each reporting healthy, with the fault marked on the link between DNS and network.
Everything reports healthy. The fault is in the path between them.

The faults that usually hide here

  • Intermittent hybrid connectivity: asymmetric routing, MTU and fragmentation, BGP flaps, firewall session timeouts, SNAT exhaustion
  • Name resolution: split-horizon DNS, forwarder loops, stale private DNS zones, conditional forwarders that differ between sites
  • Identity and certificates: Kerberos delegation and clock skew, expired or mismatched chains, Conditional Access surprises, token lifetime behaviour
  • Performance that nobody can reproduce: storage latency, noisy neighbours, throttling on Azure services, chatty applications across a WAN
  • Backups that run and do not restore, and automation that drifts from what it is supposed to enforce

What you receive

  • A written root cause with the evidence, timeline and diagrams
  • A permanent fix, applied and verified
  • The change captured as code or a runbook so it does not regress
  • Monitoring or a detection rule that would catch it earlier next time

When it is an ongoing problem

If fault-finding of this kind is a regular need, senior escalation under Managed Engineering gives your team direct access to the same engineers, with the environment already known.

Questions buyers ask

How quickly can you start?

For urgent faults, as soon as access is in place. We agree the priority and an initial response window when you call.

What access do you need?

Read-only access and logs to start, then change access through your process once the cause is proven.

What if the fault is in a vendor's product?

We gather the evidence the vendor needs and work the case with them until it is resolved.

Do you fix the problem or just diagnose it?

Both. We prove the cause, apply a permanent fix through change control and verify it.

What do we get at the end?

A written root cause, the permanent fix, the detection that would have caught it earlier and a handover session for your team.

What if it keeps happening?

Recurring faults usually point to a design problem. We can scope the redesign, or cover it under senior escalation or Managed Engineering.

Start a conversation

Find the cause, fix it properly, and make it stay fixed.

Tell us about the fault you cannot pin down.

Talk to an Engineer