Diagnose GKE Cloud TPU multi-slice training hangs (Megascale `HANG_DETECTED` logs and hang monitored events) using the ML Diagnostics `Megascale XLA (MXLA) Hang Analyzer` (`gcloud alpha mldiagnostics monitored-events`) and 1-minute Cloud Monitoring multi-slice latency metrics (`kubernetes.io/container/multislice/*`). Distinguishes XLA compiler/HLO launch divergence and host data-input stalls from TPU chip, SparseCore, ICI, or network fabric faults. Use when multi-slice TPU training jobs freeze without progressing steps, emit `HANG_DETECTED`, or stall in collective operations. Don't use for gradual step-time throughput drops without hangs (use gke-ai-troubleshooting-tpu-performance-degradation) or pod preemption/eviction restarts (use gke-ai-troubleshooting-jobset-interruption).
Permissions
Files
Diagnose GKE Cloud TPU multi-slice training hangs (Megascale `HANG_DETECTED` logs and hang monitored events) using the ML Diagnostics `Megascale XLA (MXLA) Hang Analyzer` (`gcloud alpha mldiagnostics monitored-events`) and 1-minute Cloud Monitoring multi-slice latency metrics (`kubernetes.io/container/multislice/*`). Distinguishes XLA compiler/HLO launch divergence and host data-input stalls from TPU chip, SparseCore, ICI, or network fabric faults. Use when multi-slice TPU training jobs freeze without progressing steps, emit `HANG_DETECTED`, or stall in collective operations. Don't use for gradual step-time throughput drops without hangs (use gke-ai-troubleshooting-tpu-performance-degradation) or pod preemption/eviction restarts (use gke-ai-troubleshooting-jobset-interruption).
Version history
Diagnose Cloud TPU multi-slice training hangs on Google Kubernetes Engine (GKE)
by correlating Megascale HANG_DETECTED logs and ML Diagnostics Workload
Monitoring Megascale XLA (MXLA) Hang Analyzer reports with 1-minute
multi-slice latency metrics (kubernetes.io/container/multislice/*) and GKE
node topology labels.
gcloud with
alpha component for gcloud alpha mldiagnostics) and kubectl.gcloud billing projects describe {project_id}), authenticate
(gcloud auth login), set the target project (gcloud config set project {project_id}), and ensure container.googleapis.com,
logging.googleapis.com, monitoring.googleapis.com, and
hypercomputecluster.googleapis.com are enabled.jobset and job GKE
job types, and is compatible with GKE versions 1.36.0-gke.4681000 and later
(Configure GKE for ML Diagnostics,
which also covers the cluster setup needed for on-demand profiling in Step 4
Path B). If a workload uses another framework (such as PyTorch) or another
custom resource type, gcloud alpha mldiagnostics won't list ML runs or
monitored events for it. The Megascale XLA hang analyzer and Megascale XLA
metrics require LibTPU 0.40.0 or later (see "Get started" in
Workload monitoring with ML Diagnostics).roles/hypercomputecluster.editor), the role
listed in the "IAM permissions" section of
ML Diagnostics platform,
for the ML Diagnostics CLI and API calls in this skill (ML runs, monitored
events, and on-demand profiler sessions)roles/monitoring.viewer) for the PromQL queries in Step
2roles/logging.viewer) for the Cloud Logging query in Step 1roles/container.viewer) for the kubectl get nodes query in Step 3[High Risk] steps): Kubernetes Engine Cluster Admin
(roles/container.clusterAdmin)Read-only rule: Run read-only diagnostic commands only. Never drain, delete, or re-create nodes, or run any other command that changes the cluster. Give the user any fix to apply themselves.
When you recommend a fix, link the doc section that describes it.
A Megascale hang occurs when a multi-slice worker has waited on a Megascale
communication operation for a set timeout period. The TPU logs then show a
Megascale HANG_DETECTED message. HANG_DETECTED is a catch-all signal that
the workload isn't progressing, and the cause can be in software or in hardware.
When a hang occurs, ML Diagnostics runs the Megascale XLA (MXLA) Hang Analyzer, which reports the likely cause as a code. Consult the
Megascale XLA (MXLA) Hang Analyzer
section for the definition and recommended action of each code, and route by
category:
FINGERPRINT_MISMATCH), route to
Path A: Compiler or HLO divergence.DATA_INPUT_STALL), route to
Path B: Host program queueing or data input stall.NOT_DETECTED
or reports UNKNOWN:
HANG_DETECTED logs or a hang monitored event fired, but the analyzer
report has detectionState: "NOT_DETECTED" or reports UNKNOWN ("The MXLA
hang was detected but a potential cause is not determined"), route to
Path D: No DETECTED analyzer or UNKNOWN cause.[Low Risk]Collect the target parameters. By default, query a 60-minute window `[T - 30m, T
around{issue_time}`:{project_id}: Google Cloud project ID{location}: Google Cloud region where the ML run and GKE cluster reside (for
example, us-central1){cluster_name}: GKE cluster name{namespace} / {workload_name}: Kubernetes namespace and JobSet/Pod prefix{ml_run_id}: ML Diagnostics run ID{issue_time}: Timestamp when the hang occurred (T, ISO-8601 UTC){start_time}: T - 30m{end_time}: T + 30mHANG_DETECTED logs and hang events [Low Risk]HANG_DETECTED (read-only Cloud Logging
LQL): Query k8s_container logs over [{start_time}, {end_time}] to
confirm HANG_DETECTED and identify the first stalled pods:resource.type="k8s_container"
resource.labels.project_id="{project_id}"
resource.labels.cluster_name="{cluster_name}"
"HANG_DETECTED"
timestamp >= "{start_time}" AND timestamp <= "{end_time}"
gcloud alpha mldiagnostics machine-learning-run list) to identify {ml_run_id}.gcloud alpha mldiagnostics monitored-events list and gcloud alpha mldiagnostics monitored-events describe) or
Access Workload Monitoring information through the API
to inspect the Megascale XLA (MXLA) Hang Analyzer report (detectionState,
details, and recommendedActions).
HANG_DETECTED logs or a hang monitored event fired, run Step 2 to
correlate with the 1-minute metrics; if no analyzer reports
detectionState: "DETECTED" or the analyzer reports UNKNOWN, follow
Path D: No DETECTED analyzer or UNKNOWN cause.HANG_DETECTED log or hang monitored event exists and multi-slice
latencies are normal in Step 2, rule out an MXLA hang by following
Path E: Healthy telemetry.[Low Risk]Consult the
System Metrics
section of the Workload Monitoring guide for the 1-minute multi-slice network
(kubernetes.io/container/multislice/network/*), multi-slice accelerator
(kubernetes.io/container/multislice/accelerator/*), and node duty-cycle
(kubernetes.io/node/accelerator/duty_cycle) metrics, and run read-only PromQL
queries over [{start_time}, {end_time}]:
# 1. P95 multi-slice collective end-to-end latency by pod
histogram_quantile(
0.95,
sum by (pod_name, le) (
rate(kubernetes_io:container_multislice_network_collective_end_to_end_latencies_bucket{
monitored_resource="k8s_container",
project_id="{project_id}",
cluster_name="{cluster_name}"
}[5m])
)
)
# 2. P95 multi-slice DCN transfer latency by pod
histogram_quantile(
0.95,
sum by (pod_name, le) (
rate(kubernetes_io:container_multislice_network_dcn_transfer_latencies_bucket{
monitored_resource="k8s_container",
project_id="{project_id}",
cluster_name="{cluster_name}"
}[5m])
)
)
# 3. P95 host-to-device transfer latency by pod
histogram_quantile(
0.95,
sum by (pod_name, le) (
rate(kubernetes_io:container_multislice_accelerator_host_to_device_transfer_latencies_bucket{
monitored_resource="k8s_container",
project_id="{project_id}",
cluster_name="{cluster_name}"
}[5m])
)
)
# 4. Node TPU duty cycle (Workload Monitoring detects a hang as a prolonged
# period of minimal to no TPU activity)
kubernetes_io:node_accelerator_duty_cycle{
monitored_resource="k8s_node",
project_id="{project_id}",
cluster_name="{cluster_name}"
}
[Low Risk]When the Megascale XLA (MXLA) Hang Analyzer reports culprit numeric Compute
Engine instance IDs in details or recommendedActions, map those numeric IDs
to GKE Node names and physical topology blocks using this read-only kubectl
query inspecting container.googleapis.com/instance_id:
kubectl get nodes -l cloud.google.com/gke-tpu-accelerator \
-o jsonpath='{range .items[*]}{.metadata.name}{"\tinstance_id="}{.metadata.annotations.container\.googleapis\.com/instance_id}{"\tblock="}{.metadata.labels.cloud\.google\.com/gce-topology-block}{"\tsubblock="}{.metadata.labels.cloud\.google\.com/gce-topology-subblock}{"\thost="}{.metadata.labels.cloud\.google\.com/gce-topology-host}{"\n"}{end}'
Load and follow only the reference file that matches the root-cause category from Step 1:
FINGERPRINT_MISMATCH): Read
Path A: Compiler or HLO divergence.PROGRAM_NOT_QUEUED or
DATA_INPUT_STALL): Read
Path B: Host program queueing or data input stall.UNRECOVERABLE_ERROR on specific instances): Read
Path C: Hardware or network faults.HANG_DETECTED logs or hang event fired, but the analyzer is
NOT_DETECTED or reports UNKNOWN): Read
Path D: No DETECTED analyzer or UNKNOWN cause.HANG_DETECTED logs or hang events, and steady telemetry):
Read Path E: Healthy telemetry.gcloud compute instance-groups managed commands, such as
delete, on a node pool's managed instance group. Handle nodes through GKE,
as described in
Path C: Hardware or network faults.In these kits
More from @google
Works with
Claude, Codex, Cursor & more