Hi there,
mpi_operator_job_info looks like it can get stale today. It is set when the launcher exists:
mpiJobInfoGauge.WithLabelValues(launcher.Name, mpiJob.Namespace).Set(1)
but I don’t see a matching cleanup path using DeleteLabelValues, Delete, or Reset. The MPIJob informer also does not appear to handle deletes for this metric, and syncHandler returns on NotFound without removing the series. That means a metric like this can stay exported until the operator restarts, even after the MPIJob has been deleted:
mpi_operator_job_info{namespace="...",launcher="...-launcher"} 1
I think we should consider deprecating/removing this metric. It is per-object, grows with job churn, and only exposes the launcher name and namespace. The same relationship can usually be derived from Kubernetes labels/owner refs or kube-state-metrics.
Also, I’d be interested in improving mpi-operator’s Prometheus coverage, but more around controller behavior and job lifecycle rather than per-object info gauges. Examples: workqueue depth/latency, reconcile count and duration, reconcile error count by reason, MPIJob pod/resource creation latency, job start latency, and job duration by result. I’d be happy to help contribute some of these metrics if maintainers think they fit the project.
Hi there,
mpi_operator_job_infolooks like it can get stale today. It is set when the launcher exists:but I don’t see a matching cleanup path using
DeleteLabelValues,Delete, orReset. The MPIJob informer also does not appear to handle deletes for this metric, and syncHandler returns on NotFound without removing the series. That means a metric like this can stay exported until the operator restarts, even after the MPIJob has been deleted:I think we should consider deprecating/removing this metric. It is per-object, grows with job churn, and only exposes the launcher name and namespace. The same relationship can usually be derived from Kubernetes labels/owner refs or kube-state-metrics.
Also, I’d be interested in improving mpi-operator’s Prometheus coverage, but more around controller behavior and job lifecycle rather than per-object info gauges. Examples: workqueue depth/latency, reconcile count and duration, reconcile error count by reason, MPIJob pod/resource creation latency, job start latency, and job duration by result. I’d be happy to help contribute some of these metrics if maintainers think they fit the project.