Skip to content

Remove or deprecate stale-prone mpi_operator_job_info metric #808

Description

@lordofire

Hi there,

mpi_operator_job_info looks like it can get stale today. It is set when the launcher exists:

mpiJobInfoGauge.WithLabelValues(launcher.Name, mpiJob.Namespace).Set(1)

but I don’t see a matching cleanup path using DeleteLabelValues, Delete, or Reset. The MPIJob informer also does not appear to handle deletes for this metric, and syncHandler returns on NotFound without removing the series. That means a metric like this can stay exported until the operator restarts, even after the MPIJob has been deleted:

mpi_operator_job_info{namespace="...",launcher="...-launcher"} 1

I think we should consider deprecating/removing this metric. It is per-object, grows with job churn, and only exposes the launcher name and namespace. The same relationship can usually be derived from Kubernetes labels/owner refs or kube-state-metrics.

Also, I’d be interested in improving mpi-operator’s Prometheus coverage, but more around controller behavior and job lifecycle rather than per-object info gauges. Examples: workqueue depth/latency, reconcile count and duration, reconcile error count by reason, MPIJob pod/resource creation latency, job start latency, and job duration by result. I’d be happy to help contribute some of these metrics if maintainers think they fit the project.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions