Google Professional Machine Learning Engineer: Model Monitoring
Model monitoring exists because a model can stay available while becoming less useful. Inputs change, customer behavior shifts, upstream pipelines evolve, labels arrive late, and serving code changes. Traditional infrastructure metrics are necessary but cannot show whether the statistical relationship the model learned still holds.
In Google Cloud ML, current monitoring capabilities include skew and drift analysis, endpoint monitoring for supported model types, and newer Model Monitoring v2 workflows for comparing target and baseline datasets. Some v2 capabilities remain pre-GA, so implementation should verify current launch status and supported data sources.
The objective is not to alert on every statistical movement. It is to connect meaningful model or data change to business impact and a defined response.
Separate service health from model health
Endpoint latency, error rate, saturation, and availability measure the serving system. Drift, skew, calibration, prediction distribution, and business outcomes measure the model system. Both sets matter, but they answer different questions.
Use a meaningful baseline
Drift is always relative to something. A baseline can be the training dataset, an earlier serving window, a validated reference period, or another dataset that represents expected behavior.
Monitor the features that can change decisions
Not every feature deserves the same alert threshold. High-impact features, protected attributes, key categorical distributions, and features with brittle upstream dependencies may require closer review than low-value auxiliary fields.
Treat thresholds as operational policy
Thresholds should reflect tolerance and actionability, not arbitrary defaults. Too-sensitive thresholds create alert fatigue; too-loose thresholds detect change only after users notice the impact.
Watch feature freshness and pipeline status
A stale feature table can freeze distributions and reduce apparent drift while predictions become wrong. Monitoring should therefore include the age of feature data, pipeline success, source volume, and sync lag.