Practice Exams:

Google Professional Machine Learning Engineer: Feature Management

Features are the data attributes a model uses to make a prediction. Feature management becomes an engineering problem when the same definitions must be reused across training, batch inference, and online serving without letting teams silently calculate different values for the same concept.

For Google Cloud ML, the platform surface is currently evolving. Google has deprecated Vertex AI Feature Store Legacy/V1 and optimized online serving, while newer feature-management documentation emphasizes BigQuery-backed feature data, online serving, registry-style reuse, and metadata integration. New designs should verify the current API and migration path before implementation.

The durable architectural principles remain point-in-time correctness, ownership, reuse, freshness, lineage, access control, and consistency between training and serving.

Define a feature as a data product

A reusable feature needs more than a column name. It needs a semantic definition, entity key, data type, freshness expectation, owner, valid range, and a clear explanation of how it is calculated. Without those details, a shared feature store can become a shared ambiguity store.

Names should communicate business meaning rather than pipeline implementation. A feature called customer_30d_spend is easier to govern than agg_17, and its definition can state whether refunds, tax, and currency normalization are included.

Data governance makes the same point: metadata should help people decide whether an asset is appropriate, not merely help them find it.

Keep offline and online paths aligned

Training often reads large historical datasets, while online inference needs the most recent feature values at low latency. The storage and serving paths can therefore differ even though the feature definition must stay the same.

A good architecture treats BigQuery or another governed analytical source as the authoritative history and uses an online serving layer only for the subset that needs low-latency access. Synchronization lag, failed syncs, and serving freshness become observable operational states.

Feature stores are valuable because they coordinate definitions and reuse before they improve raw serving speed.

Protect point-in-time correctness

Training data must not accidentally use information that was only known after the prediction time. If a feature pipeline joins each training example to the latest value instead of the value that existed at that historical moment, the model can learn from the future.

Point-in-time joins need reliable event timestamps and versioned history. This is especially important for account status, balances, rolling aggregates, and any feature updated frequently.

Data quality should include temporal correctness because a perfectly valid current value can still be wrong for a historical training example.

Use entities that match prediction context

The entity key determines how features are retrieved. Customer, device, merchant, session, product, and account can all be valid entities, but each creates different cardinality, freshness, and access patterns.

Avoid forcing unrelated features under one oversized entity because it looks convenient in a demo. Entity boundaries should reflect how predictions are made and who owns the underlying data.

BigQuery design can help keep offline feature tables efficient when time and entity keys shape common access patterns.

Version definitions, not just code

A feature can change semantics without changing its data type. If the definition of active_customer moves from 30 days to 60 days, old models may no longer be compatible even though the column remains boolean.

Version important feature definitions or preserve metadata that identifies the producing logic. A model registry entry should make it possible to recover which feature definitions were used during training.

Vertex pipelines can carry versioned artifacts and parameters so the feature computation is tied to the model lifecycle rather than reconstructed later from memory.

Monitor freshness separately from availability

An online feature service can respond successfully with stale data. Availability monitoring therefore needs a companion freshness signal that compares source event time, sync time, and serving time.

Different features need different service levels. A customer lifetime value might tolerate daily refresh, while fraud or inventory features may become misleading within minutes.

Model monitoring should be interpreted alongside feature freshness because drift can be caused by a stale pipeline rather than a genuine change in user behavior.

Control access to sensitive features

Central reuse can increase exposure if sensitive attributes are made broadly available. Apply IAM, dataset controls, row or column protections, and audit logging according to the classification of the feature and the decisions it supports.

Some attributes should not be used for modeling even when technically available. Responsible-AI review should examine proxies, protected characteristics, and whether a feature creates an unjustifiable or opaque decision boundary.

Responsible AI is strongest when feature choice is reviewed before training, not only after a model produces problematic outcomes.

Treat online serving as a dependency

If an application cannot generate a prediction without an online feature lookup, the feature service is on the critical request path. Latency, quotas, regional design, retries, and failure behavior need the same engineering attention as the model endpoint.

Applications should know whether to fail closed, use a safe default, fall back to cached values, or defer the decision when a feature is unavailable. The answer depends on the business risk of a stale or missing feature.

Vertex deployment should therefore be designed with the feature-serving dependency rather than as an isolated model endpoint.

Reconcile feature usage with lineage

A model incident becomes easier to investigate when the team can trace a surprising prediction back to the model version, feature definitions, source tables, pipeline run, and training snapshot that produced it.

Catalog and lineage integration can reduce the manual work of reconstructing those relationships, but lineage is only useful when resource names and ownership are consistent enough for humans to interpret.

Professional ML Engineer responsibilities now span data, models, pipelines, monitoring, and governance, which makes feature lineage part of operational ML rather than a documentation side task.

Design transformation ownership

Features are often derived through SQL, Beam, Spark, Python, or model-assisted preprocessing. The team must decide where that transformation code lives, how it is reviewed, and how the same logic is reused between historical backfills and incremental updates. Duplicating feature logic in a training notebook and an online service is a common path to training-serving inconsistency.

Treat transformation code as versioned production software. Unit tests should cover boundary conditions, missing values, categorical mappings, and time windows. When a definition changes, the model team should know whether historical training data must be recomputed or whether the new version starts only from a cutover date.

Backfill features without corrupting history

New features often require historical computation so models can train on enough examples. Backfills should use the transformation version and source snapshots appropriate to each historical period rather than applying current state blindly. Otherwise the backfill can create information that would not have existed at the time of the prediction.

Large backfills also need capacity and cost planning. Run them in bounded partitions, validate row counts and timestamps, and keep production online-serving updates from being starved by a one-time historical job.

Detect feature leakage before training

Leakage occurs when a feature contains information that would not be available when a real prediction is made, or when a proxy reveals the label too directly. Leakage can produce impressive offline metrics and disappointing production behavior. Review features from the perspective of the prediction timestamp, not from the perspective of what is easy to join in the warehouse.

Use feature documentation to record availability time, source delay, and any post-event updates. Those fields give reviewers a practical way to challenge whether the value would truly exist at inference time.

Retire features deliberately

A shared feature can outlive the model that created it. Track consumers before changing or deleting a feature, and define a deprecation window so teams can migrate. Removing a field from the online view without understanding dependencies can cause production failures far from the owning team.

Periodic cleanup is still important. A registry filled with abandoned or duplicate features makes discovery harder and increases the chance that a new model uses an outdated definition. Deprecation status should be visible alongside ownership and freshness metadata.

Feature-management architecture should also define what happens when the source system corrects historical data. Some features can be recomputed safely; others may have trained models whose historical values must remain reproducible. Decide whether corrections rewrite history, create a new feature version, or apply only prospectively, and record that policy with the feature definition.

For real-time systems, measure the feature lookup separately from the model endpoint. A 20 millisecond model with a 200 millisecond feature dependency is a 220 millisecond prediction path before application overhead. Latency budgets should allocate time explicitly to feature serving, networking, inference, and downstream decision logic.

Governance reviews should periodically identify features that encode the same concept in different ways. Consolidating duplicates reduces training-serving inconsistency and makes monitoring easier, but consolidation should preserve the semantics that existing models depend on rather than forcing a silent definition change.

The feature platform should also expose usage so owners can see which models depend on a definition before modifying it. Consumer visibility turns change management from guesswork into an explicit migration process. If the platform cannot provide that dependency graph automatically, maintain a registry field or deployment manifest that ties model versions to the feature views they require.

Related Posts

• Data & AI on Google Cloud

• Google Professional Data Engineer: BigQuery Cost Control

• Google Professional Data Engineer: BigQuery Partitioning and Clustering

• Google Professional Data Engineer: Data Governance with Dataplex

• Google Professional Data Engineer: Dataflow or Dataproc?

• Google Professional Data Engineer: Pub/Sub for Streaming Data Pipelines

• Google Professional Data Engineer: Data Quality in Google Cloud Pipelines

• Generative AI on Google Cloud

• Microsoft DP-600: Delta Tables in Microsoft Fabric

• Microsoft PL-300: DAX Context Without the Confusion