Practice Exams:

Azure Data Stores: Match the Platform to the Workload

 

Data architecture becomes confusing when teams begin with product names instead of workload behavior. “Should we use Azure SQL or Cosmos DB?” sounds like a technology question, but the useful answer depends on the data model, transaction boundaries, access patterns, latency, scale, consistency, durability, analytics needs, retention, security, and cost. Azure Storage belongs in the same conversation for object and file workloads, yet it solves a very different problem from a transactional database. The architect’s first job is to classify what the application needs the data platform to do.

The current AZ-305 exam explicitly requires candidates to recommend relational, semi-structured, and unstructured storage solutions, along with service tiers, scalability, protection, durability, integration, and analysis. That makes data-platform selection a core skill for the Azure Solutions Architect Expert certification. The important habit is to map requirements to a data model and operating pattern before comparing service features.

Start with the shape of the data and the way the application reads and writes it

A relational system organizes data around tables, keys, constraints, and transactions that preserve consistency across related records. A document or key-value workload may need flexible schemas and fast access by a partition key. Object storage is designed for blobs and files rather than row-level transactions. These models are not cosmetic differences; they determine how the application expresses relationships, enforces integrity, queries information, and scales.

Access patterns matter just as much as structure. How large are individual records or objects? Are queries known in advance? Does the application read by primary key, filter across many attributes, join related entities, scan large files, or append high-volume events? How frequently are records updated? Those questions narrow the field much more effectively than asking which Azure database is most popular. A good data store is one whose native model matches the work the application performs most often.

Azure SQL is a strong fit when relational integrity and transactional behavior are central

Azure SQL services are natural candidates for applications built around relational schemas, SQL querying, constraints, joins, and transactional consistency. Business systems that update connected records and rely on well-defined relational semantics often benefit from keeping those semantics in the database rather than recreating them in application code. The architecture then focuses on selecting the appropriate Azure SQL deployment option, service tier, compute model, availability configuration, backup strategy, and scaling approach.

The broader operational model around Azure SQL matters because “use SQL” is not the end of the design. Teams must consider workload intensity, connection behavior, storage growth, maintenance, high availability, read scaling, security, and operational responsibilities. Relational fit is the first decision; sizing and operating the selected service is the next one.

Cosmos DB fits distributed NoSQL workloads that can be designed around partitioning

Azure Cosmos DB is designed for globally distributed, highly scalable NoSQL scenarios and supports multiple APIs and consistency choices. It can be a strong fit for applications that need low-latency access at scale, flexible document-oriented data, global distribution, or traffic patterns that do not map cleanly to a traditional relational model. Those benefits depend heavily on data modeling and partition-key design because partitioning determines distribution, throughput behavior, and query efficiency.

A cloud-native Cosmos DB architecture should therefore begin with access patterns rather than with the promise of global scale. If a frequently used query fans across many partitions, or if one partition key creates a hot concentration of traffic, the service cannot compensate for a poor model automatically. Consistency requirements also matter: stronger consistency can change latency and availability tradeoffs. The architect needs to decide which guarantees the application truly requires.

Azure Storage is the right foundation for objects and files, not a cheaper imitation of a database

Blob Storage and related Azure Storage services are optimized for durable object and file data such as documents, images, backups, logs, media, archives, and large analytical files. They provide enormous scale and lifecycle options without the query and transaction model of a relational database. That makes them ideal for workloads where the application retrieves objects by name or metadata, processes files asynchronously, or keeps raw data for downstream analytics.

The concept behind a data lake illustrates the distinction. Large volumes of raw or semi-structured data can be stored economically for later processing, but a lake is not a replacement for an operational transaction store. If an application needs row-level constraints, multi-record transactions, or complex low-latency queries, those requirements point toward a database. Storage should be selected because object semantics fit, not merely because the price per gigabyte looks attractive.

Polyglot persistence is useful when boundaries are clear and expensive when they are not

Many real systems need more than one data model. A commerce platform might keep orders in a relational database, product-session state in a document store, images in Blob Storage, and analytical exports in a data lake. Using several specialized stores can reduce compromises, but every additional service creates another security model, backup process, monitoring surface, data-movement path, and consistency boundary. The benefit must exceed that operational cost.

The safest approach is to give each store a clear responsibility and define how data moves between them. Avoid keeping competing writable copies of the same business fact unless the consistency model is explicit. Derived analytical copies are different from operational sources of truth. When teams cannot say which store owns a field or how updates propagate, “best tool for the job” has turned into distributed ambiguity.

Transaction and consistency requirements define the boundaries of safe change

Applications often discover too late that a data-platform decision changed their consistency model. A relational transaction can update several related rows atomically within its supported boundary. A distributed document design may encourage aggregate boundaries and application-level coordination across partitions or services. Object storage may require entirely different patterns for versioning and concurrent writes. These differences should be part of the application design, not hidden behind a generic repository layer.

Ask what must change together, what users are allowed to observe during an update, and what happens when only part of a multi-step operation succeeds. If a business operation requires strong atomic behavior across several entities, the data model should make that requirement easy to preserve. If eventual consistency is acceptable, the architecture can exploit looser coupling and distribution. Consistency is a business behavior expressed through technology, not merely a database setting.

Scale and latency should be designed from access patterns and geography

Scale is multidimensional. Data volume can grow while request rate remains modest, or request rate can surge against a relatively small data set. Reads and writes can have different growth curves. Some users may be concentrated in one region while others are globally distributed. The correct data platform and configuration depend on which dimension drives the requirement. Architects should model throughput, storage growth, working-set size, latency targets, and regional access rather than relying on a generic “must scale” statement.

That analysis also influences caching, read replicas, geo-replication, partitioning, and data placement. A globally distributed database can reduce user latency, but it introduces decisions about replication, consistency, failover, and cost. A regional relational database with effective caching might be simpler and sufficient. The design should purchase distribution because the workload needs it, not because global capability sounds inherently more architectural.

Protection, durability, security, and governance belong in the data-store decision

Data stores are not interchangeable from a protection perspective. Backup and restore mechanisms, point-in-time recovery, replication, redundancy, encryption, private connectivity, identity integration, key management, retention controls, and auditing differ across services. The architect needs to map recovery point and recovery time objectives to the capabilities of the chosen platform and verify that the design protects against both infrastructure failure and logical mistakes such as accidental deletion or corruption.

Governance also matters. Data classification and residency can restrict regions or replication patterns. Sensitive information may require private network access and controlled administrator paths. Retention rules can influence storage tiering and lifecycle management. The cheapest or fastest store is not appropriate if its configuration cannot satisfy the organization’s control requirements. Data architecture should make security and compliance properties explicit alongside performance and scale.

Lifecycle policy should be designed at the same time. Hot operational data, historical records, legal-retention copies, backups, and analytical exports can have different performance and cost needs. Tiering, archival, deletion, and restore expectations should follow the value and obligations of the data rather than allowing every byte to remain indefinitely in the most expensive operational tier.

Operational databases are optimized for application behavior, while analytical systems are optimized for aggregation, exploration, and historical processing. Trying to make one store perform both roles can create contention, expensive queries, or awkward schemas. The architecture should decide whether analytics runs against replicas, exports, streams, a lake, a warehouse, Fabric, or another analytical platform, and how fresh the analytical copy needs to be.

Integration design should also account for change capture, batch movement, event streams, schema evolution, and failure recovery. Moving data is not free: it introduces latency, cost, governance boundaries, and another place where records can become inconsistent. A data platform decision is stronger when it includes the expected downstream consumers rather than optimizing only the first application that writes the data.

Record the workload reasoning so the platform can evolve without product loyalty

A durable architecture decision should state why a service was selected: relational transactions, document access, global distribution, object semantics, cost profile, operational model, or another requirement. It should also record the assumptions that would cause the choice to be revisited. If query patterns change, data volume grows, regulatory scope expands, or a once-regional application becomes global, the original decision may no longer be optimal.

This keeps the conversation anchored to workload behavior instead of brand preference. Azure SQL, Cosmos DB, and Azure Storage are all strong services when their native models match the problem. None is a universal data platform. AZ-305-level architecture means understanding the tradeoffs well enough to select the store that best fits the required data model, access pattern, consistency, scale, protection, and operating model—and to change that decision when the workload itself changes.

Related Posts

• Why Network Segmentation Still Stops Real Attacks

• Least Privilege as an Architecture Principle

• Availability Sets, Zones, and Scale Sets Solve Different Problems

• Entra Groups, Roles, and Access Reviews in Everyday Administration

• Spanning Tree Still Matters in a World of Faster Switches

• Network Automation Starts With Structured Data, Not Python

• Agents Need Boundaries More Than They Need More Tools

• Data Governance for RAG Pipelines That Touch Sensitive Information

• Campus Fabric Changes Segmentation

• SD-WAN Policy Turns Intent Into Path Selection