Skip to content

Design and Prepare Machine Learning Solutions:
Sample Questions and Answers

Note

The questions and answers provided in this study guide are for practice purposes only and are not official practice questions. They are intended to help you prepare for the DP-100 Microsoft certification exam. For additional preparation materials and the most up-to-date information, please refer to the official Microsoft documentation.

Topics Covered

  • Design machine learning solutions
  • Set up Azure Machine Learning workspace
  • Configure compute resources
  • Manage workspace security and access
List of References (Click to expand)
List of questions/answers (Click to expand)

Tip

Azure Machine Learning Workspace Architecture:

When a workspace is provisioned, Azure automatically creates supporting resources:

Component Purpose
Azure Storage Account Store files, notebooks, job metadata, and models
Azure Key Vault Securely manage authentication keys and credentials
Application Insights Monitor predictive services and track performance
Azure Container Registry Store Docker images for ML environments (created when needed)

Tip

Workspace Creation Methods:

Method Use Case Complexity
Azure Portal Quick setup, visual interface Low
ARM Template Infrastructure as Code, repeatable deployments Medium
Azure CLI Automation, scripting, CI/CD pipelines Medium
Python SDK Programmatic control, integration with ML workflows High

Tip

Azure ML Security Roles:

Role Permissions Best For
Owner Full access to all resources, can grant access to others Workspace administrators
Contributor Full access to all resources, cannot grant access Senior data scientists
Reader View-only access, cannot make changes Stakeholders, auditors
AzureML Data Scientist All actions except compute management and workspace settings Standard data scientists
AzureML Compute Operator Create, change, and manage compute resources MLOps engineers

Tip

Compute Types Comparison:

Compute Type Use Case Scaling Cost Model
Compute Instance Interactive development, Jupyter notebooks Fixed size Pay per hour while running
Compute Cluster Training jobs, batch inference Auto-scale (0-N nodes) Pay per node per hour
Kubernetes Cluster Production deployments, custom container orchestration Manual configuration Bring your own cluster
Attached Compute Use existing resources (VMs, Databricks) External management External billing
Serverless Compute On-demand training jobs Fully managed Pay per job execution

Tip

Workspace Assets:

Asset Type Description Example Use Case
Models Trained ML models with metadata Store trained models for deployment
Environments Runtime configurations (packages, Docker images) Ensure reproducible training and inference
Data Datasets and data references Manage training and validation datasets
Components Reusable pipeline steps Build modular ML pipelines

Tip

Compute Cluster Configuration:

Parameter Description Considerations
size VM type (CPU/GPU, memory, storage) Match compute needs to workload
max_instances Maximum number of nodes for scaling Balance performance vs cost
tier Priority level (Dedicated vs Low Priority) Low Priority = 80% cost savings but no availability guarantee

Q1: Azure Machine Learning Workspace Components

When you create an Azure Machine Learning workspace, which additional Azure resources are automatically created?

  • Azure Storage Account only ❌: Incorrect. Multiple resources are created automatically.
  • Azure Key Vault and Application Insights only ❌: Incorrect. Additional resources are also created.
  • Azure Storage Account, Azure Key Vault, Application Insights, and Azure Container Registry ✅: Correct. These four resources are automatically provisioned to support the workspace.
  • Only the workspace resource itself ❌: Incorrect. Supporting resources are automatically created.

Q2: Workspace Creation Methods

Which of the following are valid methods to create an Azure Machine Learning workspace? (Select all that apply)

  • Azure Portal UI ✅: Correct. You can create a workspace through the Azure portal interface.
  • Azure Resource Manager (ARM) template ✅: Correct. ARM templates enable Infrastructure as Code for workspace creation.
  • Azure CLI with ML extension ✅: Correct. The Azure CLI with the ML extension supports workspace creation.
  • Python SDK ✅: Correct. The Azure ML Python SDK can programmatically create workspaces.

Q3: Built-in Security Roles

You need to give a team member the ability to run experiments and manage datasets, but they should not be able to create or delete compute resources. Which role should you assign?

  • Owner ❌: Incorrect. Owner has full access including compute management.
  • Contributor ❌: Incorrect. Contributor can manage compute resources.
  • AzureML Data Scientist ✅: Correct. This role allows all workspace actions except compute management and workspace settings.
  • AzureML Compute Operator ❌: Incorrect. This role is specifically for compute management.

Q4: Compute Instance vs Compute Cluster

When should you use a Compute Instance versus a Compute Cluster?

  • Use Compute Instance for distributed training, Compute Cluster for development ❌: Incorrect. This is backwards.
  • Use Compute Instance for interactive development, Compute Cluster for scalable training jobs ✅: Correct. Compute Instances are ideal for notebooks and experimentation, while Compute Clusters are for training at scale.
  • They are interchangeable and serve the same purpose ❌: Incorrect. They have different use cases.
  • Use Compute Cluster only for inference, Compute Instance only for training ❌: Incorrect. Both can be used for training.

Q5: Customer Managed Keys

When implementing customer-managed keys (CMK) for workspace encryption, what information must you provide?

  • Only the Key Vault name ❌: Incorrect. More information is required.
  • Key Vault resource ID and key URI ✅: Correct. Both the full Key Vault resource ID and the specific key URI are required.
  • Only the key URI ❌: Incorrect. The Key Vault resource ID is also needed.
  • Key Vault name and subscription ID ❌: Incorrect. The full resource ID and key URI are needed.

Q6: Workspace Assets

Which of the following are considered assets in an Azure Machine Learning workspace? (Select all that apply)

  • Models ✅: Correct. Trained models are registered as assets in the workspace.
  • Environments ✅: Correct. Runtime environments are managed as assets.
  • Data ✅: Correct. Datasets and data references are workspace assets.
  • Components ✅: Correct. Reusable pipeline components are assets.

Q7: Azure CLI vs Python SDK

What are the main advantages of using Azure CLI with Azure Machine Learning? (Select all that apply)

  • Automate creation and configuration for repeatability ✅: Correct. CLI enables automation and scripting.
  • Ensure consistency across multiple environments ✅: Correct. Scripts can be reused across dev, test, and production.
  • Integrate with DevOps CI/CD pipelines ✅: Correct. CLI commands work well in automated pipelines.
  • Provide interactive debugging capabilities ❌: Incorrect. Interactive debugging is better suited for SDKs or IDEs.

Q8: Compute Configuration Parameters

When creating a compute cluster, which parameter controls the cost versus availability trade-off?

  • size ❌: Incorrect. Size affects performance and cost but not availability guarantees.
  • max_instances ❌: Incorrect. This controls scaling but not availability guarantees.
  • tier ✅: Correct. Setting tier to 'LowPriority' reduces costs by ~80% but provides no availability guarantee.
  • location ❌: Incorrect. Location affects latency and compliance but not the cost/availability trade-off.

Q9: Environment Management

What is the recommended approach for managing environments across multiple experiments?

  • Create a new environment for each experiment ❌: Incorrect. This creates unnecessary redundancy.
  • Create and register an environment, then reuse it across experiments ✅: Correct. Registered environments promote reproducibility and reduce duplication.
  • Always use the default environment ❌: Incorrect. Default environments may not have required packages.
  • Share Dockerfiles manually between team members ❌: Incorrect. This doesn't leverage Azure ML's environment management capabilities.

Q10: Authentication Methods

What are the three required parameters for authenticating to an Azure Machine Learning workspace using the Python SDK?

  • workspace_name, resource_group, location ❌: Incorrect. Location is not required for authentication.
  • subscription_id, resource_group, workspace_name ✅: Correct. These three parameters are required to connect to a workspace.
  • subscription_id, workspace_name, tenant_id ❌: Incorrect. Resource group is required instead of tenant_id.
  • resource_group, workspace_name, region ❌: Incorrect. Subscription_id is required instead of region.

Code Examples

Creating a Workspace with Python SDK

from azure.ai.ml import MLClient
from azure.ai.ml.entities import Workspace, CustomerManagedKey
from azure.identity import DefaultAzureCredential

# Create workspace with customer-managed key
ws = Workspace(
    name="my-workspace",
    location="eastus",
    display_name="My ML Workspace",
    description="Workspace for DP-100 preparation",
    customer_managed_key=CustomerManagedKey(
        key_vault="/subscriptions/<sub-id>/resourcegroups/<rg>/providers/microsoft.keyvault/vaults/<vault>",
        key_uri="<key-identifier-uri>"
    ),
    tags={"purpose": "DP-100", "environment": "development"}
)

# Create the workspace
ml_client = MLClient(
    credential=DefaultAzureCredential(),
    subscription_id="<subscription-id>",
    resource_group_name="<resource-group>",
    workspace_name="<workspace-name>"
)

ml_client.workspaces.begin_create(ws)

Creating a Compute Cluster

from azure.ai.ml.entities import AmlCompute

compute_cluster = AmlCompute(
    name="cpu-cluster",
    size="Standard_DS3_v2",
    min_instances=0,
    max_instances=4,
    tier="Dedicated",  # or "LowPriority" for cost savings
    idle_time_before_scale_down=120  # seconds
)

ml_client.compute.begin_create_or_update(compute_cluster)

Best Practices

  1. Security:
  2. Use customer-managed keys for sensitive workloads
  3. Implement least-privilege access with appropriate roles
  4. Enable private endpoints for network isolation

  5. Cost Optimization:

  6. Use Low Priority compute for non-critical workloads
  7. Set up auto-shutdown schedules for compute instances
  8. Configure appropriate idle time for compute clusters

  9. Organization:

  10. Use consistent naming conventions
  11. Apply tags for cost tracking and governance
  12. Create separate workspaces for different environments (dev/test/prod)

  13. Automation:

  14. Use ARM templates or CLI scripts for reproducible deployments
  15. Integrate workspace creation into CI/CD pipelines
  16. Version control your infrastructure code