Claude Skill

dotnet-mcaf-ml-ai-delivery

Apply MCAF ML/AI delivery guidance for data exploration, feasibility, experimentation, testing, responsible AI, and operating ML systems. Use when the repo includes model training, inference, data science workflows, or ML-specific delivery planning.

LLM Mart · 0 points · 0 views 0 listing impressions 0 install-command copies
Virus-scanned Reviewed automatically before listing.

Full trust report

Download postpartum-genushyacinthus29-dotnet-skills-skills_dotnet-mcaf-ml-ai-delivery-bfa4ebd.zip · 15 KB
Part of postpartum-genushyacinthus29/dotnet-skills — 80 skills

Install

skills CLI npx skills add https://github.com/Postpartum-genushyacinthus29/dotnet-skills/tree/main/skills/dotnet-mcaf-ml-ai-delivery
Claude Code claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install postpartum-genushyacinthus29-dotnet-skills@llmmart
Git git clone https://github.com/Postpartum-genushyacinthus29/dotnet-skills.git

The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole postpartum-genushyacinthus29/dotnet-skills collection as a plugin from our marketplace. Git is the plain clone.

Skill manifest

MCAF: ML/AI Delivery

Trigger On

  • the repo contains model training, inference, experimentation, or data-science workflow
  • ML work needs explicit process, testing, or responsible-AI guidance
  • delivery discussion is mixing product, data, and model concerns

Value

  • produce a concrete project delta: code, docs, config, tests, CI, or review artifact
  • reduce ambiguity through explicit planning, verification, and final validation skills
  • leave reusable project context so future tasks are faster and safer

Do Not Use For

  • generic software delivery with no ML or data-science component
  • loading all ML references when only one stage is active

Inputs

  • the current ML stage: framing, data exploration, experimentation, training, inference, or operations
  • product assumptions, data assumptions, and model assumptions
  • current verification and responsible-AI expectations

Quick Start

  1. Read the nearest AGENTS.md and confirm scope and constraints.
  2. Run this skill's Workflow through the Ralph Loop until outcomes are acceptable.
  3. Return the Required Result Format with concrete artifacts and verification evidence.

Workflow

  1. Separate product assumptions, data assumptions, and model assumptions.
  2. Keep experimentation traceable and testable.
  3. Treat responsible AI, data quality, and ML-specific verification as first-class requirements.
  4. Load only the references that match the current ML stage.

Deliver

  • clearer ML/AI delivery guidance
  • better links between data, experimentation, verification, and responsible AI
  • docs that match how the ML system is built and validated

Validate

  • the active ML stage is explicit
  • experimentation and evaluation are traceable
  • responsible-AI and data-quality requirements are not bolted on at the end

Ralph Loop

Use the Ralph Loop for every task, including docs, architecture, testing, and tooling work.

  1. Brainstorm first (mandatory):
    • analyze current state
    • define the problem, target outcome, constraints, and risks
    • generate options and think through trade-offs before committing
    • capture the recommended direction and open questions
  2. Plan second (mandatory):
    • write a detailed execution plan from the chosen direction
    • list final validation skills to run at the end, with order and reason
  3. Execute one planned step and produce a concrete delta.
  4. Review the result and capture findings with actionable next fixes.
  5. Apply fixes in small batches and rerun the relevant checks or review steps.
  6. Update the plan after each iteration.
  7. Repeat until outcomes are acceptable or only explicit exceptions remain.
  8. If a dependency is missing, bootstrap it or return status: not_applicable with explicit reason and fallback path.

Required Result Format

  • status: complete | clean | improved | configured | not_applicable | blocked
  • plan: concise plan and current iteration step
  • actions_taken: concrete changes made
  • validation_skills: final skills run, or skipped with reasons
  • verification: commands, checks, or review evidence summary
  • remaining: top unresolved items or none

For setup-only requests with no execution, return status: configured and exact next commands.

Load References

  • read references/ml-ai-projects.md first
  • open references/data-exploration.md, references/feasibility-studies.md, references/ml-fundamentals-checklist.md, references/model-experimentation.md, references/testing-data-science-and-mlops-code.md, references/responsible-ai.md, or references/ml-model-checklist.md only when that stage is active

Example Requests

  • "Define the delivery workflow for this ML feature."
  • "We need responsible-AI and testing guidance for this model."
  • "Separate product, data, and model decisions in our docs."
Files (dotnet-skills)
  • references
    • data-exploration.md 410 B
      # Data Exploration
      
      Before training or integrating a model, understand the data you actually have.
      
      ## Questions to Answer
      
      - where the data comes from
      - what permissions and restrictions apply
      - how representative it is of production use
      - what quality issues, bias, gaps, or drift risks exist
      
      ## Deliverables
      
      - data sources and ownership
      - data quality risks
      - assumptions that will affect experimentation
      
    • feasibility-studies.md 8.1 KB
      # Feasibility Studies
      
      The main goal of feasibility studies is to assess whether it is feasible to solve the problem satisfactorily using ML with the available data. We want to avoid investing too much in the solution before we have:
      
      * Sufficient evidence that a solution would be the best technical solution given the business case
      * Sufficient evidence that a solution is compatible with the problem context
      * Sufficient evidence that a solution is possible
      * Some vetted direction on what a solution should look like
      
      This effort ensures quality solutions backed by the appropriate, thorough amount of consideration and evidence.
      
      ## When are Feasibility Studies Useful?
      
      Every engagement can benefit from a feasibility study early in the project.
      
      Architectural discussions can still occur in parallel as the team works towards gaining a solid understanding and definition of what will be built.
      
      Feasibility studies can last between 4-16 weeks, depending on specific problem details, volume of data, state of the data etc. Starting with a 4-week milestone might be useful, during which it can be determined how much more time, if any, is required for completion.
      
      ## Who Collaborates on Feasibility Studies?
      
      Collaboration from individuals with diverse skill sets is desired at this stage, including data scientists, data engineers, software engineers, PMs, human experience researchers, and domain experts. It embraces the use of engineering fundamentals, with some flexibility. For example, not all experimentation requires full test coverage and code review. Experimentation is typically not part of a CI/CD pipeline. Artifacts may live in the `main` branch as a folder excluded from the CI/CD pipeline, or as a separate experimental branch, depending on customer/team preferences.
      
      ## What do Feasibility Studies Entail?
      
      ### Problem Definition and Desired Outcome
      
      * Ensure that the problem is complex enough that coding rules or manual scaling is unrealistic
      * Clear definition of the problem from business and technical perspectives
      
      ### Deep Contextual Understanding
      
      Confirm that the following questions can be answered based on what was learned during the Discovery Phase of the project. For items that can not be satisfactorily answered, undertake additional investigation to answer.
      
      * Understanding the people who are using and/or affected by the solution
      * Understanding the contextual forces at play around the problem, including goals, culture, and historical context
      * To accomplish this a researcher will:
      * Collaborate with customers and colleagues to explore the landscape of people who relate to and may be affected by the problem space being explored (Users, stakeholders, subject matter experts, etc)
      * Formulate the research question(s) to be addressed
      * Select and design research to best serve the research question(s)
      * Identify and select representative research participants across the problem space with whom to conduct the research
      * Construct a research plan and necessary preparation documents for the selected research method(s)
      * Conduct research activity with the participants via the selected method(s)
      * Synthesize, analyze, and interpret research findings
      * Where relevant, build frameworks, artifacts and processes that help explore the findings and implications of the research across the team
      * Share what was uncovered and understood, and the implications thereof across the engagement team and relevant stakeholders.
      * If the above research was conducted during the Discovery phase, it should be reviewed, and any substantial knowledge gaps should be identified and filled by following the above process.
      
      ### Data Access
      
      * Verify that the full team has access to the data
      * Set up a dedicated and/or restricted environment if required
      * Perform any required de-identification or redaction of sensitive information
      * Understand data access requirements (retention, role-based access, etc.)
      
      ### Data Discovery
      
      * Hold a [data exploration](./data-exploration.md) workshop and deep dive with domain experts
      * Understand data availability and confirm the team's access
      * Understand the data dictionary, if available
      * Understand the quality of the data. Is there already a data validation strategy in place?
      * Ensure required data is present in reasonable volumes
      * For supervised problems (most common), assess the availability of labels or data that can be used to effectively approximate labels
      * If applicable, ensure all data can be joined as required and understand how
        * Ideally obtain or create an entity relationship diagram (ERD)
      * Potentially uncover new useful data sources
      
      ### Architecture Discovery
      
      * Clear picture of existing architecture
      * Infrastructure spikes
      
      ### Concept Ideation and Iteration
      
      * Develop value proposition(s) for users and stakeholders based on the contextual understanding developed through the discovery process (e.g. key elements of value, benefits)
      * As relevant, make use of
      * Co-creation with team
      * Co-creation with users and stakeholders
      * As relevant, create vignettes, narratives or other materials to communicate the concept
      * Identify the next set of hypotheses or unknowns to be tested (see concept testing)
      * Revisit and iterate on the concept throughout discovery as understanding of the problem space evolves
      
      ### Exploratory Data Analysis (EDA)
      
      * Data deep dive
      * Understand feature and label value distributions
      * Understand correlations among features and between features and labels
      * Understand data specific problem constraints like missing values, categorical cardinality, potential for data leakage etc.
      * Identify any gaps in data that couldn't be identified in the data discovery phase
      * Pave the way of further understanding of what techniques are applicable
      * Establish a mutual understanding of what data is in or out of scope for feasibility, ensuring that the data in scope is significant for the business
      
      ### Data Pre-Processing
      
      * Happens during EDA and hypothesis testing
      * Feature engineering
      * Sampling
      * Scaling and/or discretization
      * Noise handling
      
      ### Hypothesis Testing
      
      * Design several potential solutions using theoretically applicable algorithms and techniques, starting with the simplest reasonable baseline
      * Train model(s)
      * Evaluate performance and determine if satisfactory
      * Tweak experimental solution designs based on outcomes
      * Iterate
      * Thoroughly document each step and outcome, plus any resulting hypotheses for easy following of the decision-making process
      
      ### Concept Testing
      
      * Where relevant, to test the value proposition, concepts or aspects of the experience
      * Plan user, stakeholder and expert research
      * Develop and design necessary research materials
      * Synthesize and evaluate feedback to incorporate into concept development
      * Continue to iterate and test different elements of the concept as necessary, including testing to best serve RAI goals and guidelines
      * Ensure that the proposed solution and framing are compatible with and acceptable to affected people
      * Ensure that the proposed solution and framing is compatible with existing business goals and context
      
      ### Risk Assessment
      
      * Identification and assessment of risks and constraints
      
      ### Responsible AI
      
      * Consideration of responsible AI principles
      * Understanding of users and stakeholders’ contexts, needs and concerns to inform development of RAI
      * Testing AI concept and experience elements with users and stakeholders
      * Discussion and feedback from diverse perspectives around any responsible AI concerns
      
      ## Output of a Feasibility Study
      
      The main outcome is a feasibility study report, with a recommendation on next steps:
      
      If there is not enough evidence to support the hypothesis that this problem can be solved using ML, as aligned with the pre-determined performance measures and business impact:
      
      * We detail the gaps and challenges that prevented us from reaching a positive outcome
      * We may scope down the project, if applicable
      * We may look at re-scoping the problem taking into account the findings of the feasibility study
      * We assess the possibility to collect more data or improve data quality
      
      If there is enough evidence to support the hypothesis that this problem can be solved using ML
      
      * Provide recommendations and technical assets for moving to the operationalization phase
      
    • ml-ai-projects.md 426 B
      # ML/AI Projects
      
      ML/AI delivery needs separate treatment for data, experimentation, model behaviour, and operations.
      
      ## Workstreams
      
      - problem framing
      - data exploration and quality
      - experimentation and evaluation
      - deployment and monitoring
      - responsible-AI review
      
      ## Practical Rule
      
      Do not let model work collapse into a generic software backlog item.
      Each stage needs explicit assumptions, evidence, and exit criteria.
      
    • ml-fundamentals-checklist.md 2.8 KB
      # ML Fundamentals Checklist
      
      This checklist helps ensure that our ML projects meet our ML Fundamentals. The items below are not sequential, but rather organized by different parts of an ML project.
      
      ## Data Quality and Governance
      
      - [ ] There is access to data.
      - [ ] Labels exist for dataset of interest.
      - [ ] Data quality evaluation.
      - [ ] Able to track data lineage.
      - [ ] Understanding of where the data is coming from and any policies related to data access.
      - [ ] Gather Security and Compliance requirements.
      
      ## Feasibility Study
      
      - [ ] A feasibility study was performed to assess if the data supports the proposed tasks.
      - [ ] Rigorous Exploratory data analysis was performed (including analysis of data distribution).
      - [ ] Hypotheses were tested producing sufficient evidence to either support or reject that an ML approach is feasible to solve the problem.
      - [ ] ROI estimation and risk analysis was performed for the project.
      - [ ] ML outputs/assets can be integrated within the production system.
      - [ ] Recommendations on how to proceed have been documented.
      
      ## Evaluation and Metrics
      
      - [ ] Clear definition of how performance will be measured.
      - [ ] The evaluation metrics are somewhat connected to the success criteria.
      - [ ] The metrics can be calculated with the datasets available.
      - [ ] Evaluation flow can be applied to all versions of the model.
      - [ ] Evaluation code is unit-tested and reviewed by all team members.
      - [ ] Evaluation flow facilitates further results and error analysis.
      
      ## Model Baseline
      
      - [ ] Well-defined baseline model exists and its performance is calculated. ([More details on well defined baselines](./ml-model-checklist.md))
      - [ ] The performance of other ML models can be compared with the model baseline.
      
      ## Experimentation setup
      
      - [ ] Well-defined train/test dataset with labels.
      - [ ] Reproducible and logged experiments in an environment accessible by all data scientists to quickly iterate.
      - [ ] Defined experiments/hypothesis to test.
      - [ ] Results of experiments are documented.
      - [ ] Model hyper parameters are tuned systematically.
      - [ ] Same performance evaluation metrics and consistent datasets are used when comparing candidate models.
      
      ## Production
      
      - [ ] [Model readiness checklist](./ml-model-checklist.md) reviewed.
      - [ ] Model reviews were performed (covering model debugging, reviews of training and evaluation approaches, model performance).
      - [ ] Data pipeline for inferencing, including an end-to-end tests.
      - [ ] SLAs requirements for models are gathered and documented.
      - [ ] Monitoring of data feeds and model output.
      - [ ] Ensure consistent schema is used across the system with expected input/output defined for each component of the pipelines (data processing as well as models).
      - [ ] [Responsible AI](./responsible-ai.md) reviewed.
      
    • ml-model-checklist.md 16.4 KB
      # ML Model Production Checklist
      
      The purpose of this checklist is to make sure that:
      
      - The team assessed if the model is ready for production before moving to the scoring process
      - The team has prepared a production plan for the model
      
      The checklist provides guidelines for creating this production plan. It should be used by teams/organizations that already built/trained an ML model and are now considering putting it into production.
      
      ## Checklist
      
      Before putting an individual ML model into production, the following aspects should be considered:
      
      - [ ] [Is there a well defined baseline? Is the model performing better than the baseline?](#is-there-a-well-defined-baseline-is-the-model-performing-better-than-the-baseline)
      - [ ] [Are machine learning performance metrics defined for both training and scoring?](#are-machine-learning-performance-metrics-defined-for-both-training-and-scoring)
      - [ ] [Is the model benchmarked?](#is-the-model-benchmarked)
      - [ ] [Can ground truth be obtained or inferred in production?](#can-ground-truth-be-obtained-or-inferred-in-production)
      - [ ] [Has the data distribution of training, testing and validation sets been analyzed?](#has-the-data-distribution-of-training-testing-and-validation-sets-been-analyzed)
      - [ ] [Have goals and hard limits for performance, speed of prediction and costs been established so they can be considered if trade-offs need to be made?](#have-goals-and-hard-limits-for-performance-speed-of-prediction-and-costs-been-established-so-they-can-be-considered-if-trade-offs-need-to-be-made)
      - [ ] [How will the model be integrated into other systems, and what impact will it have?](#how-will-the-model-be-integrated-into-other-systems-and-what-impact-will-it-have)
      - [ ] [How will incoming data quality be monitored?](#how-will-incoming-data-quality-be-monitored)
      - [ ] [How will drift in data characteristics be monitored?](#how-will-drift-in-data-characteristics-be-monitored)
      - [ ] [How will performance be monitored?](#how-will-performance-be-monitored)
      - [ ] [Have any ethical concerns been taken into account?](#have-any-ethical-concerns-been-taken-into-account)
      
      Please note that there might be scenarios where it is not possible to check all the items on this checklist. However, it is advised to go through all items and make informed decisions based on your specific use case.
      
      ## Will Your Model Performance be Different in Production than During the Training Phase
      
      Once deployed into production, the model might be performing much worse than expected. This poor performance could be a result of:
      
      - The data to be scored in production is significantly different from the train and test datasets
      - The feature engineering steps are different or inconsistent in production compared to the training process
      - The performance measure is not consistent (for example your test set covers several months of data where the performance metric for production has been calculated for one month of data)
      
      ### Is there a Well-Defined Baseline? Is the Model Performing Better than the Baseline?
      
      A good way to think of a model baseline is the simplest model one can come up with: either a simple threshold, a random guess or a very basic linear model. This baseline is the reference point your model needs to outperform. A well-defined baseline is different for each problem type and there is no one size fits all approach.
      
      As an example, let's consider some common types of machine learning problems:
      
      - **Classification**: Predicting between a positive and a negative class. Either the class with the most observations or a simple logistic regression model can be the baseline.
      - **Regression**: Predicting the house prices in a city. The average house price for the last year or last month, a simple linear regression model, or the previous median house price in a neighborhood could be the baseline.
      - **Image classification**: Building an image classifier to distinguish between cats and no cats in an image. If your classes are unbalanced: 70% cats and 30% no cats and if you always predict cats, your naive classifier has 70% accuracy and this can be your baseline. If your classes are balanced: 52% cats and 48% no cats, then a simple convolutional architecture can be the baseline (1 conv layer + 1 max pooling + 1 dense). Additionally, human accuracy at labelling can also be the baseline in an image classification scenario.
      
      Some questions to ask when comparing to a baseline:
      
      - How does your model compare to a random guess?
      - How does your model performance compare to applying a simple threshold?
      - How does your model compare with always predicting the most common value?
      
      > **Note**: In some cases, human parity might be too ambitious as a baseline, but this should be decided on a case by case basis. Human accuracy is one of the available options, but not the only one.
      
      Resources:
      
      - ["How To Get Baseline Results And Why They Matter" article](https://machinelearningmastery.com/how-to-get-baseline-results-and-why-they-matter/)
      - ["Always start with a stupid model, no exceptions." article](https://blog.insightdatascience.com/always-start-with-a-stupid-model-no-exceptions-3a22314b9aaa)
      
      ### Are Machine Learning Performance Metrics Defined for Both Training and Scoring?
      
      The methodology of translating the training metrics to scoring metrics should be well-defined and understood. Depending on the data type and model, the model metrics calculation might differ in production and in training. For example, the training procedure calculated metrics for a long period of time (a year, a decade) with different seasonal characteristics while the scoring procedure will calculate the metrics per a restricted time interval (for example a week, a month, a quarter). Well-defined ML performance metrics are essential in production so that a decrease or increase in model performance can be accurately detected.
      
      Things to consider:
      
      - In forecasting, if you change the period of assessing the performance, from one month to a year for example, then you might get a different result. For example, if your model is predicting sales of a product per day and the RMSE (Root Mean Squared Error) is very low for the first month the model is in production. As the model is live for longer, the RMSE is increasing, becoming 10x the RMSE for the first year compared to the first month.
      - In a classification scenario, the overall accuracy is good, but the model is performing poorly for some subgroups. For example, a classifier has an accuracy of 80% overall, but only 55% for the 20-30 age group. If this is a significant age group for the production data, then your accuracy might suffer greatly when in production.
      - In scene classification scenario, the model is trying to identify a specific scene in a video, and the model has been trained and tested (80-20 split) on 50000 segments where half are segments containing the scene and half of the segments do not contain the scene. The accuracy on the training set is 85% and 84% on the test set. However, when an entire video is scored, scores are obtained on all segments, and we expect few segments to contain the scene. The accuracy for an entire video is not comparable with the training/test set procedure in this case, hence different metrics should be considered.
      - If sampling techniques (over-sampling, under-sampling) are used to train model when classes are imbalanced, ensure the metrics used during training are comparable with the ones used in scoring.
      - If the number of samples used for training and testing is small, the performance metrics might change significantly as new data is scored.
      
      ### Is the Model Benchmarked?
      
      The trained model to be put into production is well benchmarked if machine learning performance metrics (such as accuracy, recall, RMSE or whatever is appropriate) are measured on the train and test set. Furthermore, the train and test set split should be well documented and reproducible.
      
      ### Can Ground Truth be Obtained or Inferred in Production?
      
      Without a reliable ground truth, the machine learning metrics cannot be calculated. It is important to identify if the ground truth can be obtained as the model is scoring new data by either manual or automatic means. If the ground truth cannot be obtained systematically, other proxies and methodology should be investigated in order to obtain some measure of model performance.
      
      One option is to use humans to manually label samples. One important aspect of human labelling is to take into account the human accuracy. If there are two different individuals labelling an image, the labels will likely be different for some samples. It is important to understand how the labels were obtained to assess the reliability of the ground truth (that is why we talk about human accuracy).
      
      For clarity, let's consider the following examples (by no means an exhaustive list):
      
      - **Forecasting**: Forecasting scenarios are an example of machine learning problems where the ground truth could be obtained in most cases even though a delay might occur. For example, for a model predicting the sales of ice cream in a local shop, the ground truth will be obtained as the sales are happening, but it might appear in the system at a later time than as the model prediction.
      - **Recommender systems**: For recommender system, obtaining the ground truth is a complex problem in most cases as there is no way of identifying the ideal recommendation. For a retail website for example, click/not click, buy/not buy or other user interaction with recommendation can be used as ground truth proxies.
      - **Object detection in images**: For an object detection model, as new images are scored, there are no new labels being generated automatically. One option to obtain the ground truth for the new images is to use people to manually label the images. Human labelling is costly, time-consuming and not 100% accurate, so in most cases, only a subset of images can be labelled. These samples can be chosen at random or by using active learning techniques of selecting the most informative unlabeled samples.
      
      ### Has the Data Distribution of Training, Testing and Validation Sets Been Analyzed?
      
      The data distribution of your training, test and validation (if applicable) dataset (including labels) should be analyzed to ensure they all come from the same distribution. If this is not the case, some options to consider are: re-shuffling,  re-sampling, modifying the data, more samples need to be gathered or features removed from the dataset.
      
      Significant differences in the data distributions of the different datasets can greatly impact the performance of the model. Some potential questions to ask:
      
      - How much does the training and test data represent the end result?
      - Is the distribution of each individual feature consistent across all your datasets? (i.e. same representation of age groups, gender, race etc.)
      - Is there any data lineage information? Where did the data come from? How was the data collected? Can collection and labelling be automated?
      
      Resources:
      
      - ["Splitting into train, dev and test" tutorial](http://cs230.stanford.edu/blog/split/)
      
      ### Have Goals and Hard Limits for Performance, Speed of Prediction and Costs been Established, so they can be Considered if Trade-Offs Need to be Made?
      
      Some machine learning models achieve high ML performance, but they are costly and time-consuming to run. In those cases, a less performant and cheaper model could be preferred. Hence, it is important to calculate the model performance metrics (accuracy, precision, recall, RMSE etc), but also to gather data on how expensive it will be to run the model and how long it will take to run. Once this data is gathered, an informed decision should be made on what model to productionize.
      
      System metrics to consider:
      
      - CPU/GPU/memory usage
      - Cost per prediction
      - Time taken to make a prediction
      
      ### How Will the Model be Integrated into Other Systems, and what Impact will it Have?
      
      Machine Learning models do not exist in isolation, but rather they are part of a much larger system. These systems could be old, proprietary systems or new systems being developed as a results of the creation a new machine learning model. In both of those cases, it is important to understand where the actual model is going to fit in, what output is expected from the model and how that output is going to be used by the larger system. Additionally, it is essential to decide if the model will be used for batch and/or real-time inference as production paths might differ.
      
      Possible questions to assess model impact:
      
      - Is there a human in the loop?
      - How is feedback collected through the system? (for example how do we know if a prediction is wrong)
      - Is there a fallback mechanism when things go wrong?
      - Is the system transparent that there is a model making a prediction and what data is used to make this prediction?
      - What is the cost of a wrong prediction?
      
      ### How Will Incoming Data Quality be Monitored?
      
      As data systems become increasingly complex in the mainstream, it is especially vital to employ data quality monitoring, alerting and rectification protocols. Following data validation best practices can prevent insidious issues from creeping into machine learning models that, at best, reduce the usefulness of the model, and at worst, introduce harm. Data validation, reduces the risk of data downtime (increasing headroom) and technical debt and supports long-term success of machine learning models and other applications that rely on the data.
      
      Data validation best practices include:
      
      - Employing automated data quality testing processes at each stage of the data pipeline
      - Re-routing data that fails quality tests to a separate data store for diagnosis and resolution
      - Employing end-to-end data observability on data freshness, distribution, volume, schema and lineage
      
      Note that data validation is distinct from data drift detection. Data validation detects errors in the data (ex. a datum is outside of the expected range), while data drift detection uncovers legitimate changes in the data that are truly representative of the phenomenon being modeled (ex. user preferences change). Data validation issues should trigger re-routing and rectification, while data drift should trigger adaptation or retraining of a model.
      
      Resources:
      
      - ["Data Quality Fundamentals" by Moses et al.](https://www.oreilly.com/library/view/data-quality-fundamentals/9781098112035/)
      
      ### How Will Drift in Data Characteristics be Monitored?
      
      Data drift detection uncovers legitimate changes in incoming data that are truly representative of the phenomenon being modeled,and are not erroneous (ex. user preferences change). It is imperative to understand if the new data in production will be significantly different from the data in the training phase. It is also important to check that the data distribution information can be obtained for any of the new data coming in. Drift monitoring can inform when changes are occurring and what their characteristics are (ex. abrupt vs gradual) and guide effective adaptation or retraining strategies to maintain performance.
      
      Possible questions to ask:
      
      - What are some examples of drift, or deviation from the norm, that have been experience in the past or that might be expected?
      - Is there a drift detection strategy in place? Does it align with expected types of changes?
      - Are there warnings when anomalies in input data are occurring?
      - Is there an adaptation strategy in place? Does it align with expected types of changes?
      
      Resources:
      
      - ["Learning Under Concept Drift: A Review" by Lu at al.](https://arxiv.org/pdf/2004.05785.pdf)
      - [Understanding dataset shift](https://towardsdatascience.com/understanding-dataset-shift-f2a5a262a766)
      
      ### How Will Performance be Monitored?
      
      It is important to define how the model will be monitored when it is in production and how that data is going to be used to make decisions. For example, when will a model need retraining as the performance has degraded and how to identify what are the underlying causes of this degradation could be part of this monitoring methodology.
      
      Ideally, model monitoring should be done automatically. However, if this is not possible, then there should be a manual periodical check of the model performance.
      
      Model monitoring should lead to:
      
      - Ability to identify changes in model performance
      - Warnings when anomalies in model output are occurring
      - Retraining decisions and adaptation strategy
      
      ### Have any Ethical Concerns Been Taken into Account?
      
      Every ML project goes through the [Responsible AI](responsible-ai.md) process to ensure that it upholds Microsoft's [6 Responsible AI principles](https://www.microsoft.com/en-us/ai/responsible-ai).
      
    • model-experimentation.md 463 B
      # Model Experimentation
      
      Experimentation should be traceable and comparable.
      
      ## Rules
      
      - record datasets, parameters, and evaluation outputs
      - compare against a baseline, not only against intuition
      - separate quick exploration from decisions that affect production
      - keep success criteria visible before running long experiment loops
      
      ## Smells
      
      - experiments with no saved assumptions
      - metric improvement with no product interpretation
      - cherry-picked results
      
    • responsible-ai.md 388 B
      # Responsible AI
      
      Responsible-AI work is part of delivery, not a final checklist.
      
      ## Review Areas
      
      - harm scenarios
      - bias and representational gaps
      - explainability needs
      - privacy and data minimization
      - fallback behaviour when the model is wrong or uncertain
      
      ## Rule
      
      If the model can materially affect users or decisions, document the harm model and mitigation plan before release.
      
    • testing-data-science-and-mlops-code.md 3.2 KB
      # Testing Data Science and MLOps Code
      
      Testing ML code follows the same rule as the rest of MCAF: verify real behaviour with real boundaries.
      
      ## Core Rules
      
      - test data pipelines, feature transforms, training code, and inference code through real execution paths
      - do not use mocks, fakes, stubs, or service doubles
      - do not hardcode important values inline; keep reusable test values in named constants or fixtures
      - keep datasets tiny, deterministic, and representative
      - separate fast checks from heavier end-to-end or training checks
      
      ## Data Loading
      
      Use small real files or checked-in fixtures instead of replacing file I/O with doubles.
      
      Recommended pattern:
      
      1. keep a tiny CSV, JSON, image, or parquet fixture in the test assets
      2. load it through the same public function or pipeline used by production code
      3. assert the parsed schema, row count, key values, and validation behaviour
      
      Good checks:
      
      - valid file loads successfully
      - missing file fails in the expected way
      - invalid schema is rejected clearly
      - nulls, outliers, or malformed records are handled explicitly
      
      ## Data Transformation
      
      Transformation tests should prove input-to-output behaviour with small deterministic examples.
      
      Use:
      
      - named fixture constants for sample values
      - one assertion focus per test when possible
      - parametrized tests for shape, range, normalization, encoding, or padding rules
      
      Avoid:
      
      - inline magic numbers or string literals repeated across tests
      - giant synthetic datasets that make failures unreadable
      
      ## Model Training and Inference
      
      Use tiny real models or reduced training configurations that still execute the true code path.
      
      Examples of useful checks:
      
      - the model accepts the expected input shape
      - the training step updates weights or parameters
      - inference produces the expected output shape and type
      - the saved model can be loaded and used by the real inference entry point
      - evaluation code calculates metrics correctly on a fixed fixture dataset
      
      Keep long-running verification separate from the fast inner loop by marking or grouping those suites explicitly.
      
      ## Integration and Pipeline Tests
      
      For MLOps flows, prefer narrow but real integration coverage:
      
      - real feature pipeline plus real model inference
      - real model artifact load plus API or batch entry point
      - real data validation and schema enforcement
      - real persistence of model metadata, metrics, or lineage where the system depends on it
      
      If a dependency is external, use a real sandbox or test environment with the real contract.
      
      ## Validation and Monitoring
      
      ML verification should cover more than code execution.
      
      Test or document:
      
      - schema validation
      - feature drift detection logic
      - threshold and alert calculations
      - fallback behaviour when model output is invalid or unavailable
      - reproducibility of evaluation inputs and metrics
      
      ## Test Design Rules
      
      - prefer TDD for bug fixes and non-trivial behaviour changes
      - use named constants for file names, labels, thresholds, and expected values
      - keep fixtures small enough to understand at a glance
      - prove behaviour through public interfaces, not internal implementation details
      - if a test is hard to write without doubles, treat that as a design problem and simplify the boundary
      
  • SKILL.md 4.2 KB
    ---
    name: dotnet-mcaf-ml-ai-delivery
    version: "1.0.0"
    category: "AI"
    description: "Apply MCAF ML/AI delivery guidance for data exploration, feasibility, experimentation, testing, responsible AI, and operating ML systems. Use when the repo includes model training, inference, data science workflows, or ML-specific delivery planning."
    compatibility: "Requires repository access when ML/AI docs, experiments, or delivery guidance live in the repo."
    ---
    
    # MCAF: ML/AI Delivery
    
    ## Trigger On
    
    - the repo contains model training, inference, experimentation, or data-science workflow
    - ML work needs explicit process, testing, or responsible-AI guidance
    - delivery discussion is mixing product, data, and model concerns
    
    ## Value
    
    - produce a concrete project delta: code, docs, config, tests, CI, or review artifact
    - reduce ambiguity through explicit planning, verification, and final validation skills
    - leave reusable project context so future tasks are faster and safer
    
    ## Do Not Use For
    
    - generic software delivery with no ML or data-science component
    - loading all ML references when only one stage is active
    
    ## Inputs
    
    - the current ML stage: framing, data exploration, experimentation, training, inference, or operations
    - product assumptions, data assumptions, and model assumptions
    - current verification and responsible-AI expectations
    
    ## Quick Start
    
    1. Read the nearest `AGENTS.md` and confirm scope and constraints.
    2. Run this skill's `Workflow` through the `Ralph Loop` until outcomes are acceptable.
    3. Return the `Required Result Format` with concrete artifacts and verification evidence.
    
    ## Workflow
    
    1. Separate product assumptions, data assumptions, and model assumptions.
    2. Keep experimentation traceable and testable.
    3. Treat responsible AI, data quality, and ML-specific verification as first-class requirements.
    4. Load only the references that match the current ML stage.
    
    ## Deliver
    
    - clearer ML/AI delivery guidance
    - better links between data, experimentation, verification, and responsible AI
    - docs that match how the ML system is built and validated
    
    ## Validate
    
    - the active ML stage is explicit
    - experimentation and evaluation are traceable
    - responsible-AI and data-quality requirements are not bolted on at the end
    
    ## Ralph Loop
    
    Use the Ralph Loop for every task, including docs, architecture, testing, and tooling work.
    
    1. Brainstorm first (mandatory):
       - analyze current state
       - define the problem, target outcome, constraints, and risks
       - generate options and think through trade-offs before committing
       - capture the recommended direction and open questions
    2. Plan second (mandatory):
       - write a detailed execution plan from the chosen direction
       - list final validation skills to run at the end, with order and reason
    3. Execute one planned step and produce a concrete delta.
    4. Review the result and capture findings with actionable next fixes.
    5. Apply fixes in small batches and rerun the relevant checks or review steps.
    6. Update the plan after each iteration.
    7. Repeat until outcomes are acceptable or only explicit exceptions remain.
    8. If a dependency is missing, bootstrap it or return `status: not_applicable` with explicit reason and fallback path.
    
    ### Required Result Format
    
    - `status`: `complete` | `clean` | `improved` | `configured` | `not_applicable` | `blocked`
    - `plan`: concise plan and current iteration step
    - `actions_taken`: concrete changes made
    - `validation_skills`: final skills run, or skipped with reasons
    - `verification`: commands, checks, or review evidence summary
    - `remaining`: top unresolved items or `none`
    
    For setup-only requests with no execution, return `status: configured` and exact next commands.
    
    ## Load References
    
    - read `references/ml-ai-projects.md` first
    - open `references/data-exploration.md`, `references/feasibility-studies.md`, `references/ml-fundamentals-checklist.md`, `references/model-experimentation.md`, `references/testing-data-science-and-mlops-code.md`, `references/responsible-ai.md`, or `references/ml-model-checklist.md` only when that stage is active
    
    ## Example Requests
    
    - "Define the delivery workflow for this ML feature."
    - "We need responsible-AI and testing guidance for this model."
    - "Separate product, data, and model decisions in our docs."
    

Comments (0)

Sign in to join the conversation.

No comments yet.

Reviews (0)

No reviews yet.

Related