Sam Austin on October 9, 2026

Model Cards and Datasheets: Document ML Models Responsibly

Model Cards and Datasheets: Document ML Models Responsibly
Contents

An open binder of labeled documents next to a laptop, the paper trail a model needs

Figure 1: The documentation every model should ship with

Six months after launch, a colleague asks, "Can this model handle applicants from our new region?" You open the repo and find a notebook, a pickle file, and silence. Nobody wrote down what the model was built to do.

Documentation fixes that, and two formats lead the way: model cards and datasheets for datasets. This guide shows you what goes in each, how they differ, and how to generate a real card with code I tested.

Why Documentation Matters

A model without documentation behaves like a medication without a label. It might work, but nobody knows the dose, the side effects, or who shouldn't take it. Ever inherited a model nobody could explain? You know the feeling.

Good documentation does three jobs:

  1. It sets boundaries. Readers learn what the model does and doesn't do.
  2. It exposes uneven performance. Disaggregated results show who the model serves poorly.
  3. It creates accountability. Named authors and contacts mean someone answers questions.

Model Cards: The Basics

Margaret Mitchell and colleagues introduced model cards in a 2019 paper presented at FAT*. A model card is a short document that travels with a trained model. It reports what the model is, how you should use it, and how it performs across different conditions.

The paper recommends nine sections:

  1. Model Details: who built it, when, which version, and under which license.
  2. Intended Use: the planned uses and users, plus out-of-scope uses.
  3. Factors: the groups and conditions that may affect performance.
  4. Metrics: the measures and thresholds you chose, and why.
  5. Evaluation Data: the datasets you tested on.
  6. Training Data: what you know about the data the model learned from.
  7. Quantitative Analyses: results broken down by factor and by intersections of factors.
  8. Ethical Considerations: sensitive data, risks, and mitigations.
  9. Caveats and Recommendations: what remains unresolved.

I love section 7: it forces you to report results per group, not just one flattering average.

Datasheets for Datasets: The Basics

Timnit Gebru and co-authors proposed datasheets for datasets in a 2018 preprint, with a revised version published later. They borrowed the idea from electronics, where every component ships with a datasheet describing its specs and limits.

A datasheet records why someone created a dataset, what it contains, how they collected it, and what uses it supports. The paper's goal is better communication between dataset creators and dataset users. In practice, the datasheet answers questions you'll otherwise guess at later:

  • Motivation: why did anyone build this dataset, and who funded it?
  • Composition: what do the rows represent, and which groups appear?
  • Collection process: how did the data arrive, and did people consent?
  • Recommended uses: what tasks fit, and what tasks don't?

FYI, the original paper organizes its questions around the dataset lifecycle, so you can read it like a checklist while you build — the same discipline the training-data bias checks guide ends with.

How They Fit Together

The two documents answer different questions: a datasheet describes the ingredients, a model card describes the dish.

Datasheet Model card
Documents A dataset A trained model
Main question What's in this data and how did it get here? What does this model do, and for whom does it work?
Written by Dataset creators Model developers
Key content Motivation, composition, collection, uses Intended use, disaggregated metrics, limits

Link them. Your model card's training-data section should point to the datasheet instead of repeating it. IMO, that single habit saves more confusion than any template.

Build a Model Card: What Good Content Looks Like

Templates help, but content matters more. Here's how to fill the hardest sections well.

Intended and Out-of-Scope Use

Write specific sentences. "Decision support for pre-screening loan applications, with a human reviewing every denial" beats "financial applications." Then name foreseeable misuse, like fully automated denials or deployment in regions the data never covered.

Disaggregated Metrics

Report performance per subgroup, not only overall. I ran a small test with a loan model trained on synthetic data where women make up about 30% of rows. The overall numbers hid a gap:

Group n Accuracy Recall Selection rate
Overall 1,800 0.682 0.633 0.444
Female 556 0.719 0.509 0.288
Male 1,244 0.665 0.675 0.514

Look at the female row: accuracy looks higher than for men, but recall drops to 0.509, so the model misses qualified women far more often. That story vanishes if you report only the overall column. The data is synthetic, so treat these numbers as an illustration, not a finding about real lending — and if you want to generate gaps like these on purpose first, the Fairlearn tutorial builds exactly this kind of audited pipeline.

Limitations and Caveats

State what you didn't test. Small groups produce noisy metrics, so say so. Honest uncertainty earns more trust than a clean-looking table.

Generate a Card With Code

You can build cards by hand in Markdown, or generate them from your evaluation code. Generating them keeps numbers and documentation in sync. The script below computes disaggregated metrics with pandas and scikit-learn, then writes the card:

def row(mask, name):
    return {"group": name, "n": int(mask.sum()),
            "accuracy": accuracy_score(y_te[mask], pred[mask]),
            "recall": recall_score(y_te[mask], pred[mask]),
            "precision": precision_score(y_te[mask], pred[mask]),
            "selection_rate": pred[mask].mean()}

rows = [row(np.ones(len(y_te), bool), "overall")]
rows += [row(s_te == g, g) for g in ["female", "male"]]
table = pd.DataFrame(rows).round(3)

md = f"""# {card['name']}

{card['summary']}

## Intended Use
{card['intended_use']}

## Evaluation (disaggregated)
{table.to_markdown(index=False)}

## Caveats and Recommendations
{card['caveats']}
"""
open("MODEL_CARD.md", "w").write(md)

I ran the full version end to end, and it wrote a complete card — a single make_model_card.py script holds the whole thing so you can rerun it on every retrain. Note that to_markdown needs the tabulate package.

Use the Hugging Face Tooling

If you publish on the Hugging Face Hub, its library gives you a ready-made workflow. The huggingface_hub package builds a card from metadata (ModelCardData) and a Markdown body (ModelCard), and its docs show this pattern:

from huggingface_hub import ModelCard, ModelCardData, EvalResult

card_data = ModelCardData(
    language="en",
    license="mit",
    model_name="my-cool-model",
    eval_results=EvalResult(
        task_type="image-classification",
        dataset_type="beans",
        dataset_name="Beans",
        metric_type="accuracy",
        metric_value=0.7,
    ),
)

card = ModelCard.from_template(
    card_data,
    model_id="my-cool-model",
    model_description="This model does this and that.",
    developers="Your Name",
    repo="https://github.com/you/your-repo",
)
card.push_to_hub("username/my-cool-model")

The Hub's annotated template also assigns roles: a developer fills the technical sections, a "sociotechnic" (such as an ethicist or lawyer) fills bias and risk, and a project organizer fills details and uses. I like that split — one person rarely has all the context. I couldn't run this snippet in my sandbox, so test it against your installed version.

Regulation Now Asks for This Work

Documentation isn't only good practice anymore. The EU AI Act requires technical documentation for high-risk AI systems before they reach the market, and the documentation must stay current as the system changes. Annex IV lists its contents, including general description, development and design, performance metrics, and risk management.

Dates have moved. The EU's AI Omnibus (Regulation (EU) 2026/1744) entered into force on 27 July 2026, according to one legal summary, and it pushes standalone high-risk systems to 2 December 2027 and embedded ones to 2 August 2028. Obligations for general-purpose AI models stay on their original timetable.

Don't treat a model card as a compliance document by itself. It helps, because it covers many of the same questions, but regulators expect more. Ask counsel what your system needs.

Keep Your Documentation Alive

A stale card misleads worse than no card. Build habits that keep documentation current:

  1. Version the card with the model. Store MODEL_CARD.md in the same repository and tag them together.
  2. Regenerate metrics automatically. Script the evaluation tables so they update on every retrain.
  3. Review on every change. New data, new thresholds, and new use cases all trigger an update.
  4. Name a contact. Someone must answer questions and accept corrections.
  5. Record known failures. Add incidents and fixes to the caveats section.

Common Mistakes

I've written most of these cards badly at some point:

  • Reporting only aggregate metrics. Averages hide the groups that suffer.
  • Writing vague intended use. "General purpose" tells readers nothing.
  • Hiding the limitations. Honest caveats build trust.
  • Copying a template without thinking. Empty sections look worse than missing ones.
  • Writing the card once. Documentation decays unless you maintain it.
CoverBookDescriptionGet it
Cover of “Model Cards for Model Reporting” Model Cards for Model Reporting the 2019 paper behind this guide's nine sections. View on Amazon
Datasheets for Datasetsby Gebru and colleagues the dataset-side companion, cited by the bias-checks guide in this series. View on Amazon
Cover of “Designing Machine Learning Systems” Designing Machine Learning Systemsby Chip Huyen puts documentation inside the full production loop. View on Amazon

Unlock AI That Actually Works

Get lifetime access to GPT-6 Astra, Claude Fable 5.1, Gemini 3.5, Grok 4.5, and more — all in one platform. Build websites, apps, videos, content, and digital products from a single command. No monthly fees. No tool-hopping.

Click here to get GPTAstra Max now — one-time payment, lifetime access.

Frequently Asked Questions

What is a model card and what should it contain?

A model card is a short document that travels with a trained model, reporting what the model is, how you should use it, and how it performs across different conditions. The 2019 paper by Mitchell and colleagues recommends nine sections: model details, intended use, factors, metrics, evaluation data, training data, quantitative analyses broken down by factor and their intersections, ethical considerations, and caveats and recommendations. The disaggregated analyses section is the one that forces you to report results per group, not just one flattering average.

What is the difference between a model card and a datasheet?

A datasheet describes the ingredients; a model card describes the dish. Datasheets for datasets, proposed by Gebru and co-authors in 2018, document a dataset — motivation, composition, collection process, and recommended uses — and are written by dataset creators. Model cards document a trained model — intended use, disaggregated metrics, and limits — and are written by model developers. Link them instead of duplicating: the model card's training-data section should point to the datasheet.

Why report disaggregated metrics instead of overall accuracy?

Averages hide the groups that suffer. In the synthetic loan example, overall accuracy was 0.682, but the female row showed higher accuracy (0.719) with recall down at 0.509 — the model misses qualified women far more often, and that story vanishes in the aggregate column. Report performance per subgroup and per intersection of subgroups, note the sample size in each row, and state that small groups produce noisy metrics rather than hiding the uncertainty.

Does the EU AI Act require model documentation?

Yes for high-risk systems: the EU AI Act requires technical documentation before high-risk AI systems reach the market, and the documentation must stay current as the system changes, with Annex IV listing its contents (general description, development and design, performance metrics, risk management). Dates have moved — the EU AI Omnibus (Regulation (EU) 2026/1744) entered into force on 27 July 2026 and pushes standalone high-risk systems to 2 December 2027 and embedded ones to 2 August 2028. A model card helps but is not a compliance document by itself; ask counsel what your system needs.

Wrapping This Up

Datasheets document the data, and model cards document the model. Together they tell future users what you built, why you built it, and where it breaks. The effort takes hours, and it saves weeks of confusion later.

Pick one model you already run this week. Write its intended use, out-of-scope use, and a disaggregated results table. If the table makes you uncomfortable, good. You just found something worth fixing before your users did.

What are You Looking For?

esc