Back to Blog
AI for Business October 5, 2026 9 min read

Testing an AI System for Bias: What 'Reasonable Measures' Looks Like in Practice

Canadian human rights law does not care whether you intended to discriminate. It cares about the effect. That makes untested AI in hiring or service decisions a genuine liability.

By Aparna Netheti

Testing an AI System for Bias: What 'Reasonable Measures' Looks Like in Practice

Testing an AI system for bias means checking whether its outcomes differ across groups in a way you cannot justify, and writing down what you found. That is the whole exercise. It does not require a data science team, and skipping it is riskier than most Canadian businesses realise.

Here is why. Human rights legislation across Canadian jurisdictions addresses discriminatory effect, not just intent. A hiring tool that systematically ranks one group lower is a problem even if nobody chose that outcome and nobody can explain it. "The vendor's model did it" is not a defence anyone should want to test.

What can you actually test without a research team?

Three things, in increasing order of effort. Most companies can do the first two in a week.

Outcome rates by group. Take the decisions the system produced over a period. Compare the positive outcome rate across the groups you can lawfully observe or reasonably infer. If applicants from one group advance at half the rate of another, you have something to explain, and you would rather explain it to yourself than to a tribunal.

Input review. List every feature the system sees. Look for proxies: postal code, name, graduation year, employment gaps, photograph, language of the submission. Each of these correlates with a protected characteristic. Removing a protected field while keeping five proxies for it accomplishes nothing.

Paired testing. Build matched pairs that differ in exactly one attribute, run them through, and compare. This is the technique fair-housing researchers have used for decades, it needs no model access, and it works on a vendor's black box. It is also the single most convincing artefact you can hand an auditor.

TestEffortWhat it catches
Outcome rates by groupLowSystematic disparity in real decisions
Input and proxy reviewLowFeatures that encode protected characteristics
Paired testingMediumDifferential treatment on one changed attribute
Subgroup accuracyMediumA model that works well overall and badly for a minority
Ongoing monitoringMediumDrift after deployment, which is where most problems appear

The measurement problem nobody warns you about

You often cannot measure disparity because you do not collect the demographic data that would let you.

This is a real tension in Canada. Privacy principles push you to collect less. Fairness testing needs the very attributes you avoided collecting. Both instincts are correct, and the resolution is not to collect protected characteristics into your operational systems.

What works: collect them separately, voluntarily, with clear purpose limitation, held apart from the decision system and used only for aggregate analysis. Or use paired testing, which sidesteps the problem entirely because you construct the inputs yourself. Or partner with the vendor and ask what testing they have done, on what population, and get the answer in writing.

What does not work: concluding that because you cannot measure it, you have no duty to look. That reasoning does not survive contact with a complaint.

What does "reasonable" mean here?

Proportionate to the stakes and the reach of the decision, documented, and repeated.

For a tool that drafts marketing copy, reasonable is close to nothing. For a tool that screens job applicants at volume, reasonable means testing before deployment, monitoring after, having a human able to override, and re-testing when the model or the applicant pool changes. The gap between those two is enormous, which is why the algorithmic impact assessment comes first: it tells you which end of the scale you are on.

Three practical markers that a programme is reasonable rather than decorative:

  • The test was defined before the results were seen, so the threshold was not chosen to pass.
  • Someone had the authority to stop deployment based on the result, and that person was not the system's owner.
  • The finding was recorded, including when the finding was "no material difference found", with the date and the method.

That third one matters more than it sounds. A negative result you can produce a year later is worth far more than a confident memory.

What do you do when you find something?

Not necessarily switch off the system. Four legitimate responses, and picking one deliberately is the point.

Fix the inputs, if a proxy is doing the damage. Adjust the threshold or the process, for example by sending borderline cases to human review rather than auto-rejecting. Restrict the use, so the tool ranks but does not filter. Or retire it, if the disparity is material and unexplainable.

Whichever you choose, record the reasoning, the approver, and the date in your incident log or assessment record. An organization that found a disparity, considered it, and acted is in a defensible position. An organization that never looked is not, and the difference between them is entirely documentary.

What about tools you bought?

Ask the vendor four questions and record the answers verbatim: what fairness testing has been done, on what population, what disparities were found, and what you are expected to test on your own data. A vendor that has done none of this is not disqualified, but it moves the testing obligation onto you, and you should price that in before signing. Keep the answers with the entry in your vendor inventory.

This is general information, not legal advice.

Valdra keeps bias testing attached to the system it belongs to, with the method, the date, and the decision that followed, so the record exists when someone asks a year later. The AI governance view is where that lives.

Frequently asked questions

Do Canadian businesses have to test AI systems for bias?+

No statute names bias testing as a standalone duty, but human rights legislation across Canadian jurisdictions addresses discriminatory effect rather than intent. An untested system that produces disparate outcomes in hiring or service decisions is a real exposure regardless of what anyone intended.

Can you test for bias without a data science team?+

Yes. Comparing positive outcome rates across groups, reviewing inputs for proxies such as postal code or graduation year, and running paired tests that differ on one attribute all require ordinary analytical skill and no access to the model itself.

What if we do not collect demographic data?+

Collect it separately and voluntarily with a clear limited purpose, held apart from the decision system and used only in aggregate, or use paired testing, which sidesteps the issue because you construct the inputs. Concluding you have no duty to look because you cannot measure does not hold up.

Does removing protected fields solve bias?+

No. Proxies carry the same information. Postal code, name, graduation year, employment gaps, and submission language all correlate with protected characteristics, so removing one field while keeping several proxies for it changes very little.

What should we do if a test finds a disparity?+

Choose deliberately among fixing the inputs, adjusting thresholds or routing borderline cases to human review, restricting the tool to ranking rather than filtering, or retiring it. Then record the reasoning, the approver, and the date, because acting on a finding is what makes the position defensible.

AI bias testingalgorithmic fairness Canadadiscrimination AI hiringAI audit biashuman rights AI Canadadisparate impact testing

AI governance and privacy compliance, simplified.

Valdra helps Canadian companies govern AI and meet PIPEDA and Law 25 — hosted in Canada.

Try Valdra

Our own compliance

We run our own compliance programme inside Valdra — the product we sell. Our SOC 2, ISO 27001 and ISO 42001 programmes are actively in progress; we do not claim certifications we do not yet hold.

Valdra compliance badge — click to verify
  • PIPEDA
  • Law 25 (Quebec)
  • CASL
  • Data hosted in Canada 🇨🇦
  • AI governance
View our Trust Centre

Self-declared, not audited by a third party. Click the badge to verify it is genuine and see what it covers.