Can We Trust AI Decisions? What Data Can Tell Us About Bias and Fairness
AI systems increasingly influence decisions about hiring, credit, insurance and access to services. Statistical and econometric tools can test whether different groups receive systematically different outcomes, but measuring disparities is only the beginning: researchers must determine why differences appear, whether the comparison is valid and what definition of fairness is appropriate.
Artificial intelligence increasingly helps organizations decide which job applicants receive interviews, which transactions appear suspicious, which customers receive recommendations and how people are classified or ranked. Because algorithms apply mathematical rules rather than displaying human emotions, it can be tempting to assume that automated decisions are inherently objective. Yet an AI model learns relationships from data generated by real institutions, markets and human behavior, meaning historical inequalities and measurement problems can enter the system. Removing a human decision-maker does not automatically remove bias from the decision.
This does not mean that every difference in AI outcomes represents discrimination. Two groups can receive different outcomes because their observable characteristics differ, because the model performs differently across populations, because the training data are unrepresentative or because the decision process itself contains structural disparities. Statistical analysis is therefore necessary to distinguish a raw difference from a more meaningful fairness problem. The first question is not simply “Did the groups receive different outcomes?” but “Why did those outcomes differ?”
NIST treats fairness as one of several characteristics of trustworthy AI and specifically describes trustworthy systems as being fair with harmful bias managed. Its framework also emphasizes validity, reliability, safety, security, transparency, accountability, explainability and privacy, illustrating that fairness cannot be evaluated independently from the broader system. NIST further identifies systemic, computational/statistical and human-cognitive forms of bias rather than treating bias as a single technical defect.
This is where statistics and econometrics become particularly valuable. Researchers can compare selection rates, error rates and model performance across groups, control for relevant characteristics and examine whether disparities persist after accounting for legitimate explanatory factors. Data cannot decide what society should consider fair, but it can reveal patterns that would otherwise remain hidden inside an algorithm.
Statistical Tests Can Reveal Whether Groups Receive Different Outcomes
The simplest fairness analysis begins with outcomes. Suppose an AI hiring system evaluates 10,000 applicants and recommends 60% of applicants from Group A but only 40% from Group B for interviews. The raw selection-rate gap is immediately visible, but its existence alone does not establish why it occurred.
Researchers can calculate selection rates for each group and compare their absolute or relative differences. Similar comparisons can be made for loan approvals, insurance decisions, fraud alerts, school admissions or any other process producing observable outcomes. A fairness audit begins by making disparities measurable rather than relying on assumptions about whether an algorithm “looks neutral.”
Statistical hypothesis testing can then examine whether an observed difference is plausibly attributable to sampling variation. Confidence intervals can show the uncertainty surrounding estimated differences, while sufficiently large datasets can help identify patterns across demographic groups and subgroups. NIST’s 2024 Generative AI Profile specifically recommends appropriate fairness metrics and statistical hypothesis tests for business processes with categorical or numerical outcomes that rely on generative AI.
But raw outcome rates provide only one dimension of fairness. A model can produce similar approval rates for two groups while making substantially more mistakes for one of them, or it can produce different approval rates while achieving comparable accuracy conditional on relevant characteristics. Equal outcomes and equal model performance are not the same statistical objective.
This is why analysts also examine error rates. In a binary classification problem, false positives occur when the model incorrectly predicts a positive outcome, while false negatives occur when it incorrectly rejects something that should have received a positive classification. If those errors are distributed very differently across groups, the aggregate accuracy rate can hide an important disparity.
Imagine an AI system that is 90% accurate overall. That figure sounds impressive, but it becomes far less reassuring if accuracy is 95% for one population and 70% for another. Average accuracy can conceal unequal reliability.
Researchers should therefore disaggregate performance. Precision, recall, false-positive rates, false-negative rates and calibration can be examined across relevant groups depending on the application. The appropriate metric depends heavily on what the algorithm actually does and what consequences its errors create.
Econometrics Helps Ask Whether the Difference Remains After Other Factors Are Considered
Raw comparisons can identify disparities, but econometric analysis can investigate them more deeply. Suppose an automated lending system approves fewer applicants from one demographic group. Comparing approval rates reveals the difference, but it does not determine whether the difference comes from income, debt, credit history, loan characteristics, model design or another factor.
Regression analysis can estimate the association between group membership and outcomes while controlling for observable characteristics considered relevant to the decision. Researchers might compare otherwise similar applicants based on income, credit history, debt burden and requested loan size. If a substantial group difference remains after appropriate controls are introduced, that residual disparity becomes an important signal requiring further investigation—but it is not automatically proof of unlawful discrimination.
That final qualification matters enormously. Regression results depend on the variables included, their measurement quality, the functional form of the model and assumptions about the data-generating process. Important characteristics may be unobserved, poorly measured or endogenous, meaning that statistical controls do not magically create a perfect experiment.
Researchers must also be cautious about controlling for variables that are themselves affected by historical or institutional disparities. A variable can appear statistically relevant while partly transmitting the very mechanism being investigated. A model can be mathematically sophisticated and still answer the wrong causal question.
Matched comparisons can provide another approach. Researchers can compare individuals with similar observable characteristics who differ in the demographic dimension being examined, although matching likewise depends on the quality and completeness of observed data. Where circumstances permit, experimental or quasi-experimental research designs can provide stronger causal evidence.
Counterfactual testing can also be informative. An auditor might construct otherwise similar inputs and alter a characteristic—or a feature closely associated with it—to determine whether model outputs change. NIST recommends counterfactual testing among the methods organizations can use to assess harmful bias in generative AI systems.
These approaches move fairness analysis beyond simple averages. Econometrics is useful because it forces investigators to specify what comparison they believe is meaningful and what alternative explanations must be considered. The result is not a perfect definition of fairness, but a more disciplined investigation of where disparities originate.
Bias Can Enter Through Data Even When Sensitive Variables Are Removed
One common response to concerns about AI discrimination is to remove variables such as race or gender from the model. That can be useful in some contexts, but it does not guarantee that the resulting algorithm is neutral. Other variables may contain correlated information that allows the model to reproduce similar patterns indirectly.
Location provides an intuitive example. A postal code may appear to be purely geographic information, but residential patterns can correlate with income, race, ethnicity, access to services and historical housing conditions. A sufficiently powerful machine-learning model may therefore reconstruct demographic information through combinations of apparently neutral variables.
NIST explicitly recommends examining whether model features can act as proxies for demographic group membership, including characteristics embedded in language, metadata and other inputs. It also recommends examining training and evaluation data for representativeness, subgroup coverage and systemic bias.
Historical labels create an even deeper problem. Suppose a company trains a hiring model to predict which applicants resemble employees who were historically successful within the organization. If historical recruitment or promotion patterns favored particular backgrounds, the algorithm can learn those relationships even if no developer intentionally programs discriminatory rules.
AI does not need an explicit instruction to discriminate; it can learn statistical patterns produced by earlier decisions. The resulting model may then reproduce those patterns at much greater scale because automated systems can evaluate thousands or millions of cases.
Bias can also emerge from missing data. Some populations may be poorly represented in the training dataset, causing the model to perform reliably for the majority population but unpredictably for smaller groups. Brookings has recently described a related problem as algorithmic exclusion, where insufficient data can prevent meaningful predictions for certain people or populations altogether.
Data quality therefore needs to be evaluated separately for different populations. Analysts should examine sample sizes, missing values, measurement quality, coverage and whether the deployment population resembles the population represented in training and testing data. A dataset can be enormous and still be unrepresentative of the people most affected by the algorithm.
This also explains why simply testing an AI system once before deployment is inadequate. Customer behavior changes, labor markets evolve, economic conditions shift and the population using a system can differ from the population on which it was originally evaluated. Fairness and performance can consequently deteriorate even when the underlying model code remains unchanged.
There Is No Single Mathematical Definition of Fairness
One of the most difficult facts about algorithmic fairness is that different definitions can conflict. A company might want equal selection rates across demographic groups, equal error rates, equal true-positive rates or equally calibrated predictions. Each represents a different conception of what it means for the algorithm to behave fairly.
Demographic parity generally asks whether groups receive a positive outcome at similar rates. In hiring, for example, researchers might compare the proportion of applicants from each group recommended for an interview. Large differences can reveal an important disparity, but demographic parity does not explain whether applicants had similar underlying characteristics relevant to the decision.
Equal opportunity focuses on whether qualified individuals from different groups have similar chances of receiving a positive classification. Equalized odds goes further by considering both true-positive and false-positive rates across groups. Calibration asks whether a given predicted probability has a similar meaning across populations.
NIST’s Generative AI Profile specifically lists demographic parity, equalized odds and equal opportunity among fairness metrics that may be appropriate for evaluating business processes involving AI. Crucially, NIST also recommends context-specific metrics developed with domain experts and affected communities rather than assuming that one universal statistic defines fairness.
This creates unavoidable trade-offs. A system may improve one fairness metric while worsening another, especially when underlying outcome rates differ between groups. “Make the algorithm fair” is therefore not a complete technical specification; someone must decide what type of fairness matters for the particular decision and why.
Context becomes decisive. In medical screening, missing a genuinely sick patient may have different consequences from incorrectly flagging a healthy patient. In fraud detection, false accusations and missed fraudulent transactions impose different costs, while in hiring the consequences of incorrectly excluding qualified candidates raise another set of concerns.
Fairness consequently cannot be delegated entirely to data scientists. Legal experts, economists, domain specialists, management and affected communities may all need to participate in determining which outcomes and errors matter. NIST similarly emphasizes that fairness includes equality and equity concerns and that perceptions and standards of fairness can differ across applications and contexts.
Statistics can measure a fairness criterion precisely, but statistics cannot decide which fairness criterion society should value. That remains a governance and policy judgment informed by evidence rather than replaced by it.
The question “Can we trust AI?” is too broad to answer with a universal yes or no. Trust depends on the particular system, the decision being made, the quality of the underlying data, the consequences of errors and the controls surrounding deployment. A recommendation algorithm choosing movies creates fundamentally different risks from an algorithm influencing employment, credit or medical decisions.
What data can provide is evidence. Selection rates can reveal outcome disparities, error-rate analysis can show whether the model performs differently across groups, econometric methods can test whether differences persist after relevant observable characteristics are considered, and counterfactual tests can probe how outputs respond to changes in inputs. Fairness becomes more manageable when organizations convert vague concerns about bias into measurable questions.
But measurement is not sufficient on its own. Algorithmic audits need to examine data collection, missing observations, model performance, subgroup outcomes, proxies and the way organizations actually use model outputs. Brookings’ work on employment-algorithm auditing similarly emphasizes that audits should examine selection rates, fairness metrics, subgroup performance, training data and the possibility that variables act as proxies for protected characteristics.
Monitoring also needs to continue after deployment. NIST’s AI Risk Management Framework organizes risk management around Govern, Map, Measure and Manage, emphasizing that AI risk should be addressed across the lifecycle rather than through a single pre-launch test. NIST is currently revising AI RMF 1.0, so organizations relying on it should also follow updates to the framework.
The most important conceptual shift is to stop treating algorithms as inherently neutral simply because they use mathematics. Human decisions determine which data are collected, what outcome is predicted, which variables are included, what error rate is acceptable and how the resulting score is used. AI can automate a decision process, but it cannot remove the economic, institutional and social context in which that decision occurs.
At the same time, algorithmic decision-making creates an opportunity that purely informal human judgment often does not provide. When systems generate measurable scores and outcomes, researchers can systematically compare groups, test hypotheses, identify disparities and evaluate interventions. AI can create bias, reproduce bias or sometimes reduce bias—but data give us tools to determine which of those is actually happening.
Ultimately, trustworthy AI requires more than high predictive accuracy. It requires evidence that the system works for the populations it affects, transparent decisions about which fairness objectives matter, continuous testing and clear accountability when outcomes become problematic. The right question is not whether AI deserves unconditional trust, but whether a particular AI decision system has produced enough evidence to justify the trust being placed in it.
