Tickd.ai
← The Tickd Guide

Ethics & Responsible Use

Who Decides What's 'Ethical' AI? The Bias Behind the Guardrails

Every model has values baked in — from its training data and from the people who tuned it. The interesting question is not whether bias exists but who gets to choose the correction.

Updated 9/13/2026

Every AI assistant you use has a position on contested questions. Not because it was programmed with opinions, but because two things gave it some: the text it learned from, and the humans who shaped its behaviour afterwards.

Bias from data

Models learn statistical patterns from enormous corpora of human writing. That writing over-represents some languages, countries, professions and eras, and carries our historical assumptions with it. The documented consequences are unglamorous and real: uneven quality across languages, stereotyped associations between occupations and genders, worse performance on dialects and names outside the training distribution, and unequal accuracy in vision systems across skin tones.

These are measurable. They also move — the same model family can improve on one benchmark and regress on another between versions.

Bias from tuning

Then comes the guardrail layer. Labs decide what a model refuses, how it hedges, which viewpoints it presents as settled, and what tone it adopts. Those decisions are made by policy teams and encoded through fine-tuning and RLHF, often with rater guidelines the public never sees.

This is where the harder argument lives. Refusing to help synthesise a nerve agent is uncontroversial. Refusing to summarise one side of a political argument, or presenting a contested empirical question as closed, is a judgement call with a thumb on the scale.

The competing critiques

Two criticisms are made loudly, usually by different people, and they do not cancel out.

Guardrails are too weak. Models still produce discriminatory outputs, still get jailbroken, and still get deployed into hiring, lending and moderation decisions where errors land on the people least able to appeal. On this view, self-imposed policy is no substitute for auditing and liability.

Guardrails are too heavy, and unaccountable. A small number of companies are setting de facto speech norms for hundreds of millions of users, with policies written internally, changed without notice, and reflecting a narrow cultural vantage point. Refusals are inconsistent, sometimes condescending, and rarely explained.

Both critiques describe things that are actually happening. Holding them at the same time is not a contradiction — it is just an accurate description of an immature industry.

What the labs are trying

Approaches differ, and comparing them is instructive. Some publish an explicit constitution or model spec so the intended values can be argued with in public. Some run bias benchmarks and report the results. Some let deployers adjust the safety layer for their own context. Some let users set persistent preferences on tone and directness. Deciding which of these is more honest than the others is a matter of taste, and worth forming your own view on by reading the documents themselves rather than the press coverage.

What to do as a user

Treat the assistant as an editor with opinions, not an oracle. Ask the same contested question of two or three platforms — OpenAI, Claude, Gemini, Grok — and the differences in framing will tell you more about the guardrails than any policy page. When accuracy matters, ask for sources and check them. When a refusal seems absurd, rephrase, and notice that a system whose rules you can accidentally route around is a system whose rules were never that principled to begin with.

More on the underlying mechanics: the alignment problem, explained simply.

ethicsbiasfairnessguardrails

Keep going

Build something with the prompt generator, decode the jargon in the glossary, or compare the tools on our platform deep-dives.