Using AI honestly: ethics, academic integrity, and what watermarks change
What AI actually is, what it is genuinely changing, and where the line sits between using it well and letting it do your thinking for you — plus what Claude's new invisible watermarks really prove.
Most guidance about AI and honesty is either "don't use it" or "use it responsibly" — and neither tells you what to actually do on a Tuesday night with a deadline. This is the more specific version: what these tools really are, what they're genuinely changing, where the line sits, and what Claude's new invisible watermarks do and don't prove.
The problem
"Use AI ethically" is advice nobody can act on
Almost every school, employer, and platform now has an AI policy. Most of them say some version of "use it responsibly" and stop there. That leaves you guessing at the boundary — and guessing badly, because the honest answer isn't a single rule. Using AI to explain a concept you're stuck on and using it to produce the work you'll be graded on are not the same act, even though both are "using AI."
The useful questions aren't about whether AI touched your work. They're about what it touched, whether you can stand behind the result, and whether anyone reading it would feel misled if they knew. Those are answerable. "Was AI involved?" mostly isn't — and as of August 2026, it's a question that machines have started trying to answer on their own.
First principles
What AI actually is — and why that matters here
A large language model is a prediction engine. It has read an enormous amount of text and learned which words tend to follow which other words in which contexts. When you prompt it, it produces the continuation that fits those patterns. That's it. That single fact explains nearly everything useful and everything dangerous about these tools.
It explains why the output reads so well: fluent text is exactly what it was optimised to produce. It also explains why the output can be confidently, specifically wrong. The model isn't consulting a database of facts and reporting back — it's generating text that looks like a correct answer. A fabricated citation and a real one are produced by the same process, and they look identical coming out.
So the model has no idea whether what it just told you is true. It has no sense of having guessed. There is no internal flag that goes up when it invents a statistic, a court case, or a source. This isn't a bug that a better version will quietly fix — it's what prediction means.
Why this frames the whole ethics question: If the tool can't vouch for its own output, someone has to. The moment you put your name on something, that someone is you. Most AI-related trouble — academic, professional, and legal — traces back to skipping that step rather than to using AI at all.
The effects
What genuinely changed, and what quietly got worse
The real gains are in the first draft, not the final one
Where these tools deliver is unblocking: getting a bad first draft on the page so you have something to fix, explaining a concept five different ways until one lands, translating between languages or between jargon and plain English, finding the bug you've stared past six times, and arguing against your own idea so you can see where it's weak. In every one of those, the AI is doing work that helps you do the thinking — not replacing it.
The cost is a skill you never notice losing
The struggle you skip is the struggle that teaches you. Working out how to structure an argument is how you learn to structure arguments. Sitting with a bug is how you learn to debug. Outsource that repeatedly and the work still gets done — you just stop getting better at it, and you won't feel it happening, because the output looks fine the whole time.
This is the part that policies can't police and detection tools can't catch, and it's the part that actually costs you. Nobody gets hurt more by AI-assisted shortcuts than the person taking them, and the bill arrives later — in the interview, the viva, the first week of a job where the thing you supposedly learned turns out to be missing.
The question worth asking isn't "will I get caught?" It's "in six months, will I be able to do this myself?"
The framework
Five rules that hold up in practice
These aren't about avoiding punishment. They're the working habits of people who use AI heavily and still produce trustworthy work.
Verify anything you'd be embarrassed to get wrong
Every factual claim, statistic, citation, quote, date, and name gets checked against a real source before it goes anywhere. Not because AI is unusually unreliable, but because it's unreliable in a way that doesn't announce itself — the confident tone is identical whether the model knows or is inventing.
Never submit something you couldn't explain out loud
If someone asked you to walk them through your reasoning, defend a choice, or extend the argument one step further, could you? If not, you don't understand the work well enough to put your name on it — regardless of who or what wrote it. This single test catches almost every genuine integrity problem.
Disclose when the reader would want to know
The test isn't whether a rule requires it, it's whether someone would feel misled on finding out. A translated email? Nobody cares. A personal essay about your own experience, a peer review, a piece of work being assessed as evidence of your ability? Very different. When unsure, ask — asking has never once been the thing that got someone in trouble.
Don't feed it what isn't yours to share
Client data, medical details, unreleased work, colleagues' personal information, anything under NDA — treat a chat box like a third party you haven't signed an agreement with, because that's what it is. This is the AI rule most often broken by people who are otherwise scrupulous, usually without a moment's thought.
Keep the judgement calls
Use it to generate options, stress-test reasoning, and handle the mechanical parts. Keep the decisions about what matters, what's true, what's worth saying, and what you're willing to defend. The moment those move to the model, you've stopped being the author in any sense that counts.
For students
Academic honesty, specifically
Academic rules on AI are genuinely inconsistent right now — not because institutions are careless, but because the technology moved faster than the policies. One lecturer encourages AI for brainstorming and bans it for drafting. Another allows drafting with disclosure. A third treats any use as misconduct. These are all defensible positions, and you cannot infer which one applies from common sense.
So read the actual policy for the actual course, and when it's ambiguous, ask in writing and keep the reply. "I assumed it was allowed" is not a defence anywhere. "Here's the email where I asked and was told yes" ends the conversation immediately.
The trap most students fall into: Using AI to produce work you don't yet understand, then discovering the gap during an oral defence, an exam, or a follow-up question you can't answer. The submission passes; the moment someone asks you to elaborate, it doesn't.
The distinction that actually matters
Using AI to understand — explain this proof differently, why is my code failing, what's the counterargument here, quiz me on this chapter — builds the knowledge the assessment is trying to measure. Using AI to produce — write this essay, do this problem set — replaces the evidence of learning with something that only resembles it. The first is closer to a tutor. The second defeats the purpose of the exercise, whether or not it's technically permitted.
That line isn't always clean, and pretending otherwise is dishonest. But it's a far more useful compass than "did I use AI," and it maps closely onto what most institutions are actually trying to protect.
What's new
Claude now watermarks what it writes
In August 2026, Anthropic began embedding invisible watermarks in text generated by Claude — initially in models launched on or after 2 August, with older models due to follow during the EU AI Act's transition period — applied worldwide and driven partly by compliance with that Act. It uses SynthID-Text, a method developed by Google DeepMind. Supported file types like .png, .jpg and .svg also get signed provenance metadata under the C2PA standard — the same system Adobe and Google use.
Read the primary source: Anthropic's own write-up explains the mechanism and, more usefully, is candid about what the watermark cannot establish. How Claude's text watermarking works →
The mechanism is subtler than people assume. Nothing is inserted — no hidden characters, no invisible unicode, no extra tokens. As the model generates, it constantly reaches points where several different words would work equally well. Normally that choice comes from an arbitrary random source; with watermarking, it comes from a key-based one instead. The result reads identically, but the pattern of choices is statistically verifiable if you hold the key.
What it can't do — which is the important part
Anthropic is unusually direct about the limits, and they matter more than the feature. The watermark cannot distinguish between text Claude wrote and text Claude merely edited. It's unreliable on short passages, and sparse on factual writing where exact wording is forced. Light editing tends to preserve it; a genuine full rewrite removes it. It says nothing about which user or organisation was involved. And it obviously cannot detect output from any other AI system.
Crucially, it also cannot confirm that something was written by a human. Absence of a watermark proves nothing at all — the text might be human-written, might be from another model, might be Claude output that was rewritten. A detection API is coming, but it estimates likelihood; it does not deliver verdicts.
So don't draw the wrong conclusion in either direction: This is not a cheating detector, and treating it as one will produce false accusations. It's also not a reason to feel safe — the real cost of outsourcing your thinking was never detection. Watermarking is about provenance in the information ecosystem, not policing individual students.
In practice
A checklist you can actually run
- I've read the specific AI policy for this specific course, job, or publication
- Every fact, figure, quote, and citation has been verified against a real source
- I could explain and defend every part of this out loud, unprompted
- I haven't pasted in anything confidential or anyone else's personal data
- The judgement calls — what's true, what matters, what to say — were mine
- If the reader learned exactly how I used AI here, they wouldn't feel misled
Six lines, and they cover essentially every AI-integrity failure worth worrying about. None of them require you to avoid AI. They require you to remain the author.
Start building — The strongest defence against using AI badly is understanding how it works from the inside. Our Builder track takes you from prompting to shipping real agents — the kind of knowledge that makes the honesty questions obvious rather than agonising.
Frequently asked questions
Is using AI to write an essay cheating?
It depends on the rules you agreed to, and they vary widely between institutions and even between courses. But a more useful test than "is it against the rules" is this: could you explain and defend every part of the work out loud? If not, the submission misrepresents what you know — which is the thing academic integrity rules exist to prevent, whether or not AI is mentioned in the policy.
Can teachers detect AI-generated writing?
Not reliably. Commercial AI detectors produce both false positives and false negatives at rates that make them unsafe to accuse anyone on, and they have been shown to flag non-native English writers disproportionately. Claude's watermarking is more technically robust than those detectors, but it only covers Claude output, can't tell writing from editing, and doesn't work well on short passages. In practice, teachers more often notice a mismatch between submitted work and what a student can discuss in person.
Does Claude watermark everything it writes?
From August 2026, supported Claude models embed an invisible statistical watermark in generated text, worldwide, across Claude products including the API and Claude Code. Generated files such as .png, .jpg and .svg also carry signed C2PA provenance metadata. The watermark is detectable more reliably in longer passages, and is sparse in factual text where the wording is largely forced.
Can an AI watermark be removed?
Yes — a genuine full rewrite where the wording is substantively replaced removes it, since the watermark lives in the pattern of word choices rather than in any hidden characters. Light editing usually preserves it. This is exactly why watermarking should be understood as a provenance signal for the information ecosystem rather than as an enforcement mechanism against individuals.
Is it okay to use AI for homework?
Using it to understand — explaining a concept differently, quizzing you, showing you why your code fails, arguing the counterposition — generally supports the learning the homework is meant to produce. Using it to generate the answers you submit generally replaces the evidence of that learning with something that only resembles it. Check the specific policy, and if it's ambiguous, ask in writing and keep the reply.
If my work has no watermark, does that prove I wrote it?
No. Absence of a watermark proves nothing: the text could be human-written, could come from a different AI system entirely, or could be Claude output that was rewritten. Anthropic is explicit that the watermark cannot confirm human authorship, cannot identify who produced a piece of text, and cannot establish ownership or responsibility for it.