What does Responsible AI really mean?
A simple yardstick to measure if you use AI properly or not
“How can we teach our staff to use AI properly?” is a question I am often asked by people leaders.
The question originates from concerns that employees are not using AI responsibly: they are over-delegating business decisions to AI, not thoroughly checking AI outputs, seeing their skills eroded as they defer too much to AI, and facing declining productivity when AI slop saves time for them but slows down the system as a whole.
In organisations, the usual response to these problems is to develop guidance on the “responsible use of AI”, which is typically interpreted to mean a set of policies, guidelines and frameworks that set boundaries on how AI should be used within an organisation to ensure that the use of AI is ethical, safe and productive.
I’ve worked in the area of responsible AI for years. Indeed, I even wrote a book on the topic. And I’ve found a couple of problems with this approach to responsible AI. First, policies are useless if they are not enforced. And most organisations develop the policies but not the enforcement framework. Second, and more fundamentally, responsible AI is too often described in very abstract terms. The framing is around ethics, values and corporate regulation. All of these are important, but thinking of responsible AI only at such a high level doesn’t provide any concrete guidance to staff using a generative AI tool daily, wondering what exactly is the right way to use AI.
For the most part, an average staff member won’t encounter deep ethical questions about their use of AI. More likely, they’ll be using AI to produce a document and will be more concerned with where do the boundaries lie in terms of how they can and should use AI in producing that document. Is it okay to use AI at all? If yes, is any use of AI sanctioned as long as AI outputs are checked? Or are there some ways of using AI that are off limits?
Responsible AI policies don’t address these very real questions. And it’s these very practical considerations that I aim to address in this article.
Responsible AI in Day to Day Work
Consider the following examples. Which of these do you consider to be responsible use of AI?
The reference letter. You are asked to provide a reference letter for a former colleague. Time is tight and you’re busy with other things, so you give a few dot points to AI and ask for a draft. You agree with what it produces and so send it off. Nothing is untrue, but the recipient thinks this is your considered judgement.
The tender response. You use AI to draft the bulk of a tender response, verifying every claim. You win the work, but find out later that a much smaller competitor with a better solution, who wrote their bid manually, lost out.
The innovation prize. You enter a competition to come up with ideas to improve your local community. The prize is $1000. You use AI to generate entries and pick the best one. You win.
People tend to have different reactions to these scenarios. In the case of the reference letter, some will say this use of AI is fine. The letter is accurate; AI didn’t make anything up. Others, however, will be horrified at this use of AI: a reference letter is a medium of trust between two parties and that trusted relationship shouldn’t be delegated to an AI.
For the innovation prize, does it matter that you won a prize but didn’t have the idea? Many people will say yes, some going so far as to say it’s dishonest. However, what if we modify this scenario slightly so that you don’t simply have AI generate ideas and pick one, but use AI to brainstorm different ideas, including putting your own thinking into the winning choice. Is this way of using AI okay? Or should the prize only go to a fully human-generated idea?
These dilemmas are not hypothetical. Workers are faced with situations like these every day, and need to ask themselves: when is it okay to use AI and when is it not? And does the way in which I use AI matter? If so, what is the right way?
Ways to use AI
One of the key points of this article is that the use of AI is not a black and white issue. Take AI in writing, for example. There is a great deal of backlash against any use of AI in writing beyond spellchecking, improving grammar, or research before writing begins. It’s as though there are only two modes for AI in writing: either use it in a very minimal way, or go all in and generate an entire article and use as is. Of course, most uses of AI are neither of these. In practice, AI use is much more nuanced. An author might use AI to brainstorm ideas, write some text him/herself, then ask AI to improve the flow of the text. In my view, this shouldn’t be banned. Although there are plenty of commentators out there who say I am wrong.
The point of this article isn’t to debate whether AI-assisted writing should be allowed or not. It’s to show that there are different ways of using AI in any knowledge task. And understanding how we use AI is useful as a sanity check to understand whether we can legitimately claim ideas as our own or have really thought through the reasoning ourself.
Inspired by Danny Liu’s categorisation about different ways to use AI, I’ve come up with my own four-category AI-use classification. Each category captures one way of interacting with AI. Whilst the “best” category will change given a particular context, in general, you should aim to be in the Partner category. This is where you get the best of both worlds: your ideas and thinking are still yours, but AI helps you refine them. Other categories aren’t always wrong, but, in general, I’d be aiming to be in the Partner category the majority of time.
Surrender: You accept AI output as final without checking it. You write a prompt, and other than a few tweaks, you accept AI’s first answer as the answer.
Verify: You check AI’s output, carefully catching errors, confirming claims and asking for sources, but you’re still letting AI do most of the work, although you’re careful not to trust it more than you should.
Partner. You start a conversation with AI by bringing your own ideas to the table. From there, you actively shape the process, setting constraints up front, redirecting AI mid-task, bringing your own expertise to challenge the framing (not just the facts), and iterating towards something that is better than either you or AI alone could produce.
Reject. You dismiss AI’s output wholesale and redo the work yourself, or route around it entirely out of distrust or because you consider the AI output to be substandard and unfixable.
Applying these categories to the three scenarios above illustrates that, for any given task, AI can be used in very different ways:
The reference letter. It’s unlikely that Surrender would produce a good result here. You’d need a “mega prompt” to give AI all the necessary context about the person you are providing a recommendation for. Verify could potentially produce good results, but again you’d need to include a lot in the context of the prompt. Partner means that you have a conversation with AI, where you collaboratively explore the strengths and weaknesses of the person you are recommending; only once that’s fully done would you ask AI to generate an actual letter (or write it yourself). Under Reject, you try to use AI to draft a letter, but get frustrated by its generic nature and decide not to use AI at all.
The tender response. Surrender here is risky: a generic prompt produces boilerplate that ticks compliance boxes but ignores what actually differentiates you from competitors bidding on the same tender. Verify can work if you feed AI a very detailed brief and then rigorously check every claim against your actual capability, but AI is still driving the strategy. Partner means working through your key differentiators with AI first, arguing about what the client really values, stress-testing your angle, and only drafting once that thinking is done. Under Reject, you try AI on the tender, find the response too generic to fix, and write the whole thing yourself.
The innovation prize. Surrender produces a submission that reads like every other entry: competent but forgettable, because AI has no way of knowing what’s genuinely novel about your idea from a thin prompt. Verify gets you further if you supply rich detail on the innovation and then check that AI hasn’t overstated or misdescribed the technical claims. Partner means using AI as a co-designer, asking for and giving ideas, and refining those together. Under Reject, you generate an idea with AI, find it boring, and can’t imagine how AI could help you improve it.
So, which category are you?
There isn’t a “right” answer when it comes to which AI category you should use. Most likely, you will be in a different category at different times, depending on the context. For example, if all you want is an answer to a simple question like “Which actor is the lead in Obsession?” then Surrender is perfectly fine. However, if the task is to come up with a strategy document for a 5000-strong organisation, I wouldn’t want to be in any category other than Partner. The Reject category is also perfectly fine. There are plenty of instances where I’ve found AI can’t really understand what I’m trying to achieve, and doing it myself is quicker (and frankly, less stressful).
Generally speaking, however, I would say you should aim to be in the Partner category the majority of the time. Which begs the question: which mode do you personally tend to be in most of the time?
In the absence of a scientifically rigorous instrument to measure your interactions with AI and assign a category, we can take a less rigorous (but still useful) approach: ask AI itself!
Here is a prompt you can use to ask your favourite genAI tool to assess how you interact with it. Note that this prompt will work straightaway in Claude Pro and Max, which can access your full history of conversations. Other genAI tools, depending on your setup and subscription, may not be able to immediately analyse past conversations. In the worst case, you may need to cut/paste prior conversations into the current chat; 15 should be sufficient. In any case, try the prompt first. If your AI tool can’t provide an answer, it should let you know what you need to do to help it.
What you should get back is a (very) approximate set of percentages, one per category, based on your conversation history. Obviously, take these numbers with a very large pinch of salt. You can also ask AI to give you examples of interactions you’ve had in each category. The point here isn’t to get a definitive score but to use the exercise to self-reflect on how you typically use AI and see if you want to make any changes.
Here’s Claude’s response when I entered the prompt above:
Reject is harder to measure, of course, as it often ends up with me leaving Claude altogether.
You can also get spot-checks from AI to sanity-check how you think you are interacting with it. Try this prompt at any point:
Again, depending on which AI tool you are using, it might understand straight away or need further context. This evening, I was using Claude to help me with some branding for my company. The above prompt got the following answer:
Teaching Responsible AI
I’ve taught students of all shapes and sizes over many decades in various career roles. And one thing I’ve learnt is that you have to meet students where they are, not where you think they should be. Which means presenting concepts in simple, easy to understand ways. I think the responsible AI ecosystem is getting this wrong.
When I first started researching responsible AI six years ago, the problem was a lack of guidance. Now, the problem is the opposite: there are hundreds of frameworks, principles, reports, books and podcasts out there, each promising to teach you exactly how to use AI responsibly. And yet, most of these frameworks are overly complex, try to do too much, and ultimately don’t solve the problems that AI users really have: how do I know if I’m applying AI responsibly, not in some abstract sense, but here and now?
What I’ve tried to do in this article is to redress the balance, even if only a little. Personally, I find it’s helpful to reflect on how I use AI from time to time, and, if necessary, to adapt. My simple four-category model provides a quick way of assessing AI use. And, perhaps most importantly, it pushes back against the black-and-white thinking that seems to dominate AI discussions nowadays. It’s never “all in on AI” or “avoid it like the plague”. AI is a really useful tool, but, like any tool, can be misused. So, think about which category of use you mostly fit into, and then ask if that’s where you want to be.






Excellent idea: My statistics were,
* Partner — ~60% (dominant mode): writing/policy work, iwiki.au, article series — you originate ideas, redirect on substance, catch overclaims, supply facts Claude can't know
* Verify — ~25%: research-heavy claims (provenance chains, sourcing, ATO dispute) — you check findings before accepting them
* Surrender — small, low-stakes only: mechanical fixes (Mac camera bug, mesh network, car drain) where there's little to verify
* Reject — ~0%: no instances found of discarding Claude's work wholesale or routing around it
Great piece, John. I’ve developed the SECURE GenAI Use Framework (https://secureframework.ai) and the AI Inherent Risk Scale (AIIRS https://aiirs.ai) to provide guidance to people on using AI. I’d love to hear your thoughts on these.