top of page

Your AI Can Do the Job. Should You Let It?

Writer: Jenny Kay Pollock
Jenny Kay Pollock
4 days ago
6 min read

Updated: 2 days ago

Purple graphic with bold CAN and SHOULD text, a check mark and question mark in boxes, with small star accents.

Divya Sabade is an AI Product Leader with experience across ML-powered decision systems, enterprise technology, and agentic AI. She learns by building and writing about AI products, agents, trust, and the product decisions behind deploying AI in the real world.


What building AI products taught me about autonomy, human judgment, and trust few years ago, some of the most important questions I asked when building ML-powered products was, “Can the model do this? Can it identify risky behavior? Make the right recommendation? Help someone make a better decision?"

Today, as AI moves from answering questions to taking actions, I find myself asking a different question, “Should we let it?”And increasingly, a third, “How much freedom has it earned?”

I spent years working on ML-powered decision systems, including seller risk and performance products at Walmart. We dealt with a version of this problem long before we started calling software “agents.” A model could be very good at predicting something. That didn’t automatically mean it should make the final decision.


Now AI systems can retrieve information, call tools, make decisions, and take actions across real workflows. That distinction matters even more. The more I build with agents, the more I think one of the defining AI product skills will be “Judgment,” deciding not only what AI can do, but what we should allow it to do, under what conditions, and when that should change.


Here are three lessons I keep coming back to.


1. Capability Doesn’t Equal Permission

Imagine an AI customer-support agent that can understand a refund request, inspect the order, interpret the policy, and issue the refund:


  • For a $20 refund with clear policy coverage, perhaps we let it act.

  • What about $2,000? Maybe a person should review it.

  • Now imagine the agent has successfully handled thousands of similar refunds. Perhaps we loosen that boundary.

  • Then one afternoon it attempts 50 refunds in eight minutes.


The model hasn’t necessarily changed. Its permissions may not have changed. But our willingness to let it continue probably should. The distinction I have found useful: 

Capability tells us what AI can do. Judgment determines how much freedom it has earned.

That freedom doesn’t have to be permanent. A system might act independently in one situation, require approval in another, and stop completely in a third.

That is  not an inconsistency. That is  good product design.

A Simple Test I Use

Consequence | Confidence | Reversibility

When I am  thinking about how much autonomy an AI system should have, I come back to three things:

  1. Consequence: What happens if it’s wrong?

  2. Confidence: How strong is the evidence that it will behave correctly here?

  3. Reversibility: Can we easily undo what it does?


Low consequence, strong confidence, and easy reversal? More autonomy may make sense.

High consequence, uncertainty, and an action that is difficult to reverse? Keep human judgment closer.

This isn’t a formula. It’s a way to make the trade-off explicit. And you don’t have to be an AI researcher to use it. If you are a product manager, designer, engineer, founder, or business leader deciding what AI should be allowed to do, you’re already making autonomy decisions.

2. “Human in the Loop” Doesn’t Automatically Mean Human Judgment

The obvious response to AI uncertainty is often: Put a human in the loop.

Sometimes that is exactly right. But adding an approval button doesn’t guarantee meaningful oversight.

Anthropic has shared an interesting example from Claude Code that users approved roughly 93% of permission prompts. The company has discussed the risk of approval fatigue and subsequently introduced stronger sandboxing and other controls to reduce unnecessary permission prompts.

That finding stuck with me because it exposes an easy product trap: A human click is not the same thing as human judgment.

If I approve the same low-risk action 200 times, at some point I may no longer be evaluating it. I am clearing a queue.


Instead of asking, Where should we put a human in the loop? We should be asking: Where can human judgment actually change the outcome?

For a routine, reversible action the AI has performed reliably many times, constant approval may add little value. For an unusual, consequential, or difficult-to-reverse decision, human attention becomes much more valuable. As AI gets better, the human role may shift from doing the work to approving it, supervising it, and eventually handling exceptions.


The goal isn’t to remove people. It is to spend human judgment where it matters most.


3. Trust Should Change When the Evidence Changes

Purple gauge infographic titled Trust Is a Dial, Not a Switch, with needle near Earned Autonomy and Low Trust label.

This brings me to the part of AI product development I find most interesting right now which is Evals. Before deployment, we evaluate whether an AI system behaves the way we expect. But passing an eval doesn’t mean an AI is now “trustworthy.” It tells us how the system performed under the conditions we tested.

The real world keeps introducing conditions we didn't know about such as new users, new data, new tools, strange edge cases, and unexpected sequences of actions. Anthropic’s work on agent evaluations highlights why this gets harder with agents. The final answer isn’t always enough to evaluate success. The trajectory, tool calls, intermediate behavior, and actual outcome can all matter.

Production can also reveal things that strong pre-deployment testing misses. 

OpenAI recently described a limited internal deployment of a model designed for long-running tasks where the team observed behaviors its existing evaluations hadn’t captured. The response is what interests me from a product perspective: investigate what happened, turn those observations into new evaluations, strengthen safeguards and monitoring, test again, and adjust deployment.

That suggests a different product lifecycle.

Instead of Build → Evaluate → Launch

I increasingly think about Build → Evaluate → Deploy narrowly → Observe → Learn → Adjust

If evidence becomes stronger, perhaps the system earns another tool, a higher transaction threshold, or one less approval. If failures emerge, people increasingly override it, or the environment changes, perhaps some of that autonomy should contract.


Trust isn’t a one time launch decision. It should change as the evidence changes. And every meaningful failure should teach us something. A customer complaint can become an eval. A human override can reveal a missing edge case. An unexpected tool sequence can become a regression test. A near miss can show us where a boundary needs to change. The product shouldn’t simply experience failures. It should learn how to test for them.


What This Changes for AI Product Leaders

When I started exploring these questions, I thought I was looking at several different problems:

  • When should AI stop? 

  • When should a person approve? 

  • How much autonomy should an agent receive? 

  • How should we evaluate whether it’s working?


I increasingly think they are different versions of the same product. How do we give increasingly capable AI systems more freedom without giving up our ability to understand, intervene, and change course?


For product leaders, that changes the job. We are not only defining what the AI should accomplish. We’re designing the conditions under which it can act, when it should ask for help, and what evidence should change those boundaries.


I have spent much of my career thinking about how machines make decisions. What surprises me is that the more capable AI becomes, the more important the human questions seem to become:

  • Should it act? 

  • What happens if it’s wrong? 

  • When should someone step in? 

  • What evidence would make us trust it more? 

  • And what should make us take some of that freedom back?


We will undoubtedly build more capable AI.The harder, and perhaps more important, work is developing the judgment to decide what we should let it do.


Further Reading

Anthropic, Demystifying evals for AI agents A practical look at evaluating agent outcomes, trajectories, graders, and production behavior.


OpenAI, Safety and alignment in an era of long-horizon models A useful case study in turning unexpected deployment behavior into new evaluations, safeguards, and monitoring.


Anthropic,  How we contain Claude across products Useful thinking on permission fatigue, containment, and designing boundaries within which agents can operate. Ethical AI Agents in Customer Journeys: A Practical Framework By WOMEN x AI Guest Blogger Pooja Kashyap, Conversational AI Evangelist at Conversive AI, this piece offers a practical framework for building ethical AI agents into customer journeys with fairness, transparency, human oversight, and trust at the center. How AI Agents Will Reshape the Future of Work By WOMEN x AI Guest Blogger Moha Shah, Venture Capital Leader & Innovation Operator, this piece explores how AI agents will reshape work in 2026 by automating execution, accelerating enterprise adoption, and shifting people toward judgment, orchestration, creativity, and strategy.

AI Disclosure: AI tools were used to support research and editing. The ideas, professional experiences, analysis, and final editorial decisions are the author’s.

 
 
bottom of page