Small business owner using a voice AI agent to run tasks on a laptop

ChatGPT and Claude Can Now Do Your Actual Work by Voice — Here’s What That Means for Small Business

Estimated read time: 6 minutes

For three years, talking to an AI meant asking it questions. This week, both of the companies that matter flipped that on its head. Now you talk, and the AI goes and does the thing — opens your apps, clicks the buttons, files the pull request, drafts the calendar invite. Voice stopped being a search box and started being a set of hands.

On Wednesday, Anthropic upgraded Claude’s voice mode so it can complete tasks inside apps you already live in — Gmail, Calendar, Slack, Notion, and Canva. On Thursday, OpenAI brought ChatGPT Voice to its desktop app, where it can now direct AI agents running in ChatGPT Work and Codex, look things up across websites and apps, and handle multi-step commands without you touching the keyboard. Two announcements, twenty-four hours apart, pointing at the same future. That’s not a coincidence. That’s a category being born in public.

Here’s what actually shipped, what it can do for a business with three employees instead of three thousand, and where it still falls on its face.

What OpenAI and Anthropic actually shipped

Let’s be precise, because the marketing blurs it.

OpenAI updated its ChatGPT desktop app to add ChatGPT Voice. The phone version that launched earlier this month was built for smoother conversation — better interruptions, more natural back-and-forth — but it wasn’t built to take action. The desktop version is. It’s powered by OpenAI’s ChatGPT-Live voice models, and it works with two things at once: ChatGPT Work (the agent that operates like a coworker) and Codex (the coding agent). It can also tap “computer use” skills to browse websites and poke around apps. On a Mac, a feature called Appshots lets it see what’s actually on your screen, alt-text included. In OpenAI’s own demo, a developer gave one spoken command — create a new thread, open a pull request, find the root cause of a bug — and the agent went off and did all three.

Anthropic updated Claude’s voice mode a day earlier, wiring it to its more capable Opus, Sonnet, and Haiku models so it can carry out tasks in Gmail, Calendar, Slack, Notion, and Canva. Less “write me an email” and more “reply to Dana, move our Thursday call to Friday, and drop the notes in Notion.”

The through-line: you’re no longer dictating text or asking trivia. You’re delegating a job, out loud, and an agent executes it while you keep your hands free.

Why “voice that does work” is different from Siri

If your gut reaction is “we’ve had voice assistants for a decade and they’re mostly good for setting timers,” that’s fair. It’s also why this is a bigger deal than it sounds.

The old voice assistants were command-matchers. You said one of a few thousand phrases they recognized, and they ran a pre-built response. Ask for anything off-script and you got “here’s what I found on the web.” They didn’t understand your intent so much as pattern-match your syllables.

What OpenAI and Anthropic shipped this week is different in three specific ways. First, it’s agentic — the voice layer sits on top of a model that can chain multiple steps together, so “book it, invoice them, and text me when it’s done” is one request, not three. Second, it works across your real apps instead of a walled garden, because the agent underneath already has access to your email, calendar, and workspace tools. Third, it can ask you back — when the agent hits a fork it isn’t sure about, it says so and waits, instead of guessing and quietly breaking something.

That last one matters more than the demos let on. The difference between a useful assistant and a dangerous one is whether it knows when to stop and check.

What this looks like for a small business

Strip away the developer demos — most small business owners are not filing pull requests — and the practical value shows up in the boring, repetitive work that eats your afternoons.

Think about the things you already do between other things: confirming a client appointment, moving a meeting, turning a voicemail into a follow-up email, dropping a to-do into your project tool, pulling three numbers out of a spreadsheet, cleaning up a Canva graphic before it goes out. None of these are hard. All of them require you to stop, open an app, and click through five screens. Voice-driven agents are aimed exactly at that gap — the tasks that are too small to delegate to a person but too frequent to keep doing by hand.

The solo consultant wraps a call and says, “Draft a recap email to the client, add the two action items to my Notion board, and put a 20-minute follow-up on my calendar next Tuesday.” The shop owner, hands literally full, says, “Text the 2 p.m. that we’re running fifteen minutes behind.” The one-person marketing team says, “Resize this graphic for Instagram and drop it in the Canva folder.” Each of these is a minute of clicking replaced by a sentence.

The compounding effect is the real story. A small business runs on the owner’s attention, and attention is the one input you can’t buy more of. If voice reliably claws back even thirty minutes of context-switching a day, that’s not a productivity hack — it’s a couple of extra hours a week pointed at work that actually grows the thing. That’s the same reason we keep coming back to AI as an implementation problem, not a model problem: the winner isn’t whoever has the smartest AI, it’s whoever wires it into the boring parts of the day.

The honest limits (read this part)

Now the cold water, because the gap between a launch demo and your Tuesday is wide.

It can be confidently wrong, at speed. An agent that takes action is an agent that can take the wrong action — email the wrong client, move the wrong meeting, overwrite the wrong file — and voice makes it faster to set that in motion. Text at least leaves a trail you skim before hitting send. Keep a human check on anything that leaves your business or touches money.

Screen access is a real trade. OpenAI’s Appshots lets ChatGPT see what’s on your Mac screen. That’s what makes it useful, and it’s also exactly what it sounds like. Before you turn that on, know what’s visible — client data, financials, other people’s private information — and decide whether you’re comfortable with an AI reading it. This is a genuine judgment call, not a checkbox to click past.

Voice is lousy in the real world. Half of small business happens in noisy rooms, on job sites, in cars, next to a customer. Dictating a complex multi-step command to an agent in a quiet office is one thing. Doing it over a blender is another. For a lot of owners, the phone-and-thumbs workflow will stay faster for a while.

It’s an early feature, and it acts like one. These rolled out this week. Expect misfires, confused hand-offs, and the occasional “I’ve completed that” for a thing it did not complete. Test on low-stakes work — internal notes, drafts, your own calendar — long before you let it touch a customer.

For a fuller picture of what today’s agents genuinely handle versus what’s still a demo, our plain-English rundown of AI agents is the companion to this piece.

Should you use it yet?

Cautiously, yes — as a hands-free layer over work you’d trust an eager new hire with on day one. Let it draft, summarize, reschedule your own calendar, and shuttle notes between your apps. Keep it away from anything irreversible, anything financial, and anything that reaches a client without your eyes on it first.

The bigger takeaway is the pattern, not the feature. When both OpenAI and Anthropic ship the same capability in the same week, they’re telling you where the interface is heading: away from typing prompts and toward handing off jobs. The owners who win with this won’t be the ones who talk to their AI the most. They’ll be the ones who figure out which repetitive thirty minutes to hand over — and which thirty they should never let go of. If you’re still deciding where AI fits in your operation at all, start with how to actually start and run a business with AI, then bolt voice on once the workflow underneath it is solid.

FAQ

Do I need a paid plan to use ChatGPT or Claude voice?

The most capable, agent-controlling versions generally live in the paid tiers, and OpenAI’s desktop voice ties into ChatGPT Work and Codex. Free plans get lighter voice features. Check your current plan, because both companies reshuffle tiers constantly.

Is this the same as the voice assistant on my phone?

No. Your phone’s built-in assistant mostly matches commands to pre-set actions. These new voice modes sit on top of full AI agents that can chain multiple steps and operate inside your actual apps. Different tool, different risk profile.

Can it really control my computer?

On desktop, to a degree — OpenAI’s version can direct agents, browse, and on macOS see your screen via Appshots. It’s real but early. Treat it as a capable assistant that still needs supervision, not autopilot.

Which one should a small business pick, ChatGPT or Claude?

Pick by where your work already lives. If you run on Gmail, Calendar, Slack, Notion, and Canva, Claude’s task integrations line up cleanly. If you’re deeper in the OpenAI ecosystem or do technical work, ChatGPT’s desktop voice plus Codex fits better. Most owners will settle on whichever they already pay for.

What’s the single biggest risk?

An agent taking a wrong action fast because you told it to out loud and didn’t review. Keep a human check on anything that sends, pays, publishes, or deletes.

Faceted Media Magazine covers business, AI, and entrepreneurship for the people building what’s next.