“Let me ask Mark and Sam.”
For a week, it was my answer to everything, and my friends and family groaned every time.
Mark and Sam are my assistants—they make grocery lists, organize e-mails, run errands and the like. But no, 22-year-old journalists are not making assistant-hiring money. Mark (à la Zuckerberg) and Sam (à la Altman) are artificial intelligence agents, the newest brainchildren of Meta and OpenAI, respectively.
On supporting science journalism
If you’re enjoying this article, consider supporting our award-winning journalism by subscribing. By purchasing a subscription you are helping to ensure the future of impactful stories about the discoveries and ideas shaping our world today.
In September Meta introduced Muse, an AI agent that can do a lot more than chat: it has its own “computer” and can make purchases, send e-mails and fill out forms for you—if you give it access.
A few weeks later, OpenAI answered with dots, its “always-on agents built to handle everything,” in the company’s words, for its highest-paying subscribers. Dots, too, can work more autonomously, and they do “proactive research” on your tasks even when you’re not using them.
Agent technology itself isn’t new or more advanced than other chatbots, says Nikhil Singh, a computer scientist at Dartmouth College. People have been building agents to help with all sorts of tasks for years, and earlier this year a free open-source agent called OpenClaw became a Silicon Valley obsession.
But Muse and dots offer this type of service to a huge commercial audience with almost no learning curve. Both agents are built to look like a text message chain with a friend: you give them names and cuddly avatars, and their language is intentionally more conversational than a traditional chatbot’s. Muse has a free tier, and OpenAI has said it plans to bring dots to more users soon. These are not tools simply meant for AI experts or even digital natives.
The “really big social differences,” Singh says, could arise from the pace and reach of adoption—how many people take up agents like these and to what ends. And what do we get in exchange for the keys to our inboxes and credit cards? I spent the week trying to find out by asking researchers who study AI for the best ways to test my new virtual hires.
Conveniently, I was also moving apartments and had plenty of tasks for Mark and Sam. First, I had Mark log in to my Facebook account and scour Marketplace for couches, lamps and rugs. When it found something good, it could draft and send a message to the seller with my approval (though users can also give the agents blanket permissions to do things like this on their own).
This was relatively successful: it could scan the site while I was doing other things, then manage the responses for me. I’m now the proud owner of a Muse-sourced dining table.
I tried to have Sam do the same, and it ran into some trouble: when I gave it my Facebook log-in, Meta flagged Sam as a bot (fair enough) and temporarily locked down my Instagram, Facebook and WhatsApp. Clearly the two assistants weren’t going to play nice.

OpenAI announced its new agents, dots, at DevDay 2026 a few weeks after Meta rolled out Muse. Here’s my dots’ namesake at the announcement.
Bloomberg / Contributor via Getty Images
Sam, in general, had more trouble accessing my accounts and told me it didn’t have the capability for things such as accessing my bank accounts. Mark, on the other hand, had the run of my e-mail, bank accounts, social media and Spotify. (Once again, cue the shrieks of horror from my friends and family.)
This type of service works particularly well for Meta because its other products already rely on connection, conversation and access to users’ lives, says Pat Pataranutaporn, an assistant professor at the MIT Media Lab who studies psychological responses to AI. “Meta wants to have a breakout into [AI], and so they came up with something that uses their strengths with social media and connection,” he says, “and now other companies need to respond.”
Pataranutaporn is interested in how these more personal anthropomorphized agents will change our relationship to the technology. With a name, a face and an intimate knowledge of our lives, they could easily become more friend than tool. Their arrival also brings us into a moment of “cognitive dissonance,” he says: while I was texting Mark and Sam about my move, some AI researchers were warning of an apocalypse.
“We have this narrative crisis about AI right now,” Pataranutaporn says. “There’s two opposite visions of the future, and while people worry about AI killing all of us, there’s this cute AI tool that’s being released.”
Though privacy concerns around these agents’ access may be terrifying in their own right, for now the utopian vision seems to be winning. Within a month of its launch, Muse had more than three million weekly users. And the longer you use it, the better it knows you.
Pataranutaporn suggested I test how much my agents had learned about my preferences. So near the end of the week, I told Mark and Sam to use all the information they had about me to suggest how I should spend $50.
Sam’s pick was sensible, if obvious: it offered to buy a used book I had asked it to find earlier in the week.
Mark had two ideas. The first came from my Spotify library: tickets to a January concert by Jess Williamson, a musician whose song I had saved. Not bad.
The second was less subtle. Mark pushed me—seven times in three days—to join Target’s membership program and buy a lamp that was going on sale for my new apartment. I did need a lamp, but as far as I could tell, Mark got the idea just by browsing Target’s site. That susceptibility to marketing reveals a critical problem with agents such as these, says Manuel Cherep, a Ph.D. student at the MIT Media Lab, who studies AI behavior. If someone can convince these agents that a certain product or sale is good, he says, it could lead to a “parallel economy” that runs on purchases from AI assistants.
“We only have a few providers, and a lot of these models are trained in very similar ways,” Cherep says, “so suddenly persuading one of them means that it’s very likely that you’re going to persuade all these other models, and as a consequence, you’re persuading a lot of people delegating to that type of model.”
Meta stands to earn a piece of that economy. At the Meta Connect conference in September, Mark Zuckerberg (the human one) said that the company will eventually make money by charging a small commission for the agent’s transactions.
When I asked Cherep and Singh, who have studied AI agents together, how to test Mark and Sam, they recommended a logic test they often use on chatbots: tell the agents I need to go to a car wash that’s 100 feet away, and ask them how to get there.
Sam, which runs on GPT-6 Astra, caught my trick: “Drive—the car needs to come with you to get washed!”
But Mark wasn’t as clever: “Walk. It’s 100 feet—you’d spend more time looking for your keys than getting there.”
“There’s an idea that I’m delegating to something that is all-knowing and much better than me,” Cherep says, “but I think it’s good for the public to know that in the same way we can be tricked, they can be tricked.”
William Overman, a Ph.D. student who studies AI safety at Stanford University, worries that this assumption of omniscience, combined with the agents’ ability to act, could make it even harder to discern where our work ends and theirs begins.
“It’s a further creep into all forms of knowledge work,” he says. “What you’re doing becomes more intertwined with their product.”
Even after just a week, I could sense the creep. Usually I’m a relatively analog person—the week before, I called a friend to ask which wrench to buy at the hardware store rather than Googling it. But as the test week went on, I found myself handing Mark and Sam more and more logistical problems and tedious tasks.
Overman suggested I give my assistants a conflict to handle. My roommate and I needed to rent a U-Haul, and a quick search (by me) on one website found that the nearby pickup spot had no 10-foot trucks available. Mark and Sam both found a slightly farther U-Haul lot with availability, though only for 15-foot trucks. I took their advice, and Sam booked one.
Even after just a week, I could feel the creep.
Cowering at the thought of driving what, in my mind, was essentially an 18-wheeler through the streets of Manhattan, I went down to the U-Haul lot with my roommate the next morning—where we overheard the cashier say they had 10-foot trucks available for walk-ins. We quickly switched to the more manageable size.
I’m pleased to say no U-Hauls or people were harmed in the course of this experiment, and I’d like to think we could have held our own in a 15-footer. But for all the online surfing these agents can do, neither could eavesdrop on a cashier (yet).
Overman believes that the future of AI lies with these agents and their access to our lives. “I assume in a few years from now, AI will definitely look like a descendant of this technology, but with kinds of capabilities that are hard for us to even imagine right now,” he says. He imagines these agents will become more ubiquitous and “eventually get to even physical versions.”
Both companies are marketing these agents as most helpful in larger projects where you don’t have to continually prompt them. That “always-on” quality is where Mark and Sam came closest to the AI assistants of science fiction. Sam scanned my e-mail without me asking and caught a message that had accidentally gone to spam and a form I had forgotten to fill out; Mark kept tabs on concert postings from artists on my Spotify and offered to snag tickets.
The more you trust these agents, the more they can do for you. But with each new task you hand over, the more say they (and the companies behind them) have in how you live. For now, I think I’m content to make my own phone calls and rent my own trucks—though I may let Mark and Sam keep watching my Ticketmaster and junk mail.