Ep. 53: Using Deep Research and Getting the Robots to Do Your Homework
After years of not knowing what to do with it, I finally decided to give deep research mode a try. In this episode I break down what deep research actually is, how it works under the hood, and how ChatGPT, Gemini, Claude, and Grok stack up against each other. I also walk through real use cases, prompting tips for getting a better report, and a safety reminder about hallucinations and bad citations.
Deep Research: Finally Giving It a Try
Today we’re talking about something that’s been around for a bit, but that I’d yet to actually try until last week: deep research.
This is literally a way to have the robots do your homework for you, and this is what I’m calling a 100% Maestro-approved use of AI.
I don’t like having AI speak for me, or write for me, or think for me. But I’m definitely cool with having AI do the searching for me, and that’s what we’re diving into today.
So What Is Deep Research, Exactly?
Deep Research is an AI mode that autonomously searches, reads, and synthesizes dozens to hundreds of web sources over several minutes to produce a structured, cited report on a topic.
Deep research is not new, and Google actually lead the charge with its release:
- Google (Gemini) — Launched December 2024
- OpenAI (ChatGPT) — Launched early February 2025
- Anthropic (Claude) — Launched April 2025, as “Research,” alongside Google Workspace integration; expanded and hit mobile in May 2025
- Perplexity — Launched 2025
- xAI (Grok DeepSearch) — Launched 2025
Deep research is a mode you select or turn on, and the output differs significantly from a standard chatbot reply in that the model doesn’t answer from what it already “knows.” Instead it searches the web, reads actual sources, cites them, and generates a report for you.
The model will run for roughly 5 to 30 minutes depending on the task (no need for you to sit there watching it if you don’t want to), and the final output is a structured report, usually a few pages long, that includes citations.
How It Actually Works
Under the hood:
- A lead agent reads your question and writes a research plan. (See last week’s episode for more on agents)
- It breaks that plan into chunks and hands each chunk to a little sub-agent that goes and searches its slice of the topic.
- Because the agents are searching all of those sources at once (dozens to hundreds of sources), it takes minutes instead of seconds like a standard chatbot exchange.
- Those sub-agents come back with findings plus the sources they pulled from.
- A synthesizer stitches it all into one report, and a citation pass makes sure claims point back to where they actually came from.
Better Input = Better Output
As per always with AI, the better the input you give it, the better the output it will generate for you.
A good request says what you want, what to include, what sources to prioritize, and what format.
Bad: “Tell me about AI note-taking tools.”
Better: “Compare three AI note-taking tools for solo creators. Include pricing, privacy, transcription quality. Prioritize vendor docs and reputable reviews. Give me a table.”
Bonus tactic: Some tools let you restrict the search to trusted sites, or point it at your own files, so it isn’t just grabbing whatever’s floating around on the open web.
Example Use Cases
Like many things with AI, deep research feels a bit like “here’s a solution, go find a problem,” particularly for folks who don’t have jobs that involve researching things.
I highlight this because on the other end of the spectrum we have folks like my friend Khe, who works in finance. He’s been using deep research since it was first released in 2024 because so much of his work involves researching things like company information, numbers, and publicly available stats. It’s a great use case and he’s noted time and time again how much time it’s saved him.
For folks like us (I’m very much assuming you and I are in similar boats), there are less obvious use cases. Admittedly, I only used it so I’d have some experience with it for this episode, but here are some ideas of things you could use it for:
- Vetting a tool or software before you buy, the real comparison with pricing and privacy, not a listicle
- Scoping a new offer or program by researching what competitors charge and how they position
- Pre-call or pre-pitch prep, a briefing on a person, brand, or company before a sales call or podcast interview
- Figuring out a new platform or channel (“is Substack worth it for someone like me,” “what’s actually working on YouTube in my niche”)
- Finding the questions people actually ask about their niche
- Big purchases, the car, the stroller, the mattress, the thing you’re overthinking
- Health insurance and plan comparisons
The throughline for folks in less research-heavy industries is that we’re most likely going to use deep research to produce a report that saves us hours of tab hopping and creates a brief we can actually act on, not some 40-page analysis that we’re going to present to a boss that will never even read it.
The Tool Landscape
Every major AI has a version of deep research, and each has a slightly different flavor:
- ChatGPT: the comprehensive one, longest reports, most sources, slower.
- Gemini: fast, broad because it’s plugged into Google Search, exports straight to Google Docs, has a real free tier.
- Claude: fewer sources but strong at reasoning through sources that contradict each other, cleaner writing.
- Grok: pulls live from X.
A heads up: running deep research eats way more of your usage than a normal chat, so use it wisely.
How to Actually Use Research
Depending on the AI tool you’re using, there will be a mode selector somewhere, usually in or near the message box or prompt field. Select Research (or Deep Research, or whatever it’s called on that tool) and prompt as normal.
As per always, I recommend having AI create the actual prompt for you. Don’t be surprised if it’s very comprehensive. Additionally, some models will show you a research plan before it starts.
If you’re able to choose, I suggest using the strongest reasoning model offered by that LLM. (I used Opus 4.8 for my search; Fable felt like overkill). Remember, you’re asking it to plan a multi-step job, judge which sources are worth trusting, and synthesize conflicting points into something coherent. Given those steps, hopefully you can see why this mode is more token intensive than a regular chat.
Again, the model will run anywhere from a few minutes up to 30-plus minutes depending on the task assigned. You absolutely do not need to be there to watch it or nudge it along, and it will produce a comprehensive report when it’s done.
Safety Reminder
The model can still hallucinate (aka make things up). It can cite a bad source confidently. It can and will miss paywalled or gated content that’s actually important.
Read the citations and verify the the things.
How I Used AI This Week
Each episode I share a quick example of how I used AI that week.
This week I used Claude in research mode to highlight content gaps in the AI learning space and suggest episode topics for Prompting Curiosity.
My review: The report was honestly very good. Claude did its job, cited solid sources, and organized everything cleanly.
The asterisk is that I didn’t love the results, not because they were bad, but because I don’t want to make episodes about the topics it surfaced (100% a me thing).
Da Wrap-up
Despite not loving the results it generated in the report I asked for, I do feel that deep research mode is a really helpful tool, and when a use case arises, I will gladly let the robots do my homework for me.
As always, endlessly grateful for you and your curiosity.
Catch you next Thursday.
Maestro out.
