WRITELOOP

CREATING MY OWN DEEP RESEARCH SOLUTION

I wasn’t fully satisfied with the Deep Research solutions I found in the market, and then I decided to leverage my own one using Hermes Agent (that I use on an isolated VM on dedicated hardware as my personal assistant). This has helped me to use AI for research not as a passive consumer, but as someone that can tweak the results to reach the standards I value and still preserve my critical thinking when reading the reports it produces. I have iterated a lot in the last few months to a point that now I am feeling more satisfied with its results. Here is how I have achieved that.

2026 October 1

The Dillemma

The deep research tools I tried never quite gave me what I wanted. They produced polished summaries, but when I asked “how do you know this?” the answer was nowhere in the report, so I set out to build a system whose work I could poke at, trace, and argue with instead of a summary I had to take on faith.

That system is a deep research skill I built on Hermes, the AI agent I run in an isolated VM on dedicated hardware as my personal assistant. The first versions worked, but they still did not meet the standard I care about, so I kept iterating for the last few months, and only recently has the output reached the point where I genuinely like the reports it hands me (I have even shared some with friends to get some extra opinions). This post is how it works, told the way I came to understand it, one mechanism at a time, and it ends with the checks you can run yourself on any finished report.

Which was the goal?

The goal was simple to state: I should never have to take the report’s word for anything. Facts carry their sources, inferences are labeled as inferences, and a score tells me how solid the whole base is. Stated that way it is only a promise, but everything below exists to keep it.

The features I needed

  • It builds an HTML artifact from markdown and JSON sources, so the result presents beautifully when I need to share it with other people;

  • It marks which parts of the report are facts, cited with their sources, and which are inference, meaning what the model worked out from those facts, then shows the ratio between them in a nice HTML widget;

  • It produces a “Confidence Score” that grades the evidence behind the report, also shown as a nice HTML widget;

  • I can use any combination of models for the research and for the contester, which refines the research results one step further.

What does this look like?

A list of features proves nothing on its own, so here is a detailed view of each element of the final HTML artifact, with some pictures to give a broader idea of what can be achieved:

1) Left Panel and Widgets

Title, header, table of contents, methodology, executive summary, plus the widgets for fact mix, confidence, report info and sources. The screenshot below comes from a real example:

Section 1

2) Findings

Section 2

3) Post-Findings

Recommendation, Contradiction and debates, the confidence score in detail, and the knowledge gaps:

Section 3

4) References

A list of all the references used to produce the report. Each one is numbered, contains a clickable link so I can verify the facts, has the date it was published and updated, a score (showing how much it contributed to the report), and the verification (if it was corroborated, single-source, etc…).

Section 4

The pictures show what the report looks like, but they say nothing about how each panel earns its claims, so here is the procedure behind them.

What is Hermes Agent and what does it have to do with all that?

Hermes is the AI agent I run on a side machine. I give it a job in plain language and it uses a skill, which is just a documented procedure, to finish it. A skill in this setup is exactly that: a written procedure that Hermes reads and executes the same way every time.

Hermes runs the deep research following a pipeline and ends with one report in 3 artifacts: a markdown document, a machine readable JSON file, and a web page meant for human consumption. The skill contains the instructions, but is has also some scripts to add some determinism. E.g. an script does some automated checks to confirm those 3 artifacts are in sync before the report is handed to me.

How a research run works

Planning. Hermes breaks the question into three to seven smaller ones: what the established facts are, where sources disagree, what the timeline looks like, what evidence actually exists. Then it writes a search plan of five to ten queries, phrased differently on purpose (academic, news, official announcements) so it does not only meet sources of a single kind.

First round: wide, and read in full. Several searches run at once and collect 15 to 50 candidate pages, from which Hermes picks the 8 to 15 most promising and reads each one in full. For every page it records the publication date and the last update date, because a number from 2019 and the same number from last month are not the same fact.

Second round: close the gaps. What the first round taught it decides the second: two to five targeted searches that fill the holes, chase a claim back to whoever made it first, and look specifically for evidence contradicting what was found. Contradictory evidence is kept and reported, never quietly dropped.

Scoring every source. Each source gets a relevance score from 0 to 100, which is a grade of how useful it was to this report. Five factors decide it, in order: how directly it addresses the claim, the authority of the publisher, how recent it is, how far independent sources agree with it, and how deeply it covers the topic. The ranges:

  • 90 to 100: answers the core question directly, usually from a primary source;
  • 70 to 89: strongly relevant and credible;
  • 50 to 69: moderate context;
  • 25 to 49: marginal or dated, listed anyway for transparency;
  • 0 to 24: weak, kept with a caveat instead of hidden.

Writing up. Findings get grouped into what answers the question, what supports it, where sources disagree, background, and what stays unknown. Every report then has the same parts: an executive summary, tagged key findings, a contradictions section, knowledge gaps, and a reference list showing each source’s score and verification status.

The independent fact check

The stage that changed how much I trust the output comes from the independent fact check. Every central claim goes to a contester (reviewer) that begins from zero: a fresh AI instance with no memory of the searches and no sight of the pages this research collected, given one instruction, do not accept this claim, try to prove it wrong. Two rules make that review independent.

Source independence. A source counts as independent only if its origin does not trace back to a source already supporting the same claim. Five sites retelling one press release count as one source instead of five.

Model independence. The reviewer runs on a different AI model than the one that did the research. Which model reviewed is read from the Hermes configuration and written into the methodology of every report, so you can check it. If the two ever share a model, the report says so and flags the extra risk of correlated mistakes. For example, GLM 5.3 Flash makes the research, then DeepSeek 4.1 Flash acts as the reviewer.

Each claim comes back with one of the verdicts below:

  • Corroborated: at least two independent sources agree.
  • Single-source: exactly one independent source exists.
  • Contradicted: independent sources disagree, so the report shows both positions.
  • Not found: a targeted search located no independent source at all.

Where the reviewer and the research disagree, both readings land in the contradictions section.

The fact and inference tags

Every claim in the report is tagged before delivery, and there are only two tags:

  • FACT: Hermes read it directly in a source, a quotation, an official document, a data file, or a number it measured live while doing the research.
  • INFERENCE: Hermes concluded it, a reading, an implication, a synthesis of several facts. When it cannot decide which a statement is, it tags it INFERENCE.

Three rules about the taggging:

  • A tag sits in front of its claim, in its own paragraph, so one paragraph never hides two differently tagged claims.
  • The whole report’s split is counted (facts VS inference), and that count shows as a widget on the web page next to the confidence score and inside the JSON file.
  • A script does automated check and compares the tags in the text, the counts in the data file, and the dial, and if any of them disagree the report fails and is not delivered.

For you as a reader, the split is simple: a FACT is something the AI saw, an INFERENCE is something the AI worked out. A section heavy in inferences is exactly where your own judgment earns its keep.

The confidence score

The tags show what each claim is made of, but they say nothing about how solid the whole base under it is. So every report states a confidence score on a 1 to 10 scale, written as “Score: 7 out of 10, yellow tier”. I don’t use percentages because they would pretend to a precision the evidence does not have. The three ranges:

  • 1 to 5, red: evidence gaps or a contested claim the report leans on. Read the report as a hypothesis, or as a recommendation to proceed with caution.
  • 6 to 7, yellow: the direction is corroborated by independent sources, but details may still shift.
  • 8 to 10, green: corroborated across independent sources and verified live, where the open gaps are non structural, meaning they do not touch the conclusions.

Six inputs make up the number, and all of them are visible in the report:

  1. How broadly the central claims are corroborated;
  2. The share of primary versus secondary sources;
  3. How much was verified live, meaning Hermes checked the current source during the research instead of an article about it;
  4. The reviewer’s verdicts;
  5. How many claims were contradicted;
  6. And how many knowledge gaps remain open.

The score has a legend of the ranges in plain language and two boundary paragraphs:

  • why the number is not one point lower: what would justify a downgrade and why that condition does not apply;
  • why it is not one point higher: what the next level demands and what specifically blocks it.

The automated check refuses to publish the web page if the score is missing its legend or either boundary paragraph. That is why I consider the audit trail for the number itself: if you think the score is wrong or insufficient, the two paragraphs tell you exactly what evidence would improve the report.

What this gives you as a reader

The three mechanisms together are the closest thing I have to a readable “thought process” of an AI: The fact and inference tags show what it saw versus what it concluded, the verification tags show which conclusions were checked by a second model and which rest on one source, and the score shows how solid the whole base is.

So, you do not have to trust the report as a package. You can focus where the report itself admits weakness, the Single-source and Contradicted findings and the paragraphs heavy with inferences, while treating the green, corroborated parts as already tested.

How can you verify the report yourself

That is what helped me bring confidence on it during the many iterations I did to reach this point: open any finished report and confirm it has all five parts: the methodology that says how the research was done, the tagged key findings, the confidence score with both boundary paragraphs, the contradictions section, and the reference list with scores and verification badges. Then take one SINGLE-SOURCE finding and do a manual search for one independent second source on that claim yourself, and take one INFERENCE paragraph and ask yourself whether the facts underneath it actually support the conclusion you just read.

Important notes

  • The score grades evidence, not conclusions. Seven out of ten does not mean the conclusions are 70 percent likely. It describes the state of the evidence and nothing more.
  • Single-source stays single-source. A SINGLE-SOURCE finding can still be correct, and it still deserves one more search before you repeat it to someone else.
  • Independent review reduces shared blind spots; it does not remove them. Two different models can make the same mistake. What is left unknown is listed under knowledge gaps.
  • If the independent check cannot run, the report says so. The methodology records the fallback rather than pretending a review happened.
  • Dates expire. A score reflects the evidence on the day the research ran; anything that moved since then could change the research applicability.

Complete Report Picture

Here is a complete picture of the report:

Complete Report Picture