The Josh Bersin Company
The answer lives in this podcast

Answer extracted from the The Josh Bersin Company podcast — listen to the full episode below.

🎧 Listen to the episode on Listenly

How The Josh Bersin Company Restructured Research for AI to Read Accurately

The Josh Bersin Company realized their traditionally structured research—written like books with introductions, tables of contents, and conclusions—was fundamentally inefficient for AI systems. They built ARC (Agent Ready Corpus), a new architecture that extracts research content into machine-readable chunks and indexes it with eight dimensions of metadata tagging, allowing large language models to read and cite sources precisely without hallucinating or guessing.

Agent Ready Corpus (ARC) is a proprietary architecture developed by The Josh Bersin Company that restructures research data into a format optimized for AI consumption. Rather than delivering content in traditional document format, ARC breaks down nearly 1,800 case studies, maturity models, benchmarks, and frameworks into machine-readable segments, each tagged across eight distinct dimensions to enable large language models to understand context, retrieve accurate citations, and deliver precise answers without reliance on generative hallucination.

From Document Format to Machine-First Structure

The core problem was structural. Bersin's company had spent 28 years accumulating research across 95 functional capability areas in human capital management—everything from performance management to pay equity to employee wellbeing. Yet when large language models tried to extract value from this corpus, the traditional book-like format with narrative introductions and conclusions actually made it harder for AI to isolate and cite factual claims without inventing answers.

The solution was not to write more reports. As Bersin explains in the Jupiter Release episode, the company stripped away boilerplate entirely and rebuilt their entire knowledge base from scratch into what they call the Agent Ready Corpus.

Eight Dimensions of Tagging for Precision

The ARC architecture doesn't simply chunk text randomly. Each piece of research is tagged across eight different dimensions of metadata, creating a multi-layered index that lets LLMs understand not just what a fact is, but what context it applies to—industry, company size, geography, business challenge, and more. This granular tagging is what prevents hallucination.

The result is concrete: when a user queries Galileo (the Josh Bersin Company's AI-powered research platform) about performance management in retail companies in Germany, the system doesn't guess. It finds the exact relevant case studies and benchmarks from their corpus of nearly 1,800 real-world examples, pulled from approximately 120 countries, and returns them with proper citations. The system is reported to be 10 to 100 times more efficient at token usage compared to general-purpose LLMs querying the same research in unstructured format.

"We essentially built a new architecture that takes the essence of Galileo and carefully indexed it and tagged it in eight different dimensions so that the LLM can read it extremely well and does not have to hallucinate."

Josh Bersin — Founder and CEO, The Josh Bersin Company. Bersin has studied human capital practices and organizational management for nearly 30 years, developing maturity models, benchmarks, and nearly 1,800 case studies across 95 functional capability areas. His research serves over 2,000 companies seeking evidence-based solutions tailored to their specific industry, location, and company size.

The deeper insight Bersin explores in this episode reveals how 18 months of close collaboration with Microsoft shaped the optimization of this architecture specifically for Copilot integration—a detail that underscores the industrial maturity of the approach.

This shift represents a fundamental rethinking of how research institutions can serve AI-driven workflows. Rather than asking AI to understand human-authored content, The Josh Bersin Company engineered their content to be natively understandable by AI, without sacrificing the depth or specificity that made their research valuable in the first place.

See also

Why did the strategy shift from building a specialized AI agent to creating a portable intelligence layer across multiple platforms?

Customers indicated they already had AI platforms like ChatGPT, Gemini, Claude, or Microsoft Copilot and wanted Galileo integrated there rather than as a standalone agent.

What opportunity did the emergence of large language models like ChatGPT create for HR research organizations?

When ChatGPT came out, Bersin saw an opportunity to build a search engine where users could query an entire corpus of knowledge and get answers, then find supporting case studies and sources.

How has the business model for monetizing human capital research evolved over the past decade?

The Josh Bersin Company initially used a research membership model similar to Gartner, selling access to a library of reports, models, and frameworks. Over time, the model has evolved to integrate AI-powered search and portability across enterprise platforms.

Listen to the episode on Listenly