What ten months produced
- 1Nearly nine in ten of the company uses it every month, and the number is still rising. 94 of 108 people in July, up from 2 last October, with no month yet showing a decline.
- 2It has produced 13,497 finished documents. More than one per conversation. Four in ten are Word files, spreadsheets and decks that went on to be used as real work.
- 3It is doing work that could not otherwise have been done. A URAC accreditation submission, an ORSA model, a reserve method validated against six months of actual claim runout, a 21-state employee handbook.
- 4Total value created is approaching $1M a year: $254,000 in time recovered ($2,702 per active user), a comparable band of work that would not have happened otherwise, avoided spend, and roughly $300K a year of list-rate model compute delivered for under $10K a month.
Adoption climbed for ten straight months
The platform went in front of the whole company in October. Ten months later 87% of employees use it every month and adoption is still climbing. It has produced 13,497 finished work products across 10,472 conversations, which is more than one deliverable per conversation. The most valuable behaviour found in the record is that it pushes back: it withdrew its own reserving recommendation after testing showed it was $1.65M short, read the plan’s own policy wording more strictly than the team had, and declined a shortcut on member notices that the Texas delivery rule would not allow.
Messages per month
Ten months of growth with no sign of levelling off. July was the largest month, at 21,237 messages.
People using it each month
From 2 people to 94, out of 108 employees. Shown separately from volume so the two are not confused.
Five things worth your attention
People leave with finished work
13,497 documents against 10,472 conversations. Four in ten recent files are Office or PDF documents: 562 Word files, 570 spreadsheets, 285 decks, 241 PDFs. Staff are not asking questions and reading answers; they are walking away with the document they needed.
Ordinary staff are doing analyst work
The code sandbox is the most-used feature in the product, with 105,470 uses by 86 of 108 employees, nearly six times the size of the engineering team. SQL is the most common thing it writes, at 989 files: claims specialists, care coordinators and underwriters are pulling answers out of the data warehouse without knowing how to query it. Nine people outside engineering have built working internal tools.
It is already doing regulated work
The record shows work touching more than 60 named standards and rules: 15 URAC standards, the Texas insurance code, NAIC and ORSA and RBC, actuarial standards of practice, the No Surprises Act, mental health parity, ERISA and Form 5500, and HIPAA. Regulated member data passes through the platform every day, under the controls that work requires.
It corrects itself when the evidence changes
All 43,763 responses the platform has sent were read, and a sample of each behaviour graded by hand. Roughly 108 conversations contain a genuine self-correction, about three quarters of them driven by evidence rather than by the user pushing back. Outright refusal is rarer, at under 1% of conversations, but the verified instances are substantial: declining to pass an accreditation standard whose wording did not hold up, and flagging that a proposed shortcut on member notices would not satisfy the Texas delivery rule.
Adoption keeps spreading on its own
Engineering started it. Operations is now the heaviest user, at 4,740 messages and 22 active people in July, and still growing. Six of ten functions were using it within six weeks, and the first group of users is at 95% still active by month eight.
How many of the 108 employees used each capability
Breadth of use, not volume. The code sandbox reaching 86 people is the result worth noticing; the engineering team is 15. The provider directory, in orange, is Evry's own directory queried in plain English by a third of the company.
Seven cases, from the record
Every case is named and verifiable against the production record. The first three were read in full, start to finish.
1. Setting the IBNR reserve
CFO & Chief Actuary · 4 weeks · 40 messages
The platform first rebuilt the existing IBNR workbook and matched it to the dollar, which established that it understood the current method before proposing changes. It then tested three ways of selecting completion factors and checked each against six months of actual claim runout. The model it eventually recommended came within 0.3% of what actually happened. An earlier recommendation of its own was 54% low, and it raised that itself:
"The 7-of-9 method I recommended was $1.65M short."
It finished by drafting the year-end statutory memorandum with the relevant actuarial standards cited.
Why it matters. Statutory financial reporting, and the only hard accuracy measurement anywhere in ten months of usage.
2. URAC accreditation and TDI response
Quality & Population Health · 394 messages
Citations were written directly into the original policy documents at page and paragraph level, in the format URAC's submission system expects. It handled URAC's follow-up questions in the same session and pulled denial rates from Power BI for an attestation the Chief Medical Officer signed. Reviewing one standard ahead of submission, it read the plan's policy language more strictly than the team had and explained why the wording would not satisfy the reviewer, which is the kind of thing far cheaper to catch before a submission than after one.
Why it matters. Accreditation and regulator-facing work, plus evidence that it holds a position when the user wants a different answer.
3. Auditing 27 required member notices
Communications · 10 messages · one sitting
There was no single place recording whether all 27 required member notices were actually being delivered. The platform read the 2026 Member Handbook, all 110,000 words of it, plus the Welcome Packet, checked both against the list and produced a one-page Word audit of exactly where each notice is delivered, in a form the compliance team could act on. When offered a shortcut of linking to a notice rather than mailing it, it declined, and cited the Texas rule that makes linking a form of electronic delivery requiring the member’s consent first.
Why it matters. A complete delivery audit in ten minutes, against a manual review that would have taken days.
4. Candidate fraud screening, including a pattern nobody had seen before
People and talent · Apr–Jul 2026 · ~30 chats
Every engineering applicant now goes through an open-source check before anyone spends time on an interview: the résumé and the public profile are cross-referenced against each other and against the employers named, and the platform returns a verdict with its reasoning. Sessions run long, up to 168 messages, because the recruiters argue with the output and make it justify itself.
Three patterns show up, in rising order of difficulty: an invented persona with no verifiable footprint; a real work history with the titles rewritten; and one the team had not encountered before: a fake candidate fronted by the genuine profile of a real, unrelated engineer. Profile harvesting is the hardest of the three to catch because each signal on its own checks out.
Why it matters. The HR team rates this the platform's most successful use in their function, and it found a new fraud pattern rather than only applying a known checklist.
5. A claims quality system, built in ninety days
Claims QA · Apr–Jul 2026 · ~45 chats
A claims analyst who joined in April built the department's quality apparatus from nothing: a deduction-based 100-point weighted audit score tied to specific desktop-procedure steps, an Excel audit workbook with per-claim tabs, specialist and management dashboards, a result-code reference sheet, dollar-based review tiers, a random 20% sample selector, and a monthly error-trend report with root-cause write-ups. The department now audits roughly 600 claims a month against it.
Why it matters. Audit infrastructure of the kind normally bought as a consulting engagement, delivered by a new employee in their first ninety days.
6. Onboarding a trading partner's enrollment files
Engineering · Mar–Jun 2026 · 66 and 102 messages
Enrollment files arrive as X12 834 transactions and must match the companion guide before a group can go live. A developer uploaded the guide and a partner’s test file and had the platform validate it segment by segment. Four successive corrected files were re-validated with a diff-style "what got fixed, what is still open" report, and a second thread ran a live three-employer onboarding through the same channel.
Why it matters. File-format onboarding is the standing bottleneck between winning a group and enrolling it, and it normally waits on a specialist.
7. Reconciling two renewal rating methods
Underwriting · Jul 2026 · 24 messages
A renewal workbook carried two rating builds that diverged once credibility weight shifted toward the group's own experience. The platform traced both builds end to end and decomposed the entire 41.14 PMPM gap into three named drivers, showed the direction flip came from the two builds comparing against different manual rates, and re-derived the medical/pharmacy trend weighting against two published industry surveys.
Why it matters. Actuarial peer review of a live pricing model, at the speed of the renewal cycle, before the number goes to a broker.
What else gets built
A sample of recurring work across the company, all of it from the record.
| Function | What was produced | Scale or outcome |
|---|---|---|
| Utilization management | Medical-necessity write-ups mapping a clinical record to the named guideline, in SBAR form | ~197 review sessions in seven months from one nurse |
| Care coordination | Member letters rewritten to a sixth-grade reading level, and translated into Spanish | ~450 correspondence sessions in seven months |
| Customer support | Salesforce call notes generated from raw call transcripts against locked house templates | ~700 notes in five weeks from one representative |
| Support operations | A weekly schedule-adherence report built from a prose rule set and run against raw data | Replaced a manual four-hour weekly build |
| Claims | A 20-day new-hire training checklist with trainer rotation and per-day procedure references | ~20 training-material sessions |
| Claims | Desktop procedures authored and revised against the plan's own documents | ~35 authoring and lookup sessions |
| Finance | Form 5500 and Schedule A prepared for a group insurance client from live enrollment data | 58-message session, 14 distinct tools |
| Actuarial | The annual reinsurance report built tab by tab against the carrier's template | 240-message build |
| Underwriting | Prospect experience workbooks assembled from incumbent-carrier PDFs and emitted as Excel | ~250 workbooks in seven months from one analyst |
| Underwriting | A calendar-year versus plan-year accumulator study across the 36-group book | One 376-message session |
| Data engineering | A provider-file ingestion pipeline: 16 fixed-width source files into one 55-column master extract | 246-message build, repeats per network |
| Data engineering | A stored procedure assigning Major Diagnostic Category and DRG to the claims table | 120-message build |
| Sales | RFP questionnaires completed from prior approved answers under an explicit no-fabrication rule | ~65 RFP sessions in 8.5 months |
| Sales | A sales associate built and corrected their own reusable RFP-intake skill, then ran it as a one-word prompt | Hired mid-July, in use by day 10 |
| Account management | Renewal packages (loss ratio, membership, average age, service-area share) into a broker-ready email | 40+ packages in seven months from one AE |
| People | The employee handbook reconciled for a remote workforce across 21 states, with per-state accrual and leave rules | ~30 policies in roughly two weeks, an annual cycle |
| People | Annual headcount and compensation analysis, benchmarked by role against the market | 41 benchmarking sessions |
| Provider network | A department operating system: outreach tracker with owners, intake form, policy template and written SOP | Five-person team, about four weeks |
| Provider network | Proposed contract rates converted to percentage-of-Medicare against the local fee schedule, inline during negotiation | Sourced and enriched in the same session |
| Marketing | A brand and terminology style guide that assembled itself out of standing corrections during live work | ~214-turn thread over eleven weeks |
What it is worth
Three things that are easy to blur together are kept separate here. Only the first belongs in a payback calculation.
Hours displaced
Work already being done that now takes less time. The only line that belongs in a payback calculation.
~3,900 hrs · $254K
Capability that did not exist
Work that would never have happened otherwise: ORSA, URAC evidence, the 21-state handbook, the IBNR back-test. Genuinely valuable, but not a cost saving, and kept out of the payback figure.
~2,600 hrs · $250–500K
Compute delivered below market
July's 5.3B tokens price out at roughly $25K at list rates, uncached. The actual bill was under $10K. At the July run rate that is ~$300K a year of model compute delivered for about a third of its list price.
~$300K / yr
Total value created
The lines above are distinct and do not overlap: $254K displaced + $250–500K new capability + $87K avoided + ~$300K compute. For payback, use only the first line: $2,702 per active user per year, across 94 monthly users.
$0.9–1.1M / yr
Two automatic levers keep the running cost down as volume climbs: the router reserves the frontier tier for about 4% of requests, and prompt caching serves repeated context at a tenth of the input rate.
How it spread, and what it took
Messages per month by function
Engineering leads early and flattens after March. Operations starts slowest and finishes largest. Actuarial appears in January and steps up sharply in April when one person found a use that fit.
| Function | Nov 25 | Feb 26 | Apr 26 | Jul 26 | Users, Jul |
|---|---|---|---|---|---|
| Operations, claims, support | 264 | 516 | 1,248 | 4,740 | 22 |
| Clinical, UM, care | 258 | 1,532 | 2,086 | 4,033 | 13 |
| Engineering | 90 | 1,770 | 3,490 | 3,574 | 13 |
| Actuarial, finance | 0 | 392 | 2,228 | 2,475 | 4 |
| Sales, growth | 1,020 | 1,012 | 958 | 1,799 | 9 |
| Executive, operations | 176 | 826 | 720 | 1,245 | 5 |
| People, HR | 388 | 618 | 562 | 1,451 | 10 |
| Marketing, communications | 0 | 0 | 428 | 655 | 2 |
| Provider network | 88 | 76 | 110 | 546 | 7 |
Engineering starts it, operations inherits it
Engineering went first and peaked in March. Operations was the last big function to pick it up and is now well ahead of everyone else, still climbing. An organisation that starts with its engineering team should keep measuring past that first group; the return builds about two quarters later, in operations.
The late functions did the best work
Actuarial started three months after everyone else and marketing six months after. Both produced the strongest results in this study. Arriving late did not predict low value; those teams turned up with harder problems.
It reaches everyone quickly and it sticks
Six of ten functions were live within six weeks of launch, and the first group of users is 95% still active at month eight. The habit also changes shape as it settles: from 19.5 messages per conversation in the first week to 6.6 by month four, the signature of something becoming routine rather than novel.
Ten months in, it is still spreading sideways
July's growth came from operations, sales, the provider network team and staff who had only just been onboarded, while the two deepest workflows were close to flat. Sessions of ten turns or more grew faster than volume did, and retrieval from Evry's own systems roughly doubled in a month.
What the platform actually processed
Everything above counts conversations; this counts the work inside them. Between January and the end of July the platform processed 12.5 billion tokens, the equivalent of reading and writing several thousand full-length novels, all of it about claims, members, benefits, reserves and regulation. July alone was 42% of it, at 5.30 billion tokens: 49 times January and 104% above June.
Tokens processed per month
Input plus output, as reported by the model provider. Every month since January has been larger than the one before it.
Average tokens behind a single request
Requests roughly doubled between January and July; tokens rose 49 times. The difference is how much context each request now carries; one request today is a whole working session.
A growing share of it runs without anyone typing
Autopilot went live in June. In July it processed 1.36 billion tokens across 5,780 runs, roughly 26% of everything the platform did that month, with no one in the chat window. A chat assistant is bounded by how much people are willing to type. Delegated work is not, and in its second month it is already close to a third of everything processed.
Nobody at Evry picks a model
The platform has run 24 distinct models from four providers, across five generation changes in the two main tiers in nine months. Almost none of that was visible to the people using it: 99% of conversations run on auto, where a classifier sends the ordinary majority to the balanced tier and reserves the frontier tier for the genuinely hard requests.
Model generations by tier
Each bar is one model in production. The bars butt up against each other: in every case the successor's first request and the predecessor's last fall on the same day. Five same-day handovers, no migration project, no user retraining.
| Model | First seen | Last seen | Requests | Typical context |
|---|---|---|---|---|
| Sonnet 4.5 | 11-24 | 02-17 | 7,420 | 6k |
| Opus 4.5 | 12-15 | 02-14 | 77 | 16k |
| Haiku 4.5 | 12-18 | 07-31 | 16,321 | 32k |
| Sonnet 4.6 | 02-17 | 06-30 | 32,440 | 6k |
| Opus 4.6 | 02-18 | 04-16 | 449 | 8k |
| Opus 4.7 | 04-16 | 05-29 | 460 | 67k |
| Opus 4.8 | 05-29 | 07-24 | 2,206 | 11k |
| Sonnet 5 | 06-30 | 07-31 | 20,800 | 43k |
| Opus 5 | 07-24 | 07-31 | 79 | 133k |
Each generation absorbed roughly twice the context
Typical context went 25k on Sonnet 4.5, 50k on Sonnet 4.6, and 93k on Sonnet 5. That is the same people bringing bigger problems as the ceiling lifted. July, the first full month on Sonnet 5, processed more than the platform’s entire first five months of 2026 put together (5.30B against 4.58B).
The bill stays flat while volume doubles
July’s 5.30 billion tokens price out at roughly $25,000 at list rates; the actual bill was under $10,000. Three automatic levers: routing keeps the frontier tier to ~4% of requests, prompt caching serves repeated context at a tenth of the input rate, and ~40% of all tokens are the platform coordinating itself, exactly the traffic caching and routing compress hardest.
Where the findings are likely to carry
Accreditation and regulator evidence travels furthest
It is the only high-value workflow that runs entirely on documents a plan already holds: no member data, no integration work, no security review before anyone can see whether it works. Every other strong use case needs the organisation’s own data loaded first.
Utilization management is the depth, not the start
It is the highest-volume use internally and the deepest evidence that the platform works on clinical operations, but it only lands once the platform has already been trusted on something lower-stakes. Accreditation work comes first; clinical operations follow.
What proved hard to reproduce
- It is joined up to the systems of record. SharePoint, Outlook, Teams, GitHub, the claims warehouse and a purpose-built provider-network tool, all reachable in one conversation.
- The workflows are encoded, not remembered. Skills staff assembled themselves, like the reusable RFP intake and the claims quality apparatus, keep working without the person who wrote them.
- The routing layer absorbs model change. Five generation changes in nine months without a migration project. Nothing had to be re-platformed to stay current.
Capability is easy to demonstrate and hard to trust. What is rarer is a production transcript of a system withdrawing a $1.65M recommendation after its own back-test, declining to pass a standard the team wanted passed, or catching that a proposed shortcut would not meet a state insurance rule. In front of a risk committee that carries further than an accuracy benchmark, because it can be read end to end.
How the numbers were produced
Every figure comes from a full census of production records between 1 October 2025 and 31 July 2026, not a sample or a survey. Ten analysts each examined one department, nine flagship conversations were read in full, and four further passes reconciled the results. Both sides of every conversation were read: 10,472 conversations, 87,684 messages and all 43,763 responses the platform sent, around 27 million tokens of output. Behavioural rates such as self-correction were detected automatically, then a sample of each graded by hand and the rate adjusted to the measured precision. Time savings are modelled from stated assumptions about how long each task takes by hand; the reserving work is measured against six months of actual claim runout. Every script is committed and the whole analysis can be re-run. Supporting analysis is available under NDA.