Literature review 2023 - 2026 Vision + LLMs in service work Workshop preparation
Capturing and Scaling Tribal Knowledge with Video Intelligence

Seeing Is Not Knowing

Three years of academic and business literature on multimodal AI in service delivery, read against a single question: what is actually scarce? The finding is consistent across benchmarks, field experiments, and deployments - the camera stopped being the constraint around 2023. The constraint that replaced it is the organization's ability to say what the camera is looking at, in a corpus it owns and can edit at field speed.

Prepared for
Smarter Services Executive Symposium
Sept 14-16, 2026 - co-hosted workshop with SightCall
Workshop subject
Class E in the taxonomy below - capture and reuse of tribal knowledge from video
Window
Publications and deployments, 2023 - Sept 2026
Verticals
Industrial & heavy equipment · Insurance & claims · Telecom, utilities & home services · Regulated service
Stance
Written through the annotation-layer argument, not neutral - marked passages are position, not finding
1 - The premise

Everything got easier except the part that was hard

In 2023 a technician could not point a phone at a hydraulic manifold and get a useful sentence back. By 2025 they could. That change is real, it is well documented, and it is not where the difficulty lives.

What a general-purpose vision-language model does now - read a data plate, describe a leak, name a component, follow a photographed procedure - was research-grade three years ago and is table stakes today. Every vendor in the field service software market has shipped some version of it. The multimodal turn that MIT Technology Review flagged in mid-2024 arrived on schedule.

But the record of the last three years divides sharply, and it divides along a line that has nothing to do with model quality. Capabilities where the set of possible answers is small, enumerated, and owned by the deploying organization have scaled to millions of transactions and are producing measurable financial results. Capabilities where the answer space is open - or where the reference corpus belongs to a different team, a different vendor, or a different revision cycle - remain stuck at pilot, and the peer-reviewed evidence on why is unusually blunt.

The insurance industry did not win at visual AI because insurers had better models. They won because a repairable vehicle has a finite parts and operations catalog that the carrier and its estimating platform jointly control, and because a damage estimate is a number that a business already knew how to act on. Field service, for the most part, has neither condition.

Annotation - the thesis this review is written through

Visual AI's real product in service is not the diagnosis. It is that, for the first time, capturing what a technician saw and did costs the technician nothing. The diagnosis is the demo. The capture is the asset - and only if it lands somewhere you own, control, author, edit, and reuse for your own purpose.

A note on where this workshop sits in the literature

The workshop's subject - capturing and scaling tribal knowledge with video intelligence - maps onto exactly one class in the taxonomy that follows, and it is the class with no published measured outcome anywhere in the three-year record. That is not a warning. Every vendor in the category now sells this capability; not one has published a number for it, and no peer-reviewed study has measured it in a service setting. You are not late to something. You are early to something nobody has instrumented.

It is also worth separating the two verbs in the title, because they fail differently. Capturing is now cheap and largely solved - the enabling change of the last three years. Scaling is not a bigger version of capturing; it is a distribution, adjudication and maintenance problem that the capture tools do not address and mostly do not acknowledge. The pattern underneath and The Adjudication Layer take that apart.

2 - Capability taxonomy

Six classes, three of which work

Grouping the literature by what the system is actually asked to do - rather than by vendor category or by model - produces six classes with sharply different evidence positions. Maturity below is evidence maturity: how well supported the capability is by measured outcomes outside the vendor's own marketing, not how many products claim it.

Class What the system is asked to do Where the evidence stands Evidence maturity
AIdentification
& capture
Read what is physically in front of the camera - serial plates, model codes, error displays, gauge faces, part labels - and match it to an asset record. Solved. Bounded output space, ground truth already in the deploying organization's own systems, immediate error correction. This is the capability quietly carrying most of the ROI claims in service software today.
Production
BBounded
classification
Decide whether what is pictured falls into one of a fixed, owned catalog of conditions - damaged / not, worn to spec, this defect type, this severity. Production at very large scale in insurance and grid inspection, where the catalog is finite and the deployer owns it. Degrades quickly as the catalog opens up or crosses organizational boundaries.
Production, bounded
CHuman-in-loop
visual triage
Put a live camera between a customer or technician and a remote expert; use AI for capture, on-screen annotation, and session summarization while a human holds the decision. The strongest deployment record in service, because the AI is not making the call. Best published outcomes come from dispatch-avoidance settings where the decision is effectively binary. Nearly all figures are vendor-reported.
Scaled, vendor-measured
DVisual document
retrieval
Treat the page image - schematic, exploded view, troubleshooting flowchart - as the retrievable unit, rather than the text extracted from it. Genuinely new and genuinely better: 20-40% end-to-end gain over text-parsed pipelines in the ICLR'25 work. Almost no field deployment evidence yet. The most under-exploited capability in this list.
Proven in lab
ESession-to-
record
Turn captured video and photos into a structured work record, a knowledge article, an annotation on a procedure, a correction to a manual. Every major vendor now claims it; nobody has published a measured outcome. The highest-leverage class in the taxonomy and the least evidenced - and the subject of this workshop. Capture is solved; distribution, adjudication and decay are not.
Claimed, unmeasured
FOpen-ended
diagnosis
Point a camera at an unfamiliar machine in an unfamiliar state and receive a diagnosis and repair path. Not close. Benchmark ceilings sit well below industrial tolerance, and the failure modes are not the ones a service organization can supervise - see The ceiling. Demos of this class are demos of class A or C with narration.
Research
Annotation

Read the maturity column and the ownership column together and they are the same column. A. Identification & capture B. Bounded classification and C. Human-in-loop visual triage work because the deploying organization owns the answer key. D. Visual document retrieval and E. Session-to- record stall not because the models are weak but because the corpus they would read from and write to belongs to somebody else - an OEM, an engineering group, a documentation team on an eighteen-month revision cycle. F. Open-ended diagnosis fails on both counts at once.

3 - Deployment record

What has actually shipped, and who is counting

The table below is the measured deployment record across the four verticals, with provenance marked. This distinction matters more than usual here: the visual-AI category is heavily vendor-published, and almost none of the field service figures in circulation have an independent methodology behind them. Where a number is vendor-reported, it is not therefore wrong - it is unaudited, and should be used as an existence proof rather than a benchmark.

Vertical Deployment Reported outcome Provenance
Insurance
& claims
CCC Intelligent Solutions - computer-vision damage analysis across the US auto claims networkClass A + B 14M+ unique claims processed through 2022, roughly 3× the pre-pandemic volume; 100+ of 150+ carriers applying AI somewhere in the claim lifecycle; 28% of repairable claims photo-initiated. Later products extend to impact-severity prediction from photos. Vendor-reportedCCC, Feb 2023 & Oct 2023. Volume figures are auditable in principle; accuracy is not published.
Insurance
& claims
Tractable - AI appraisal feeding straight-through processing; North American availability via Mitchell; carrier partnerships incl. The HartfordClass B Straight-through settlement of qualifying claims in minutes rather than days. No independently published accuracy or dispute-rate data across the window. Vendor-reportedTractable press releases
Telecom &
home services
TechSee at ADT - visual assistance for security-system support and installationClass C 2.6M+ virtual sessions; truck-roll avoidance above 75% for sessions using visual assistance. The single most-cited dispatch-avoidance result in the category. Vendor-reportedTechSee, early 2025. Avoidance denominator is not defined publicly.
Telecom &
contact centre
SightCall VISION at HELPLINE - Smart OCR serial/device capture ahead of support contact, 17 contact-centre sites, 200+ clientsClass A Removes verbal serial-number exchange and repetitive diagnostic questioning; reported as "several minutes or more" per interaction. No quantified outcome published. Vendor-reportedSightCall case study, Mar 2025
Utilities Buzz Solutions PowerGUARD; AiDash grid inspection platform - drone, helicopter and fixed-camera imagery for asset and vegetation conditionClass B Replaces "months of manual image review"; detects corrosion, broken components, vegetation encroachment, wildlife intrusion. Accuracy figures are not published by either vendor. Vendor-reportedNVIDIA customer story, Mar 2025; AiDash, Oct 2024
Industrial &
heavy equipment
John Deere Operations Center ProService - natural-language search across the full operator, technical and repair manual set for a specific machineClass D, text-first Deere reports step-by-step guidance returned in 15-30 seconds against a machine-specific manual library. Roughly 75% of early sales were at the $195 entry tier; the renewal rate after trial is not disclosed. Vendor-reportedAnnounced Sept 2026; platform launched July 2025
Industrial &
manufacturing
Squint - capture and guidance of plant-floor procedures from what the worker's camera seesClass E $40M Series B, Aug 2025, on the thesis that procedure knowledge should be captured from execution rather than authored ahead of it. No published operational outcomes. Vendor-reportedFunding coverage, Aug 2025
Aviation
(regulated)
Delta TechOps and peers - FAA-approved drone-based airframe inspectionClass B under regulatory oversight Demonstrates the regulated pattern: automated visual capture is approvable; automated judgement is not. The AI narrows where a certified human looks; the signature stays human. Trade pressAviation Week, Nov 7 2024 and Mar 24 2025
Field service
(sector-wide)
BCG, Future of Field Service with AI - executive perspective across the service lifecycle 15-20% revenue impact and 5-10 pt gross-margin improvement at full deployment; 20-30% productivity lift from smart dispatch; 24% mileage cost reduction; +14% work orders per hour. Explicitly frames the work as 70% people and process, 30% technology. Analyst estimateBCG, Mar 18 2025 - modelled opportunity, not observed averages
Field service
(sector-wide)
TSIA State of Field Services; Geotab State of Field Service; Service Council Voice of the Field Service Engineer 71.4% of field service organizations investing in AI-guided troubleshooting, but only 10.7% measure AI ROI by training impact (TSIA). 93% report partial AI implementation and 75% claim improved first-time fix (Geotab). 52% of technician time goes to paperwork and searching for information; ~45% of engineers are not planning to stay, and only 28% of departures are retirements (Service Council). Industry surveyTSIA Jan 2026; Geotab Jun 2025 (no methodology disclosed); Service Council 2024-2025 Future of Field Service Article
Peer-reviewed methods and data reviewed Analyst / trade modelled or reported, method partial Industry survey self-reported, method often undisclosed Vendor-reported existence proof, not benchmark
Annotation

Notice what is missing from this table: a single published, measured outcome for class E - the workshop's own subject. Every organization in this market is being sold session-to-record capture, and not one of them can cite a number for it. That is not a reason to wait. It is the reason the first organization to instrument it properly gets to define the benchmark - and that instrument is an internal one, because the corpus being built is internal.

4 - The ceiling

What the peer-reviewed record says the models cannot do

The academic literature over this window is unusually useful, because it keeps finding the same thing from different directions: multimodal models are strong at describing and weak at discriminating, and the gap is largest exactly where industrial work lives.

74.9%

Best-model accuracy on MMAD, an industrial anomaly benchmark of 39,672 questions over 8,366 images across seven inspection subtasks. The authors' own verdict: this "falls far short of industrial requirements."

Jiang et al., ICLR 2025 · peer-reviewed benchmark
58.1%

Average accuracy of four leading VLMs on seven trivially simple visual tasks - do two circles overlap, how many times do these lines cross, which letter is circled. Human performance is effectively 100%.

Rahmanzadehgervi et al., ACCV 2024 · peer-reviewed
77-94%

Hallucination rate across every approach tested - LLM-only, retrieval-only, and three RAG variants - on open-ended fault-code recommendation from a real vehicle technical manual covering 224 malfunctions across 22 systems.

Lukens et al., PHM Society 2025 · peer-reviewed
86% vs 57-64%

In the same study, when the task was constrained to a fixed label set, plain retrieval beat every LLM-based configuration - 86% hit rate against 70% for LLM-only and 57-64% for the RAG variants.

Lukens et al., PHM Society 2025 · peer-reviewed
87-96%

Hit rate for all methods in that study when scored at system level rather than fault-code level. The models reliably narrow to the right subsystem and unreliably name the fault - a precise description of what to automate and what not to.

Lukens et al., PHM Society 2025 · peer-reviewed
+20-40%

End-to-end gain from embedding document page images directly rather than parsing them to text first. Layout, schematics and figures carry information that text extraction silently discards.

Yu et al., VisRAG, ICLR 2025 · peer-reviewed

Four findings worth carrying into the room verbatim

The models see poorly at close range. The ACCV result is the one that reorients people. These are not adversarial images or edge cases; they are the visual equivalent of asking whether two lines touch. A system that scores 58% on that is not going to reliably judge whether a wear surface has crossed a tolerance line, no matter how fluently it describes the part.

Retrieval beat generation on the task that mattered. The Lukens result deserves more attention than it has received in the trade press. On a bounded diagnostic task against real technical documentation, adding an LLM to retrieval made results worse - and the RAG configurations, the architecture every vendor is selling, came in below both baselines. The paper's own framing is that task construction dominates model choice.

Video benchmarks have been measuring the wrong thing. Apple's 2025 work decomposing video-LLM benchmarks found that they conflate questions answerable from world knowledge alone, questions answerable from a single frame, and questions that actually require ordered frames. Overall scores hide the temporal weakness. For anyone evaluating a "video understanding" claim in a service context - did the technician do the steps in the right order? - this is the question to ask the vendor.

Diagram structure is a live research problem, not a solved one. Recent IFAC work on extracting procedural knowledge from industrial troubleshooting flowcharts - where spatial layout and technical language jointly carry the meaning - finds model-specific trade-offs between layout sensitivity and semantic robustness, and needs layout-aware prompting to get either. The troubleshooting flowchart in an OEM manual is not a document a model reads; it is a document a model has to be taught to read.

Annotation

Every one of these ceilings is a ceiling on unsupervised open-ended judgement. None of them is a ceiling on capture, on narrowing, on structuring, or on retrieval. The literature is not saying the technology is immature. It is saying the technology is competent at the half of the job the industry finds boring and incompetent at the half it keeps demoing.

5 - The pattern underneath

The scarce asset is the corpus you can edit

Put the deployment record and the benchmark record next to each other and one variable explains both: whether the deploying organization owns and can write to the knowledge the system reasons over.

A field service organization's knowledge sits in four kinds of places, and they are not equivalent:

The term tribal knowledge is doing something specific and worth pausing on. It does not mean undocumented. It means knowledge that lives in a social structure - held by a group, transmitted by telling, validated by whether it worked for someone you trust. Orr's finding was not that the technicians had failed to write things down. It was that the war story carried context the documentation format could not hold: which machine, in which room, in which season, presenting how. Strip the context out to make it a knowledge article and you delete the part that made it usable.

That is the real difficulty in the workshop's title, and it is not a capture difficulty.

The business literature has converged on the same point from the strategy side. Tung and Roussiere argue in California Management Review that as agentic AI commoditizes, "the real differentiator is not the data or even the models, but the tacit knowledge embedded in the judgment of their people" - and their prescription is a semantic layer that structures how decisions get made, not a bigger corpus. The MIT NANDA study that produced the widely-quoted 95% figure locates enterprise failure in the same place: pilots die because the systems have no persistent memory, cannot be shaped to a workflow, and do not learn from correction. Their respondents' complaint - "it's useful the first week, but then it just repeats the same mistakes" - is a complaint about the absence of a writable knowledge layer, stated in user language.

Why the single-source-of-truth answer does not ship

The instinctive response to conflicting knowledge is to fix the source. In a service organization this fails predictably, and the constraint is not legal ownership - it is write access. An OEM manual may sit entirely inside your own company and still be unreachable, because the correction path runs through another team's revision cycle. Field truth cannot reach the source at field speed. A company-wide single source of truth dies on multi-use every time: the moment two groups need the same record to serve different purposes, the governance cost of reconciling them exceeds the value of having reconciled them.

The workable alternative is to take the other group's source of truth read-only and bind corrections to it by reference rather than by edit - an annotation layer, authored by the group that has the field knowledge, available to the source owner, and explicitly accepting that the two systems will sometimes give different answers. That acceptance is the hard part, and it is a governance decision, not a technical one.

Annotation - the position

This is where vision changes the economics rather than the capability. An annotation layer has always been the right architecture and has always failed on input cost: nobody writes the note. What is new since 2023 is that the note can be a thirty-second video the technician was already going to record, structured automatically into something that binds to a specific figure on a specific page of a manual nobody can edit.

Class E "Session-to-record" is not a feature. It is the only capability in the taxonomy that builds an asset instead of consuming one.

What "scaling" actually demands

Capture is the easy verb. Scaling captured tribal knowledge is four separate problems, and the tooling in this category solves roughly one of them:

Only the first of these is a technology purchase. The other three are governance, and they are, in my view, the reason class E session-to-record has no published outcomes - not because nobody has captured anything, but because captured knowledge that isn't distributed, adjudicated, maintained and attributed produces no measurable effect to publish.

The retrieval consequence

If the annotation layer is visual, retrieval has to be visual too - and this is where the VisRAG result stops being an academic curiosity. Text-parsed pipelines discard exactly the content that carries diagnostic meaning in service documentation: the exploded view, the callout number, the flowchart branch, the photo of the correct assembly. Embedding the page image directly recovers 20-40% of end-to-end performance. A technician's photograph of a real assembly and the manual's figure of the intended assembly are, for the first time, comparable objects in the same index.

You do not need a thousand words to find the image that gets you to solution.

6 The Adjudication Layer

Adjudication, which vision makes harder

There is a moment in every guided-diagnosis system that the literature almost never addresses directly: the technician is holding two sources that disagree. The manual says one thing. The service note says another. The photograph in front of them matches neither. Someone has to choose.

Technicians adjudicate constantly, without naming it. Manuals are written cheaply and to a "good enough" standard on purpose, and the human absorbs the disconnect - that absorption is invisible labour that no system currently accounts for. Adding a camera does not remove the conflict; it adds a third, highly credible source to it. A photograph is evidence, and evidence that contradicts documentation is a governance event, not a retrieval result.

Two design commitments follow, and both cut against how these systems are currently built:

The human-factors literature explains why confident wrongness is the specific risk. Reviews of automation bias in human-AI collaboration find that people under time pressure accept plausible machine output at rates well above its accuracy - and the trust data is not stable in the other direction either: Deloitte's index, reported in HBR, found trust in company-provided generative AI fell 31% between May and July 2025, and trust in autonomous systems fell 89% over the same window. Neither over-reliance nor collapse of confidence is a good operating state, and both are produced by the same thing: a system that cannot say why it thinks what it thinks.

Annotation

Engineers designing these systems consistently underestimate how complicated the moment of choice is. It is not a ranking problem. It is the entire job, and it is the part of troubleshooting that experience actually buys.

7 - Implications

What this means for a service organization in 2026

The expertise curve does not flatten the way the pitch says it does. Two 2025 field experiments bound this precisely. In the P&G study, 776 professionals working with AI dissolved functional silos - R&D and Commercial staff produced comparably balanced solutions, and AI-equipped individuals matched the output of human pairs working without it. But the IG Group experiment found a hard limit: across 78 workers grouped by expertise distance, AI equalized conceptualization quality (4.05-4.18 out of 5 across all groups) while leaving a 13% execution-quality gap between occupational insiders and distant outsiders. AI provides the map; navigating the terrain is still domain knowledge. For a service organization staring at a technician shortage, that is the difference between compressing time-to-proficiency and eliminating the requirement for proficiency. Only the first is on offer.

Instrument class E before you need to defend it. Nobody has published a measured outcome for session-to-record capture. If you deploy it without a baseline - knowledge article production rate, mean time to first correct article after a new failure mode appears, proportion of resolutions traceable to a captured session - you will be unable to distinguish a working system from an expensive video archive two years from now.

Bound the answer space deliberately. The Lukens finding is the practical one: the same corpus and the same models produced 77-94% hallucination on the open task and 86% accuracy on the constrained one. Constraining the output to an enumerated set you control is not a limitation of the deployment; it is the deployment.

Budget as BCG frames it, not as the vendor does. 70% people and process, 30% technology. The technology in this category is now the cheap part, and the organizations that fail will fail on the write-access question, the adjudication question, and the technician-trust question - none of which appear on a platform comparison matrix.

Treat retirement-driven knowledge loss as the wrong frame. Only 28% of departing field service engineers are retiring; the rest leave for burnout, unclear paths, and disengagement, and 52% of a technician's time already goes to paperwork and searching for information. Knowledge is not walking out slowly at 65. It is churning out continuously at every age, and the capture mechanism has to survive that turnover rate - which means it cannot depend on anyone volunteering to write documentation.

8 - Practical Considerations

Gates and Steps

A readiness frame, and a prompt sequence for Capturing and Scaling Tribal Knowledge with Video Intelligence. Gates 1-4 test whether a capture use case will work at all; gate 5 tests whether it will still be working in two years, and is the one most companies skip.

The first four gates - apply to any proposed visual AI use case

GATE 1

Do you own the corpus?

Not "do you have access to it" - can you write to it, this quarter, without another team's approval?

Fail → you need an annotation layer before you need a model.
GATE 2

Is the answer space enumerated?

Can you write down every acceptable output on one page? Fault codes, part numbers, dispatch / no-dispatch, pass / fail.

Fail → expect the open-task hallucination range, not the constrained-task accuracy.
GATE 3

Does a decision change?

Name the decision the output alters and the person who makes it. If the answer is "it gives them information," there is no measurable outcome.

Fail → you have a demo, not a deployment.
GATE 4

Is capture free to the technician?

Any capture step that costs the technician time competes with the 52% of their day already lost to paperwork - and loses.

Fail → adoption decays to zero within two quarters regardless of quality.

The future gate - apply to any proposed visual AI use case

GATE 5 - THE SCALING GATE

What happens when it goes stale or contradicts?

Name the mechanism that retires a captured item when the build changes, and the mechanism that handles two captures that disagree. Both, specifically.

Fail → you are building a second source of confidently wrong answers with a two-year fuse.

Have your team review these questions to aid in the design

  1. Where does your field truth go to die? Everyone has a place where a technician's correction gets written down and never reaches the source. Name it out loud. This surfaces the write-access problem without needing the theory first.
  2. What's the last thing a technician told you that contradicted your documentation - and what happened to it? Gets the team from abstraction to a specific artifact within two minutes.
  3. If your best technician retired tomorrow, what would you actually lose? Press past "experience." Push for the specific discriminating judgement - the thing they notice that isn't in any manual.
  4. What is your organization's answer key, and who owns it? The bridge to the five gates. Insurance had one. Ask what yours would be.
  5. You've captured a thousand sessions. Now what? Force the team past the capture pitch. Who reads them, when, and what happens the first time two of them disagree. This is where the title's second verb gets tested.
  6. What would you have to be able to measure to know that captured knowledge was paying off? Nobody has published this - there is no benchmark for class E, session-to-record, anywhere in the literature. The team has to draft this to create a business case.
  7. Where are you currently asking a model to make a judgement you would not let a first-year technician make unsupervised? The closing question. It reframes the ceiling findings as a staffing policy the team already understands and already knows how to enforce.
9 - Sources & method

What was reviewed and how it was graded

Coverage window: January 2023 through September 2026. Priority was given to peer-reviewed benchmarks and field experiments, then to analyst and business-press work with disclosed method, then to vendor-published deployment records used only as existence proofs. Where a figure reached this review through a secondary source, that is marked. Two known gaps: the Service Council's 2025 Voice of the Field Service Engineer data set is member-gated and its figures here come via secondary citation, and no independent audit of any visual-AI dispatch-avoidance claim was located in the window.

Business press and analyst

Prepared September 2026 for Capturing and Scaling Tribal Knowledge with Video Intelligence, a co-hosted workshop at the Smarter Services Executive Symposium. Sections 1-4 and 7 report the literature; sections 5 and 6 and all passages marked Annotation argue a position and should be read as such. Julian Orr, Talking About Machines: An Ethnography of a Modern Job (Cornell University Press, 1996) sits outside the review window and is cited as the origin of the argument, not as evidence within it.