Internal research report · Saudi Arabia

The Clean Claim Playbook

Why Saudi insurers refuse clinic claims, how their AI adjudication engines actually decide, and a concrete plan — process, people, and a technical build — to get a partner clinic from ~20% rejections to single digits.

September 2026 Scope: KSA private outpatient clinics Method: 9-track research sweep, ~130 sources Status: Working draft for the clinic partnership
Executive summary

The denial problem is real, but it is mostly an engineering problem — and that is good news

15–25%
Typical first-pass rejection rate at Saudi clinics & SME hospitals
SAR 3–4.5B
Estimated value denied to Saudi providers every year
94%
Of NPHIES claims decided in under 30 minutes — by software, not people
73×
Higher odds of documentation denials without a clinical-documentation (CDI) function
2.54%
Final rejection rate achieved by the best-run RCM operation in KSA (market avg 9.5%)
~70%
Of denied claims are recoverable when resubmitted properly and on time

The situation. Since NPHIES went live in 2021, every eligibility check, pre-authorization, claim, and payment between Saudi providers and insurers flows through one national FHIR platform, and roughly 94% of claims are adjudicated in under 30 minutes. That speed means one thing: your claims are being judged by rule engines and machine-learning models, not by people. Tawuniya runs SAS fraud-detection analytics and states in its annual report that it uses machine learning on health claims; Bupa Arabia reported live AI models and 56 robotic-process bots as far back as 2021; TPA-side engines like Munich Re's SHIELD ship with 37,000 adjudication rules plus a million Milliman rules and literally advertise that they "justify denials." The arms race your clinic partner feels is real and it will not slow down.

The reframe the data forces. Across every study we found — a 13,974-claim market study, a 13,467-claim tertiary-hospital audit, a two-hospital CDI comparison, CHI's own 2022 numbers — the dominant denial causes are on the provider side: missing medical-necessity justification (40.5% of analyzed denials), coding errors (27.8%, with 26.8% of primary diagnoses miscoded in one audit), incomplete documentation, expired or mismatched pre-authorizations, and eligibility lapses. Kuwait's single-payer scheme rejects 3.85% of claims; the best-run outsourced RCM portfolio in KSA lands at 2.54% final rejection across 120,000 claims a month. The gap between 20% and 3% is not payer malice — it is process and tooling. Roughly 85% of denials are preventable, and about 70% of the rest are recoverable.

The strategy in one sentence. The payer runs a four-gate machine (platform validation → benefit rules → clinical rules → statistical profiling), so the winning move is to run the same machine first, on your side, before the claim ever leaves the clinic — a "pre-adjudication" layer that combines a deterministic rules engine (mirroring NPHIES's published validation and denial codes, which are public), a denial-risk model trained on the clinic's own adjudication history, and an LLM layer that checks documentation completeness and drafts medical-necessity justifications for physician sign-off. Around that engine sits a small, disciplined operating routine: eligibility checked at every visit, pre-auth locked to billed codes, a 48-hour resubmission SLA, and a weekly denial council that feeds every refusal back to the person who can prevent the next one.

Why words become practice here. Three things make this executable rather than aspirational, all specific to KSA: (1) the payer's rulebook is largely published — NPHIES exposes its ~95 standardized denial-reason codes, its structural validation codes, its FHIR profiles, and CHI publishes the coding standards and benefit schedule, so the deterministic layer can be built from public specs; (2) the clocks are regulated — 60-minute pre-auth with deemed approval on silence, 30-day settlement, 15-day resubmission, a conciliation center that settles ~69% of disputes in ~12 days — so a clinic that runs on calendar discipline has enforceable leverage; and (3) every denial comes back with a machine-readable reason code, which is free, labeled training data for the risk model and a free diagnosis of the clinic's own process.

What we recommend. A 90-day pilot at the partner clinic in three phases: instrument and measure (weeks 1–4), deploy the deterministic scrubber plus front-desk and physician workflow changes (weeks 5–8), then switch on risk scoring, the LLM documentation checker, and an appeals blitz on the backlog (weeks 9–13). Targets: first-pass denial below 10% by day 90, below 5% within 12 months, final rejection below 3%, and a documented before/after that would be the first named public case study in the Kingdom — itself a commercial asset. The full roadmap is Part VII.

How to read this report

Parts I–III explain the system and the adversary — read them to understand why each fix works. Parts IV–V are the operational playbook a clinic can run with no new software. Parts VI–VIII are the build: architecture, models, compliance, vendors. Part IX is what will change under your feet. Every load-bearing number is sourced in the sources section; vendor-published figures are flagged as such.

Part I

The machine you are up against

1.1 Who owns what (and why it changed in 2024)

Two regulators now split the field, and sending a complaint to the wrong one costs weeks. The Council of Health Insurance (CHI, formerly CCHI) owns everything about the content of claims: the Unified Policy and Essential Benefit Package (coverage floor of SAR 1M per year for large-company policies since October 2022), the Unified Contract between insurers and providers, the coding standards, the denial-code taxonomy, provider accreditation — and NPHIES itself. On 4 March 2024 the new Insurance Authority (IA) took over licensing, qualification, and conduct complaints for insurers and TPAs. Practical rule of thumb: claims, coding, platform, and provider-side disputes → CHI; insurer misconduct, licensing, and payment-delay complaints → IA.

NPHIES (المنصة الوطنية لتبادل المعلومات الصحية) launched in January 2021, is operated with Lean Business Services (a PIF company), and is mandatory: a provider that is not connected cannot submit claims or get paid. As of late 2025 it connects roughly 6,600 provider organizations, 28 insurers and 8 TPAs serving ~14M insured, and has processed 130M+ transactions. Technically it is an HL7 FHIR R4 message broker: eligibility, prior authorization, claims, attachments, payment notices, and reconciliation all travel as FHIR message bundles. A crucial subtlety: NPHIES validates but does not adjudicate — its own implementation guide calls it a "smart courier." It rejects structurally invalid claims at the gate (before any payer sees them); the actual approve/deny decision happens inside each payer's or TPA's engine, and comes back through the platform as standardized response codes.

1.2 The claim lifecycle and its regulated clocks

Almost every deadline in the Saudi claims cycle is written into CHI's Unified Contract — and several of them are weapons for the provider if the clinic keeps evidence. The one most clinics under-use: if the insurer does not answer a pre-authorization request within 60 minutes, the request is deemed approved (Art. 3.4 / 4.4h) — but only if you can prove when the payer received it, so timestamped capture of every auth request is not bureaucracy, it is money. Insurers also cannot retroactively cancel an approval for a service not yet rendered, cannot pay staff commissions on denied claims (Art. 2.3), and cannot override the treating physician's clinical decision (Art. 4.19) — approval verifies cost-effectiveness only.

The regulated clock — one claim, seven deadlines
CHI Unified Contract, Chapter 4 Art. 5 & Chapter 3. Miss a green deadline and the clinic loses leverage; hit them all and unpaid claims become enforceable disputes.
Visit booked Care delivered Claim submitted Payer settles Resubmission Final settlement Dispute eligibility ~5 s pre-auth ≤ 60 min ≤ 30 days ≤ 30 days ≤ 15 days ≤ 15 days ≤ 30 wk-days check coverage, benefits, network silence after 60 min = deemed approval from claim maturity; late = deniable pay, or rejection statement w/ reasons denials + supporting documents payer reviews, pays, مخالصة CHI disputes track; conciliation ≤ 1 yr Every deadline on this line is contractual (Unified Contract Art. 3.4, 4.4, 5.1–5.11) — a clinic that logs timestamps can enforce all of it
Note on versions: an earlier CCHI circular (3531/1681, effective Oct 2018) set the cadence at 45/45/22/22 business days; the current Unified Contract text says 30/30/15/15 days. Your partner clinic should read its own signed contract — the operative windows are contractual — and treat the shorter cycle as the planning default.

Three more lifecycle rules that clinics routinely lose money on: emergencies need no pre-auth but the insurer must be notified within 24 hours — late notification without excuse forfeits payment entirely; outpatient episodes likely to exceed SAR 500 need prior approval; and claims on NPHIES are immutable — you cannot edit one, only cancel and resubmit, which is why getting it right pre-submission beats any correction workflow.

1.3 The coding stack every claim must speak

KSA claims are coded in a specific, versioned dialect, and payer engines check it mechanically:

The mandated coding & documentation stack
LayerStandardWhat breaks claims
DiagnosesICD-10-AM (Australian Modification), 10th ed. — mandatory since 1 Jan 2020"Unspecified" codes when documentation supports 5th/6th-character specificity; chronic conditions without type/complication
Procedures & servicesSaudi Billing System (SBS) — ACHI-based, extended for outpatient care. V3.0 mandatory since 1 Jan 2026: 9,937 active codes, 412 V2 codes retiredRetired V2 codes, wrong principal-procedure sequencing, drugs/supplies coded as procedures
Coding rulesACS (Australian Coding Standards) + SBSCS (Saudi standards)Sequencing violations, misuse of "unlisted" codes
Case-mix (admitted)AR-DRG v9.0Comorbidities not documented → wrong DRG → underpayment or denial
DrugsSFDA GTIN codes, checked against the CHI (Dhaman) Drug Formulary — which links each drug to covered ICD-10-AM diagnosesDrug billed without a formulary-linked diagnosis; non-SFDA-priced items; missing days-supply on outpatient lines
DevicesSFDA GMDN codesUnregistered/uncoded consumables
Claim dataMinimum Data Set (MDS v3.1) + UCAF/DCAF 2.0 forms, physician-signed and stampedMissing vital signs, chief complaint, HPI; unsigned forms. One market study found 2 of 3 EHR systems non-compliant with MDS

The workforce reality behind that table is the quiet crisis: CHI's audit framework expects HIMAA/AHIMA/AAPC-credentialed coders, a KSA-specific certification (CCP-KSA) now exists, and yet hospitals run at roughly half recommended coder staffing — in small clinics, physicians or billing clerks usually code themselves. A Najran hospital audit found 26.8% of primary diagnoses miscoded. This is exactly the gap software can close: coding assistance and validation is the single most under-served, highest-frequency pain in the segment.

1.4 How big the problem actually is

Rejection numbers vary by source and by definition (first-pass vs final), so the honest picture is a range — but every serious measurement puts unmanaged Saudi providers well above global peers, and well above what managed operations prove is achievable:

First-pass rejection rates — KSA vs the world, vs what is achievable
Share of claims rejected on first submission, %. Green = Saudi market measurements; grey = comparators; the two bottom bars show post-resubmission (final) rates.
CHI official, NPHIES launch (Jan 2022) 33% — CHI Secretary General, claims rejected for inaccurate/incomplete data 33% Saudi SME clinics & hospitals (typical) 20–25% — Glance market study of small/medium providers 20–25% CHI official, after 2022 fixes 22% — after NPHIES code updates and weekly payer follow-up 22% Tertiary hospital audit (2023) 18.8% — King Fahad Hospital of the University, all 2023 claims 18.8% UAE market average 12–18% — UAE industry sources 12–18% US hospitals (Kodiak, 2024) 11.8% initial denial rate, 2,100+ US hospitals 11.8% Kuwait AFYA single payer (4.4M claims) 3.85% — 4.44M claims, 2016–2023 3.85% FINAL REJECTION (AFTER RESUBMISSION CYCLES) KSA market average, final 9.52% — market average final rejection (AccuMed benchmark) 9.52% Best-run KSA RCM portfolio, final 2.54% — AccuMed, 120,000+ claims/month across 1,300+ provider contracts 2.54% ◂ <5% best-practice target Sources: CHI via Al Riyadh (2022) · Glance Care market study · Al-Kahtani et al. 2025 · Kodiak Solutions 2024 · Frontiers Pub. Health 2025 · AccuMed/Cirrus. Definitions differ across sources; treat as ranges, not precision.

The financial translation for a clinic: at 4,000 claims/month with an average claim of SAR 1,800 and a 12% rejection rate, roughly SAR 864,000/month sits in the resubmission queue and ~SAR 259,000/month is never recovered (industry model). System-wide, CHI reported SAR 16.3B of claims flowing to insurers in a single month of 2022, with payer loss ratios of 82–92% — which is precisely why payers scrutinize aggressively: their own margins are thin.

Part II

Why claims are actually refused

Denials in KSA come in two fundamentally different species, and mixing them up is the most common analytical mistake clinics make:

  • Technical rejections happen at the NPHIES gate, before any payer sees the claim: blank mandatory fields, invalid or retired codes, missing MDS elements, a pre-auth reference that doesn't match, wrong date formats, an expired practitioner license. NPHIES publishes 100+ numeric validation-error codes for these. They are 100% preventable by software — no clinical judgment involved. Industry estimates put 5–8% of all claims in this bucket alone.
  • Adjudication denials come from the payer's engine after validation, and NPHIES standardizes every one of them into five published code families. This taxonomy is the single most useful artifact in the entire ecosystem, because each family has a different owner and a different fix:
The five NPHIES denial-code families — and who fixes each one
FamilyMeaningTypical codesRoot owner in the clinicPrevention / recovery
SE Supporting evidence — documentation inadequate or missing SE-1-1 vitals missing · SE-1-2 HPI inadequate · SE-1-4 physical exam · SE-1-6 investigation results Physician at point of care Documentation templates + completeness check before claim close; recover by attaching the named element via the Communication transaction
MN Medical necessity — service "not clinically justified per guidelines, without additional supporting diagnosis" MN-1-1, MN-1-2 Physician + coder together Link every service to a justifying diagnosis at order time; recover with a physician-signed justification citing a clinical guideline
AD Administrative — internal inconsistency AD-1-4 diagnosis↔service mismatch · AD-3-5/6 age/gender implausible · AD-2-4 duplicate · AD-2-5 late submission · AD-3-2 use bundled code Coder / billing desk Deterministic pre-submission edits catch essentially all of these; AD-2-5 (late filing) is usually unappealable — prevent, don't fight
BE Billing errors — money & authorization mechanics BE-1-4 pre-auth required, not obtained · BE-1-5 claim inconsistent with what was authorized · BE-1-6 calculation discrepancy · BE-1-1 co-pay not collected Front desk + billing Lock billed codes to the authorized set; validate price lists; collect deductibles at visit
CV Coverage — the policy genuinely doesn't cover it CV-1-3/1-4 diagnosis/service not covered · CV-2-1 not a covered member · CV-3-2 too frequent · CV-4-9 included in another service Front desk (eligibility & benefits) + patient conversation Benefit check before service; for genuinely uncovered care, patient signs and self-pays before the procedure (Unified Contract Art. 9.5) — this family should shrink to near zero as a "denial" and become a front-desk conversation instead

What actually fills those buckets? The two best claim-level datasets in the Kingdom agree on the shape even where the labels differ:

Market study — 13,974 claims, 26 providers
Glance Care, Q1 2022, share of analyzed denials
No medical-necessity justification 40.5% of analyzed denied claims 40.5% Coding errors 27.8% of analyzed denied claims 27.8% Policy non-adherence 2.9% of analyzed denied claims 2.9% Remainder: eligibility, technical & other causes (not itemized)
Tertiary hospital — 1,117 denials analyzed
Al-Kahtani et al. 2025, all 2023 claims of one academic center
Non-coverage 35% — clinical condition not covered by the policy 35% System / technical errors 27% — including one payer rejecting 100% of paper-submitted claims 27% Missing medical data 20% — incomplete clinical documentation 20% Other 18%: duplicates, limits, referrals, personal data

Read together, the two studies say the same thing from different angles: roughly two-thirds of denials trace to documentation, coding, and data quality — things the clinic controls — and most of the rest is coverage friction that belongs at the front desk, not in a claim. Three further findings sharpen the picture:

  • Documentation capability is the single biggest lever. In a controlled two-hospital comparison, the hospital without a clinical-documentation-improvement (CDI) function had 73× the odds of documentation-caused denials (86.2% of its denials, vs 8% at the CDI hospital).
  • Payer behavior is payer-specific and time-specific. In the tertiary-hospital study, one insurer accounted for 65% of denials, rejection probability varied significantly by month (seasonal payer behavior), and that same payer rejected 100% of paper-submitted claims in early 2023 because it only accepted its own portal. A per-payer submission playbook is not optional.
  • Outpatient is where denials live: 93.6% of denials in the academic-center study were outpatient claims — exactly a clinic's book of business.
The claim that survives is the boring claim

Every payer engine is looking for a reason to stop the claim. A claim survives when it is internally consistent (diagnosis ↔ procedure ↔ drug ↔ age ↔ gender ↔ specialty), matches its authorization character-for-character, carries every MDS element and attachment the service type requires, uses maximally specific current-version codes, and arrives inside the window. Nothing clever. Boring, complete, consistent, on time — at scale, every time. That is a systems property, not a heroics property.

Part III

Inside the payer's AI — what the machine actually checks

"The insurance companies built AI to refuse claims" is half-true, and the precise half matters for strategy. What Saudi payers have built is best modeled as three deterministic gates plus one statistical layer. The deterministic gates produce the instant denials your clinic sees daily; the statistical layer mostly doesn't deny individual claims — it profiles the clinic and decides how much scrutiny (audits, information requests, deductions) everything you submit will get.

The four gates a claim must pass — and the mirror the clinic should build
Top row: the payer-side pipeline. Bottom: the pre-adjudication mirror, running the same checks before submission.
Claim FHIR bundle GATE 1 · NPHIES Platform validation schema · MDS · code sets 100+ numeric error codes GATE 2 · PAYER Benefit & eligibility membership · limits · exclusions timely filing GATE 3 · PAYER Clinical rules engine auth match · dx↔service↔drug 37k+ rules (SHIELD-class) GATE 4 · PAYER ML FWA profiling utilization outliers, clustering profiles the clinic, not the claim instant technical reject CV-* denials BE / AD / MN / SE denials audits · info requests · deductions · future scrutiny THE MIRROR — CLINIC-SIDE PRE-ADJUDICATION, BEFORE SUBMISSION FHIR + MDS validator kills gate-1 rejects Eligibility + benefit check kills gate-2 denials Rules + risk model + LLM doc check kills gate-3 denials Utilization self-monitoring keeps the profile clean Every check the payer runs is knowable in advance: gates 1–2 from published NPHIES specs and real-time eligibility; gate 3 from the public denial-code taxonomy plus the clinic's own history; gate 4 from watching your utilization vs specialty norms. A denial is the payer telling you which rule you failed. Feed every ClaimResponse back into the mirror.
Documented payer-side automation in KSA
ActorWhat is publicly documentedWhat it means for the clinic
NPHIES (platform)94% of claims decided <30 min; validates schema, codes, MDS before payer sees the claim; ~95 standardized denial-reason codes in 5 familiesFirst-pass cleanliness is a software problem; the taxonomy is your free rulebook
TawuniyaSAS Detection & Investigation for FWA (2020); annual report states ML on health-claims fraud controls; pre-auth automation; 15+ RPA bots (100+ planned); ran the InsurAI startup challenge on claims automation ($225k prizes, 2025)The Kingdom's most aggressive adjudication automation; also the payer with the largest denial share in the one academic study (65% — partly payer-mix). Deserves its own submission playbook
Bupa Arabia4 live AI models + 56 RPAs reported in 2021; 98% of reimbursement claims processed digitally; "hundreds of thousands of approvals in an hour"; dedicated AI-enabled claims automation under a 450-person pre-auth & claims divisionVolume auto-adjudication: consistency and exact auth matching matter more than narrative
TPAs & enginesMunich Re SHIELD: 37,000 rules + 1M Milliman rules, HDBSCAN clustering + LightGBM, "automates claims medical adjudication… and justifies denials"; GlobeMed Saudi automated pharmacy rules engine; Waseel OTD validates in real time against 18 insurers' medical policiesEven mid-tier payers rent world-class engines — assume every claim faces one
RegulatorsNo KSA rule restricts payer AI in adjudication (no US-style scandal either); IA requires settlement in 15 days (individual) / 45 days (business) with penalties for unlawful rejection; CHI fined insurers SAR 2.7M in one 2022 round; FWA estimated at 10–12% of claimsDon't argue "the AI denied me" — argue rule compliance, documentation, and deadlines. The procedural guardrails are your leverage, not AI ethics

3.1 The adversarial checklist — what gate 3 keys on

Synthesizing the published rulebooks (NPHIES codes, SHIELD's rule categories, Abu Dhabi's published adjudication rules — which codify the same GCC practice), the clinical engines check, in rough order of denial volume:

  1. Authorization exact-match: the preAuthRef exists, is inside its validity period, and the billed codes are identical to the authorized codes. Any drift → BE-1-5, regardless of clinical merit. (Corollary: lock billing codes at authorization time; if the plan changes mid-procedure, re-authorize before billing.)
  2. Internal consistency: diagnosis plausible for age and gender; service plausible for the diagnosis, the provider type, the clinician specialty, and the encounter type (AD-1-*, AD-3-*).
  3. Bundling & frequency edits: no separate billing of components included in another billed service (CV-4-9, AD-3-2); service frequency within policy norms (CV-3-2).
  4. Medication logic: drug's GTIN maps to a formulary entry whose linked diagnoses include one on the claim; dose and days-supply plausible.
  5. Medical-necessity thresholds: does the documentation attached to the claim justify the service against clinical guidelines (MN-1-1)? This is where the 40.5% lives — and note carefully: it is adjudicated from what is written, not from what was done. The care was probably justified; the claim didn't say so.

Gate 4 is different in kind. Statistical FWA engines cluster providers against peer groups: tests-per-visit ratios, repeat-visit patterns, services on national overuse lists (Saudi Arabia runs a Choosing Wisely program targeting vitamin-D panels, unnecessary imaging, PPIs — services on those lists attract utilization rules), coding-distribution anomalies. A clinic that drifts from its specialty's norms gets everything scrutinized harder — more information requests, more deductions, slower payment — even when individual claims are clean. This is why the playbook in Part IV includes monitoring your own utilization statistics: it is cheaper to see your outlier pattern before the payer does.

What not to import from the US debate

There is no Saudi equivalent of the US NaviHealth/Cigna AI-denial scandals, no regulator restriction on payer AI, and a generally pro-automation regulatory posture. Appeals framed as "your algorithm wrongly denied me" have no legal hook in KSA today. Appeals framed as "here is the denial code, here is the evidence answering it, here is the contractual clause and the deadline you missed" win — ~70% of properly-worked resubmissions recover.

Part IV

The operating playbook — what the clinic runs, with or without new software

International denial-management evidence is unambiguous about where to spend effort: ~44% of denials originate at the front end before any claim exists, registration/eligibility alone is the #1 root cause (~24–27%), 84–86% of denials are avoidable, and once a denial happens, a fifth to half of the avoidable value is never recovered — while 35–65% of denied claims are never even reworked because nobody has capacity. Prevention beats rework economically every time (fixing front-end registration data alone cut denials 67% in one published study). The playbook below adapts that evidence to the KSA mechanics from Parts I–III. None of it requires the Part VI build — the build makes it scale.

4.1 Front end — before the patient is seen

  • Eligibility at every visit, not every policy. NPHIES answers in ~5 seconds; run it at booking and at arrival (coverage lapses mid-month; treating a lapsed patient makes the clinic liable). Record the eligibility response reference on the encounter — claims must cite it.
  • Benefit & network check: class, benefit caps remaining, dental/optical/maternity riders, deductible. Anything not covered → patient conversation and signed self-pay consent before service (Art. 9.5). A coverage denial after the fact is a failure of this desk, not of the payer.
  • Authorization discipline: outpatient episode likely > SAR 500 → pre-auth before treating. Physician completes the request → forwarded within 15 minutes → payer clock starts. Timestamp everything; at minute 61 with no answer, the service is deemed approved and the log is your proof. Answer payer queries within 30 minutes (also contractual).
  • One intake checklist, named owners, cross-training. Small-clinic denial prevention is ownership discipline: one person owns eligibility, one owns authorizations (can be the same person), everyone uses the same checklist, and absence cover is explicit.

4.2 Point of care — while the patient is in the room

  • The UCAF is the claim. Vital signs, chief complaint, history of present illness, examination findings, diagnosis, management plan — completed and physician-signed during the encounter. Every SE-family denial code names one of these fields; an empty vitals box is a future SE-1-1.
  • Every order carries its "why". The habit that kills MN denials: each lab, image, drug, and procedure is linked at order time to the diagnosis that justifies it. If a service has no supporting diagnosis, the payer's engine will notice even when the medicine was sound.
  • Specificity beats speed: "diabetes" is a denial; "T2DM with diabetic polyneuropathy" is a paid claim and, under DRG payment, correct revenue. Templates per common presentation (the clinic's top 20 visit types) make this near-free at the point of care.
  • Emergency? Treat, then notify the insurer within 24 hours. Put the notification on a checklist with an owner — a missed notification forfeits the whole claim.

4.3 Mid-cycle — coding & claim assembly

  • Certified coding review (in-house CCP-KSA/AAPC-credentialed, or outsourced): ICD-10-AM to full character depth, SBS V3.0 codes only (412 V2 codes are dead since Jan 2026), correct sequencing, no drugs coded as procedures, dental via the ADA→SBS mapping.
  • Pre-submission scrub against the machine edits: diagnosis↔service↔age↔gender↔specialty consistency, bundling, duplicates, auth-code match, days-supply on outpatient drug lines, attachment presence per service type. Part VI automates this; until then, a printed edit checklist on the biller's desk catches the top ten.
  • CDI-lite: a designated reviewer (nurse or senior coder) queries the physician on incomplete notes before claim close, not after denial. This is the practice the 73× study validates; at clinic scale it is one person-hour a day, prioritized on high-value and historically-denied service types.

4.4 Back end — after submission

  • Denial triage within 48 hours of the rejection statement, sorted by value × recoverability, not first-in-first-out. Route by code family: SE → attach the named evidence; MN → physician justification; BE/AD → fix and resubmit; CV → verify benefit, else write off consciously and fix the front desk.
  • Resubmit correctly on NPHIES: a new claim referencing the original (related.claim, relationship prior), containing all original line items plus the new evidence; use the Communication transaction when only information was requested. Never re-fire the same bundle with new IDs — after 30 days of payer silence, the correct move is NPHIES support with the queued bundle IDs.
  • Reconcile cash monthly against NPHIES PaymentReconciliation details; chase every deduction line, watch predecessor entries for clawbacks, and demand the مخالصة (final clearance) each cycle so deductions surface inside the 15-day window instead of accumulating into an annual write-off fight.
  • Weekly denial council (30 minutes): the top-10 denial Pareto by code, payer, physician, and service; every recurring denial gets a named root-cause owner and a workflow change. This meeting — not any model — is where denial rates actually fall, because it closes the loop to the person who can prevent the next one.

4.5 The numbers to run the clinic by

Denial-management KPIs — definitions and targets
KPIDefinitionClinic today (measure it)90-day target12-month target
First-pass acceptance (clean-claim) rateClaims paid without rework ÷ claims submittedlikely 75–85%≥ 90%≥ 95%
Initial denial rate (count & SAR value)First-response denials ÷ submissions, 3-month trailinglikely 15–25%< 10%< 5%
Final rejection rateDenials never recovered ÷ submissionsoften ~8–10%< 5%< 3% (best-in-class: 2.5%)
Denial overturn rateResubmissions paid ÷ resubmissions filed≥ 60%≥ 70%
Time to resubmitRejection statement → resubmission filed≤ 5 days≤ 2 days
Days in A/ROutstanding receivables ÷ avg daily billingoften 60–90+< 60< 40
Net collection rateCollected ÷ contractually collectible≥ 92%≥ 96%
Denial write-offs as % of revenueValue abandoned ÷ net revenueoften 3–8%< 3%< 1.5%

Two governance notes. First, measure by payer — the academic evidence shows payer identity and even calendar month predict denial; a blended average hides the one insurer generating 60% of the pain. Second, publish per-physician denial feedback gently but visibly; the documented adoption barrier is not physician unwillingness (a 430-physician Saudi survey shows two-thirds already think about insurance in clinical decisions) but the absence of structured support telling them what to write.

Part V

The appeals & escalation machine

Appeals are high-yield and under-used everywhere: internationally ~70% of worked appeals overturn, yet most denials are never appealed for lack of capacity. In KSA the resubmission right is contractual, the evidence requirements map one-to-one to denial codes, and there are three escalation tiers above the payer with real teeth. The design goal: make appealing so cheap and templated that every recoverable denial gets worked.

5.1 The evidence map — answer the code, not the vibe

Denial code → the evidence that overturns it
Code contestedWhat the payer is claimingThe winning attachment set
MN-1-1Service not clinically justified per guidelinesPhysician-signed justification letter naming the clinical practice guideline; the supporting diagnosis (added/coded); relevant investigation results showing indication
SE-1-1 / 1-2 / 1-4 / 1-6Vitals / HPI / exam / investigation results inadequateThe completed UCAF section or the actual lab/radiology report — supply exactly the named element, via the Communication transaction
BE-1-4 / BE-1-5No authorization / claim differs from authorizationThe preAuthRef, NPHIES transaction ID, and timestamp log; for deemed approvals: proof of the 60-minute silence; if codes genuinely drifted, corrected claim after cancel
AD-1-4 / AD-3-*Diagnosis/service inconsistencyCorrected coding with documentation supporting the linkage; often a recode-and-resubmit rather than an argument
CV-1-3 / 1-4 / 3-2Not covered / too frequentPolicy schedule extract showing the benefit; eligibility response; for frequency: clinical justification of recurrence. If truly excluded — don't appeal; fix intake
AD-2-5Submitted lateUsually unwinnable unless you hold proof of timely submission or a documented reasonable excuse. Prevent upstream

5.2 The appeal letter that works

There is no official CHI template; the structure below is synthesized from the contractual mechanism and practitioner sources, and it is deliberately mechanical so an LLM can draft it and a physician can sign it in one minute:

  1. Identifiers: NPHIES claim ID, member ID, authorization reference, service date, rejection-statement date (starts the 15-day clock — cite it).
  2. The exact code contested, quoted with its official English/Arabic definition.
  3. The answer: two or three sentences of clinical narrative mapping each attached exhibit to the code.
  4. The legal anchor: the Unified Contract article involved (Art. 5.4 resubmission right; Art. 3.4 deemed approval; Art. 4.19 physician's clinical authority), and the request: re-adjudication with a revised ClaimResponse.
  5. Physician signature and stamp on the clinical content; indexed exhibits.

5.3 When the payer stalls — the escalation ladder

Escalation tiers above the payer (post-March 2024 map)
TierBodyUse it forClockWhat the data says
1CHI e-complaint (chi.gov.sa / my.gov.sa)Platform, coding, provider-payer process violationsResponse ~10 business daysCheap, fast, creates a paper trail
2CHI Conciliation & Settlements Center (مركز الصلح والتسويات)Financial disputes over denied/unpaid claimsFile within 1 year of the compensation due date — hard jurisdiction limit~69% settlement rate, ~12 days average, binding outcomes, no courts (Q1 2023 data)
3CHI Secretary General → Art. 14 violations committeeLaw/contract violations by the payerComplaint within 90 days of the dispute; committee decisions appealable to Board of Grievances in 60 business daysThe formal enforcement route; CHI does fine insurers
4Insurance Authority / Insurance Disputes CommitteesInsurer conduct, licensing, systematic payment delayIDC appeals only above SAR 50,000; decisions finalThe post-2024 home for insurer-conduct complaints; IA rules require settlement in 15/45 days with penalties for unlawful rejection
The one-year cliff

The Conciliation Center refuses disputes more than one year past the compensation due date. Old denials rot into worthlessness quietly. Part of the 90-day pilot (Part VII) is an aging audit of the existing denial backlog — anything approaching twelve months goes to conciliation now or is written off consciously, never by default.

Part VI

The technical blueprint — what you actually build

The product is a pre-adjudication engine: a layer that sits between the clinic's EMR/HIS and the NPHIES submission point, sees every claim before it leaves, and either passes it, fixes it, or routes it to a human with a specific reason. Three layers, built in this order because each one de-risks the next:

System architecture — the pre-adjudication engine and its feedback loop
Deterministic first, ML second, LLM third; every ClaimResponse becomes training data and new rules.
CLINIC EMR / HIS Front desk Physician Biller / coder encounters · orders · notes eligibility & auth alerts documentation prompts risk-ranked worklist PRE-ADJUDICATION ENGINE 1 · Deterministic validator & rules FHIR/MDS profiles · SBS-ICD matrices · auth cache benefit limits · formulary · bundling · dedup 2 · Denial-risk scorer (per payer) gradient-boosted trees on the clinic's own history score = P(denial) × claim value → review queue 3 · LLM layer (human-gated) UCAF completeness check · necessity justification appeal drafts · RAG over policies & CHI docs physician signs everything clinical Certified gateway or direct (certified) NPHIES FHIR $process-message Payer engine adjudication clean claim Denial analytics store every ClaimResponse + denial code + payment reconciliation → Pareto dashboards · rule updates · model retraining ClaimResponse · denial codes · remittance weekly denial council reads this The loop is the product: each denial makes the next claim harder to deny.

6.1 Layer 1 — deterministic validation (weeks, not months, and the highest ROI)

Everything gate 1 and gate 2 check is published, machine-readable, and therefore reproducible offline:

  • FHIR profile validation against the NPHIES implementation guide's StructureDefinitions (all downloadable from portal.nphies.sa/ig): MustSupport population, episode identifier and per-item invoice-number extensions, days-supply on outpatient medication lines, encounter types, attachment constraints (10MB, title + date). This alone eliminates the 5–8% of claims that die on structural validation. Open-source tooling for NPHIES profile validation already exists as a starting point.
  • The edit engine: ICD-10-AM↔SBS compatibility and age/gender/specialty plausibility matrices (versioned — SBS V3.0 now, with a release process for V4); bundling and duplicate edits; Dhaman-formulary diagnosis↔drug checks; CHI benefit-schedule caps (dental/optical/maternity riders); per-payer quirks learned from history (submission channels, attachment preferences).
  • The authorization cache: every preAuthRef + validity period + authorized code set, with a hard block on claim assembly when billed codes drift from the authorized set (BE-1-5 is entirely preventable this way) and expiry alerts before the preAuthPeriod lapses.
  • The free dry-run channel: NPHIES batch claims return validation-only responses with deferred adjudication — a sanctioned way to test claim batches against the platform's actual validator before committing. Use it nightly.

6.2 Layer 2 — denial-risk scoring (needs data, so start collecting on day 1)

The literature is consistent: gradient-boosted trees on encounter + claim features predict denials well (published AUCs 0.83–0.91; a Jackson Health dual-expert XGBoost/Random-Forest ensemble hit 97.6% holdout accuracy scoring encounters before discharge; Google's Deep Claim beat baselines by 22% recall at 95% precision on 2.9M claims). Two Saudi-specific facts make this tractable and valuable: NPHIES returns labeled denial codes on every rejection — supervised labels for free — and no published Saudi denial-prediction model exists, so the partner clinic's data is both an operational asset and a publishable first.

  • Features: payer/plan, physician, department, service codes and ICD↔SBS combos, claim value, auth flag, patient age/gender, visit type, submission month (the validated Saudi predictor set), plus historical physician- and payer-level denial ratios.
  • Labels: denial yes/no and code family (a 6-class problem — the family predicts the fix). Handle the 75–85/15–25 class imbalance explicitly (class weights or a biased-expert ensemble).
  • Operating point: the score routes, it does not block — score = P(denial) × claim value sorts the biller's review queue so scarce human attention lands where money is at risk. Retrain on a rolling window; expect drift at every SBS release, payer-policy change, and new plan year.
  • Cold start: weeks 1–8 run rules-only while the historical extract (12–24 months of claims + ClaimResponses + reconciliations from the HIS/gateway) is assembled and labeled. A clinic doing 3–5k claims/month yields 40–100k labeled claims from two years of history — enough for a respectable first model.

6.3 Layer 3 — the LLM layer (highest leverage on the biggest denial bucket, strictest guardrails)

The 40.5% medical-necessity bucket and the SE documentation family are language problems, and the published evidence says LLMs handle them well within limits: GPT-4-class models drafted radiotherapy appeal letters physicians rated clear, faithful to the supplied clinical history, and near submission-ready — but consistently failed at citing real literature. Multi-agent RAG over payer medical policies reaches ~95% accuracy on necessity-checklist determinations. So:

  • Documentation completeness check at claim close: does this UCAF answer every SE-code element and support every billed service? Output: a specific fix list for the physician ("no supporting diagnosis linked to the vitamin-D assay"), not a rewrite.
  • Necessity justification drafts generated only from the encounter's own facts, with retrieval over a curated store of clinical guidelines, CHI policy documents, and the payer's stated criteria. Citations come from the retrieval store, never from generation — the one empirically validated failure mode.
  • Appeal-letter generation from the Part V template: denial code + claim facts + evidence index in, signed-ready letter out. Internationally this is the most proven use (appeal prep cut >90%, adopters overturning 40% more).
  • Guardrails: sentence-level grounding checks against the retrieved context, an LLM-judge quality gate, and a hard rule that nothing clinical reaches NPHIES without physician sign-off. Published clinical-safety frameworks get hallucination to ~1.5% — the human gate covers the rest.

6.4 Data protection & hosting — the non-negotiables

PDPL (in force since Sept 2023, enforced since Sept 2024, SDAIA supervising) classifies health data as sensitive: it is excluded from the easiest cross-border transfer exemptions, and transfers need Saudi SCCs/BCRs plus documented risk assessments. The pragmatic architecture: all PHI processing and LLM inference in-Kingdom (Saudi cloud regions or on-prem), de-identification before any external analytics, DPAs aligned with the SDAIA Generative-AI guidelines with any model provider, and explicit patient-consent language in the clinic's intake pack. Design this in from day one — retrofitting residency is expensive.

6.5 Integration path — how you reach NPHIES

Two routes, best used in sequence: (a) ride a certified gateway first — deploy as a validation layer between the HIS and an existing certified vendor (Waseel-class connectivity, or the clinic's current e-claims channel), which requires no NPHIES certification and gets the pilot live in weeks; (b) pursue direct System-Vendor certification later (NPHIES Academy onboarding + sandbox conformance testing) once volume and product maturity justify owning the pipe. The engine's design is identical either way — it validates the exact FHIR bundle that will be submitted.

Build order, stated plainly

Deterministic rules first (weeks, kills ~a third of denials, needs no data), risk model second (needs the historical extract, kills the queue-triage problem), LLM third (kills the biggest single bucket, needs the guardrails running). Teams that start with the model or the LLM ship demos; teams that start with the validator ship revenue.

Part VII

The 90-day pilot with the partner clinic

The pilot's twin goals: cut the clinic's denials measurably, and produce the Kingdom's first named, documented before/after — because our vendor research found that no KSA claims vendor publishes one, which makes a credible case study a commercial moat in itself. Baseline honestly, change one layer at a time, attribute effects.

90-day plan — three phases, parallel workstreams
PhaseProcess & peopleTechnologyExit criteria
Phase 1 · Instrument
weeks 1–4
Baseline audit: 6–12 months of denials coded by NPHIES family, payer, physician, service. Aging audit of the backlog (the 1-year conciliation cliff). Name the checkpoint owners. Start the weekly denial council. Front-desk checklist live (eligibility at every visit, auth timestamp log, self-pay consent for uncovered services). Historical extract from HIS/gateway: claims + ClaimResponses + reconciliations. Denial-analytics dashboard v1 (even a spreadsheet). Begin FHIR validator against NPHIES profiles. Baseline denial rate known by family & payer; backlog triaged; council running; data pipeline flowing.
Phase 2 · Prevent
weeks 5–8
UCAF templates for the clinic's top-20 visit types with necessity prompts. CDI-lite review on high-value/high-denial services. Coder upskilling on ICD-10-AM specificity + SBS V3. 48-hour resubmission SLA on new denials using the Part V evidence map. Deterministic scrubber live on 100% of claims (report-only week 5, blocking from week 6). Auth cache with code-drift blocking. Nightly batch dry-run against NPHIES validation. Payer-specific edit packs for the top 3 payers. Technical rejections ≈ 0; first-pass denial down 5+ points; resubmission time < 5 days.
Phase 3 · Predict & recover
weeks 9–13
Appeals blitz on the recoverable backlog (value-ranked). Physician denial-feedback reports. Monthly reconciliation review with the مخالصة demanded each cycle. Risk model v1 scoring the review queue. LLM documentation-completeness checks + appeal drafting (physician-gated). Measure: model precision on the queue, overturn rate on drafted appeals. First-pass denial < 10%; overturn rate ≥ 60% on worked appeals; documented before/after ready to publish.

The economics that justify it. Take a clinic at 3,000 claims/month, SAR 900 average outpatient claim, 20% first-pass denial: SAR 540k/month enters the denial pipeline; at the typical 70% recovery that still leaves ~SAR 162k/month written off, plus the rework cost of ~500 denials (international benchmarks: $25–$118 per reworked claim). Cutting first-pass denials to 8% and recovering 80% of the remainder brings the monthly write-off under SAR 45k — a ~SAR 1.4M/year swing for a mid-size clinic, before counting faster cash (days-in-A/R improvements of 25–40% are documented for RCM automation) and saved rework hours. Scale that across the ~6,600 NPHIES-connected providers and the product thesis writes itself.

What to measure religiously during the pilot

One metric per layer, attributed: (1) technical-rejection rate → validator; (2) first-pass denial rate by family → process + rules; (3) queue precision (share of flagged claims that would have denied) → model; (4) overturn rate on drafted appeals → LLM; (5) days-to-cash → the whole system. Publish the methodology with the case study — an auditable number is worth ten marketing claims in this market.

Part VIII

Vendor landscape — what exists, what doesn't, where you fit

The KSA claims/RCM field today
PlayerWhat they areNotable factsRelationship to your build
WaseelDominant NPHIES connectivity + RCM SaaSClaims 98% of providers connected, 17–18 insurers; WRCM at SAR 1,499–1,999/mo; post-denial "AI Denial Management" moduleIntegrate, don't compete. Their AI module classifies denials after the fact; your wedge is pre-submission prevention on top of their pipe
AccuMedLargest outsourced RCM (GCC)1,300+ provider contracts, 120k+ claims/mo, 2.54% final rejection; clients incl. Johns Hopkins Aramco, Saudi GermanProof the target is reachable by process; competes for full outsourcing, not for software-assisted in-house teams
KlaimClaims financing fintech$26M raised (2025) incl. a SAR 60M Saudi fund; advances up to 90% of claim value in 24h; Dr. Sulaiman Al Habib a clientNatural partner: financing needs claim-quality scoring — exactly what your engine produces
Glance CareCDI / denial analyticsAuthors of the 13,974-claim study; CDSS + documentation focusClosest analytical competitor; no published pre-submission prediction
Agent.saArabic-first AI claims agentClaims 50%+ rejection reduction, 91% automation — vendor-stated, no named customersValidates demand for the category; unproven delivery is your differentiation opportunity
Clinic SaaS (Cirrus/Anytime, Athir, Nimbo, Health Cluster…)NPHIES-ready EMR/HISGrowing field of certified clinic systems; VIDA (Al Habib's Cloud Solutions) at the enterprise endDistribution channels — the engine should plug into whichever HIS the clinic already runs
Offshore coding (MedCodex etc.)Remote ICD-10-AM/ACHI codingFills the certified-coder shortage from IndiaSignals the coding gap your automation addresses; potential hybrid partner

The gaps nobody fills — confirmed across the sweep, and together they define the wedge:

  • No self-serve, clinic-priced pre-submission denial prediction trained on a provider's own NPHIES adjudication history (US-style predictive editors have no KSA equivalent).
  • No public NPHIES-certified vendor directory and no transparent RCM pricing — publishing clear per-claim or tiered pricing is itself a differentiator.
  • No published, named before/after case study by any vendor — the pilot's documented result would be the first.
  • Almost no Arabic-native denial-analytics UX for billing teams.
  • No instrumentation of deduction (خصومات) patterns and payer payment behavior from reconciliation data — measuring what is currently invisible to every clinic.
Part IX

Risks, pitfalls, and what will change under your feet

Execution risks — and their controls

  • Physician resistance to documentation prompts → make prompts visit-type-specific and short; show each physician their own denial money monthly; never add a form that doesn't remove a denial.
  • Blocking claims too early → run every new rule in report-only mode first; a false-positive blocker erodes trust faster than denials do.
  • Model overfit to one payer/period → per-payer models, rolling retraining, drift monitoring pinned to SBS releases and plan years.
  • LLM clinical errors → grounding checks + physician sign-off gate, never optional; citations only from the curated store.
  • Data residency violations → in-Kingdom inference from day one; PDPL risk assessments documented before any external processing.

Strategic risks — worth naming honestly

  • Never optimize into fraud territory. The line is bright: coding to maximal documented specificity is the job; upcoding, unbundling, or template-generating justifications for services not truly indicated is FWA — payers' clustering engines are built to catch exactly that, penalties escalate, and one flagged pattern poisons the clinic's whole profile. The engine must encode this ethic: it makes true claims defensible, it never makes weak claims look strong.
  • Payer engines will keep tightening (Tawuniya is actively recruiting AI startups for claims automation). Expect the equilibrium to move; the feedback loop is the durable asset, not any single rule set.
  • Vendor claims in this market are mostly unaudited — including some numbers in this report marked as vendor-sourced. Your case study must be the auditable exception.

What's coming, 2026–2028

  • Value-based payment is official policy: CHI's 2025–2027 strategy (47 initiatives) rolls out AR-DRG-based reimbursement and bundled-payment pilots (cataract, diabetes, childbirth, bariatric, knee). Under DRG payment, weak documentation stops meaning "denied" and starts meaning "underpaid on every admitted case" — the same CDI muscle pays twice.
  • SBS is a living standard: V3.0 became mandatory January 2026 and retired 412 codes; version discipline in the rules engine is a permanent feature, not a migration.
  • The market doubles: ~14M insured today, targeting 22–25M by 2030 as Vision 2030 privatizes 290 hospitals and 2,300 PHCs; premiums hit SAR 84.3B in 2025 (+10.7% YoY). More insured lives, more claims-dependent providers, more demand for exactly this capability.
  • A single national payer (CNHI) is being positioned — if it lands, adjudication standardizes further (Kuwait's single payer rejects 3.85%), which reduces denial chaos but raises the bar on structured data quality. Either way, clean-claim capability wins.
Appendix

Sources & method

Method. This report synthesizes a nine-track parallel research sweep (September 2026): NPHIES/CHI regulation, denial statistics, payer-side AI, international denial-management practice, the KSA vendor field, technical architecture & ML/LLM literature, coding standards, appeals & disputes, and academic studies — ~400 searches and source fetches across English and Arabic material. Confidence convention: figures from CHI, NPHIES specifications, and peer-reviewed studies are treated as solid; vendor-published figures (Glance, HealthOrbit, MedCodex, Agent.sa, AccuMed) are directional and flagged in context; where sources conflict (e.g., 30/15-day vs 45/22-business-day settlement cadences; denial-code counts across IG versions), both are reported.

Working research document for the clinic-partnership initiative · Compiled September 2026 · Figures are as reported by their sources at research time; regulated deadlines should be verified against the clinic's signed contracts and current CHI circulars before being relied on operationally.