AI that works on Arabic documents.

Most document automation is demonstrated on clean English invoices. Yours are in Arabic, or in both, and they are scanned. This is what changes, what we can evidence, and where it still breaks.

Why Arabic is a separate problem.

Arabic is cursive, letters change shape by position, and diacritics carry meaning that a scanner often loses. Text runs right to left while the part numbers, currencies and company names inside it frequently run left to right. Vendors benchmark on printed Modern Standard Arabic and your documents are handwritten, stamped, faxed, or all three.

None of that makes the work impossible. It makes the difference between a demo and a production system much wider than it is in English, which is precisely the gap this firm exists to close. It also means the first honest number you get from us will be lower than the one a vendor showed you.

What we do, and where each of them stops.

Arabic and bilingual document extraction

Invoices, purchase orders, delivery notes and contracts that arrive in Arabic, in English, or in one document containing both. The extraction is measured on your documents rather than on a benchmark, because a benchmark has never seen your suppliers.

Where it stops Off-the-shelf OCR degrades on Arabic more than vendors admit, and it degrades furthest on exactly what you have: handwriting, stamps, scanned fax, and tables. Expect the first measurement to be worse than the demo, and expect to be shown it.

Mixed-script fields and the numbers problem

An Arabic document routinely carries Latin part numbers, English company names and Western digits inside a right-to-left line. Getting that wrong silently transposes an amount or a reference, which is the failure mode nobody notices until reconciliation.

Where it stops Arabic-Indic and Western digits both appear in Gulf business documents and the same figure can be written either way. This is a validation rule, not a model capability, and it is the kind of thing that has to be specified rather than hoped for.

Dialect in customer-facing text

Support tickets, WhatsApp messages and complaints arrive in Khaleeji or Levantine, not in Modern Standard Arabic. Triage and routing are tractable on dialect. Generation in dialect is a separate decision and usually the wrong one for a regulated business.

Where it stops We will not claim dialect coverage we have not measured on your traffic. The honest answer for a new dialect is a measurement, and it takes about a week.

Retrieval over an Arabic corpus

Answering from your own Arabic policies, contracts and manuals, with the passage it came from attached. Arabic morphology breaks naive chunking and retrieval in ways that are invisible until somebody asks a question the system answers confidently and wrongly.

Where it stops This needs an evaluation set written in Arabic by somebody who knows the domain. If that cannot be resourced on your side, say so early, because without it there is no way to tell whether the system works.

Does it have to be an Arabic model?

Usually not, and sometimes yes for reasons that have nothing to do with quality. Here is the landscape and our position on each, which is a position about selection rather than a ranking.

  • FalconTechnology Innovation Institute, Abu Dhabi Open weights, which is the whole argument. If a workload has to run inside your own tenancy or air-gapped, an open model you can host is not competing with a frontier API, it is the only option in the room.
  • JaisG42 and MBZUAI, Abu Dhabi Built Arabic-first rather than translated into Arabic. Worth measuring against a frontier model on your own documents, and worth believing the measurement rather than the positioning.
  • ALLaMSDAIA, Saudi Arabia Where a Saudi public-sector or GRE procurement expresses a preference for a national model, this is usually the one meant. That is a procurement fact before it is a technical one, and it is a legitimate reason to choose it.
  • FanarQatar Computing Research Institute The same argument as ALLaM, in Qatar.

The selection rule is the same one we use everywhere: measure the candidates on your documents, write down what counts as good enough before measuring, and keep the reasoning so a future team can revisit the decision when the landscape moves again. It will.

A frontier model wins on quality more often than regional positioning suggests. A model you can host wins whenever the workload cannot leave your tenancy. Those are different questions and conflating them is how organisations end up with a model they cannot deploy.

How we would prove it on your documents.

Send a sample under NDA. We measure extraction accuracy field by field on your own documents, in the terms your process cares about rather than as a benchmark score, and we show you the failures rather than the average.

That measurement is the deliverable, whichever way it comes out. If the accuracy will not clear the threshold your process needs, that is the finding and we will say so in writing, which is the same commitment every other phase on this site carries.

Where the documents would sit while we do it

Bring us ten of your worst documents.

Not your cleanest. The ones that jam the process now are the ones that decide whether this works.

You'll leave with 2 to 3 scored use cases, an effort estimate, and an honest cost range, whether or not we work together.

Loading the calendar.

Contact us.

A person reads every message and a person answers it, within one business day.

Used to reply to you.

Or call +1 (844) 844-0097, or write to hello@abstractionadvisors.com.