Field notes5 min read
The Arabic in the filing cabinet
Newspaper Arabic is easy. A customs form from 2009, a call recording in Gulf dialect and a report that switches to English mid-sentence are the job. Three use cases from the refinery floor.
There is a moment in almost every project when someone opens the archive. Sometimes it is a room of boxes. Sometimes it is a shared drive with fourteen thousand Word files named final, final2 and final-final. Once it was a cupboard of cassette tapes. The question is always the same, and it is asked a little hopefully: can the model read this? The honest answer, on that day, is not yet.
Three kinds of Arabic in one building
The Arabic that large language models learn from the internet is the Arabic of news sites and encyclopaedias. Modern Standard, carefully written, by people whose job is to write. Institutions do not produce that. They produce three other things, usually all at once.
Dialect: a complaint taken down from a phone call, a field note from an inspector, a message from a citizen, in the Gulf dialect, with its own words and no agreed spelling. A model trained on newspapers reads it the way a foreigner with a textbook would, which is to say it gets the gist and misses the point.
Mixture: a report that changes language in the middle of a sentence. An English technical term inside an Arabic clause, then a product code, a place name, a unit. The model has to hold both languages at once. The general models translate the English term into an Arabic word nobody in the building uses, and the engineers stop trusting the output on the first page.
Paper: decades of records that exist only as scans. Stamped forms, handwritten margins, tables typed on a machine and photocopied twice. Reading them is a vision problem before it is a language problem, and if the two are solved separately the model learns from the scanner's mistakes.
A customs archive, twenty years of forms
- The situation
- An authority holds two decades of declaration forms as scans. Stamped, annotated by hand, some photocopied so many times the boxes have drifted. Nobody can search them. A question about a shipment from 2011 means a person and a box and an afternoon.
- What we build
- A reading pipeline before any language model: find the page, recover the layout, read the print and the handwriting, keep the table as a table. Every field comes out with a confidence, and the low-confidence ones are routed to a person instead of being guessed. Then a language model is fine-tuned on the recovered text, so a question in plain Arabic finds the right form.
- What the reviewers see
- The scan on the left, the extracted fields on the right, and the doubtful ones highlighted. They correct what is wrong. Every correction becomes a training example, and after the first few thousand the model stops making that class of mistake. This loop is the product we call Refinery. The archive becomes the training set; nobody buys data.
- What it needs from you
- The scans, in whatever resolution they exist. Three reviewers for a month, people who know the forms. A written definition of what correct means for the five fields that matter most, because two experienced clerks will disagree about the sixth.
The archive is the training set. Nobody has to buy data.
Call recordings in dialect
- The situation
- A service centre has recorded its calls for years, for compliance. The recordings are in Gulf Arabic, with English words for anything technical, and nobody has ever listened to more than a sample. Supervisors know complaints are rising about one thing and cannot say what.
- What we build
- Speech-to-text trained on the centre's own calls, so it knows the dialect, the product names and the way callers switch to English for numbers. Then a search across every transcript, and a short summary per call. A supervisor can type a phrase and find every call in the last year where a customer said it.
- What changed
- Complaint categories come out of the calls instead of a drop-down menu the agents never used properly. The training team hears the ten worst calls of the month without listening to a thousand. The recordings never leave the building; the model was trained on the centre's servers and runs there.
- What it needs from you
- The recordings and the consent policy that governs them. About a hundred hours transcribed by people first, which is the ground truth the model learns the dialect from. A glossary of product names, which takes an afternoon and saves a month.
Engineering reports that switch languages
- The situation
- An energy company's reports are written in Arabic prose with English terms, part numbers and units, and the tables carry the numbers that matter. A general model asked to summarise them translates the English terms, rounds the units, and produces something fluent and wrong.
- What we build
- A model trained on five years of the company's own reports, which keeps English terms as English, keeps every number with its unit attached, and learns that a certain abbreviation means a certain valve on a certain line. It extracts the figures into a table an engineer can check against the original in a minute.
- What it needs from you
- The reports, and the engineers' own glossary, which every department has and no department has written down. One engineer to review the first hundred extractions. That review is where the model learns the difference between fluent and correct.
How we know it works
We hold back a set of the institution's own files, chosen by the institution, that the model never sees during training. When the model is ready it reads those, and we count how often it is right, field by field, question by question. That number is the one we report. Public benchmarks are useful for choosing a starting model and useless for deciding whether a customs clerk can trust the output. When the model is wrong, the wrong answers are collected, reviewed and fed back, and the number goes up in the next run. It usually goes up quickly, because the mistakes an institution's documents provoke are consistent, and consistent mistakes are learnable.
Enquiries and interviews: info@goldenlineai.com