RPA and AI: How Intelligent Document Processing Transforms Workflows

Most companies launch an intelligent document processing (IDP) project with the wrong problem statement: ‘our data entry is too slow.’ That framing is technically accurate and strategically incomplete. The real bottleneck is rarely the speed of entry itself — it is the absence of clearly defined rules governing what happens to a document after it is read. AI-assisted document processing systems do not resolve that ambiguity; they accelerate it into plain sight. Before any bot reads its first invoice, the more important question is not what the bot will read, but what the reading will trigger.Consider a practical case: a mid-sized construction and contracting company based in Ankara (312 employees, multiple active project sites across Central Anatolia) began an IDP pilot at the end of 2021. The firm was manually entering subcontractor invoices, progress-claim documents, and material delivery forms into its ERP system — roughly 1,400 documents per month, consuming the majority of three accounting staff members’ working hours. The pilot integrated an OCR-based reading engine with an RPA workflow: documents would arrive, be read, have fields extracted, and be pushed into the ERP automatically. In the first six weeks, recognition accuracy was satisfactory. Then the project stalled. Every subcontractor used a different invoice layout, some documents included handwritten margin notes, and — most critically — the firm’s internal approval rules had never been written down anywhere. The question of which progress claim connected to which project code was answered each time by the same senior accountant. She held the rule; the system did not.This is the pattern IDP projects fall into when technical success fails to produce operational value. AI-assisted document processing rests on three distinct layers: document classification (is this an invoice, a delivery note, a contract addendum?), field extraction (reading structured data such as amounts, dates, vendor names, VAT line items), and business rule mapping (determining which ERP record, which approval queue, and which cost centre the extracted data feeds into). Mature solutions exist for the first two layers; machine learning-based OCR engines, once trained, handle varied formats and languages with high accuracy. Tools supporting Turkish character sets and the domestic e-invoice XML standard are accessible in 2022 at a reasonable cost. The third layer is where the gap appears: if the rules are not written, the bot has nothing to apply.A contrasting example from healthcare distribution illustrates what adequate preparation looks like. A medical device distributor in Izmir (378 employees, contracted with the Social Security Institution) invested in an IDP solution for processing supplier invoices and product compliance certificates. Before any system was configured, the project team ran an eight-week process-mapping exercise. For each document type, the question ‘what happens when this document arrives?’ was answered in writing, and the answers were codified as business rules inside the system. Eleven months after pilot completion, monthly document processing capacity had grown to roughly four times the previous manual throughput, and the time accounting staff spent on document entry had dropped to approximately 19 percent of total working hours. One clarification is essential: this outcome cannot be directly compared with the Ankara construction case. The Izmir firm invested in preparation; the Ankara firm did not. The technology stack was comparable. The operational maturity was not.Format diversity is a practical problem that vendors frequently understate. A Turkish exporter’s document set can include domestic e-invoice XML files, English-language letters of credit, German bills of lading, and handwritten Arabic delivery notes — sometimes within the same shipment. Modern IDP engines can handle multilingual sets, but each new format type requires a separate labelled training dataset, a validation cycle, and defined exception-handling rules. The assumption that ‘once the system is installed, all my documents are processed’ collapses at the first edge case. Managing this is not a one-time setup; it is an ongoing configuration discipline. For SMEs, this maintenance burden tends to be invisible during the sales cycle and surfaces as ‘the system doesn’t understand it’ complaints around the sixth month of operation.From a data protection standpoint, document processing projects require specific attention under Turkey’s Personal Data Protection Law (KVKK). Contracts, invoices, and forms routinely carry personal data: the name of a signatory, a national identity number, contact details. Routing these through a cloud-based IDP service requires a data processing agreement and a lawful transfer mechanism. The technical and administrative safeguards under KVKK Article 12 apply to automated document handling just as they apply to any other processing activity. In practice, if a foreign-origin IDP service is used, the firm must determine upfront which data fields are fed into the engine and which are masked before transmission. A national identity number on an invoice can be anonymised before the document leaves the local environment. Skipping this step creates both legal exposure and a future data-cleansing burden that is far more expensive than designing the pipeline correctly from the start.For an SME manager ready to act, a workable framework runs as follows. First, build a document inventory: how many documents per month, how many distinct formats, which languages, and where do they originate? Without this baseline, choosing an IDP engine is guesswork. Second, write the business rules: for each document type, define in a flowchart what the system must do after reading — ‘we already know’ is not a written rule. Third, run a KVKK scan: identify which documents contain personal data and how that data will be handled inside the processing pipeline. Fourth, start a constrained pilot: begin with a single document type that has the highest volume and the most standardised format, typically supplier invoices, rather than the entire portfolio. Fifth, track the exception rate: during the pilot, monitor how frequently the system flags a document as unreadable and routes it to a human. If that rate exceeds 26 percent, the training dataset is insufficient and more labelled examples are needed. The Ankara construction firm skipped all five steps; the Izmir medical distributor completed them. The difference between those two outcomes was not the technology purchased — it was the work done before the technology was switched on.

This article was originally published in Turkish by Gökhan MERCANOĞLU on April 4, 2022. The English edition has been reviewed and edited by the author.


When real-time reporting succeeds, it does not merely put more information on a screen; it gives management clearer decisions. Silos decrease, responsibility becomes visible, and measurable progress starts. Therefore, the issue is not tool selection but rebuilding operating discipline through technology.


Gökhan Mercanoğlu
İş Zekâsı ve Raporlama