In the last week of December 2022, something shifted in my inbox. Within forty-eight hours of ChatGPT’s launch, the tone of messages from Turkish business contacts changed noticeably: ‘Can we do this too?’, ‘Let’s move before our competitors do’, ‘Management approved the budget in today’s meeting.’ A procurement manager at a mid-sized food producer in Konya called to ask for an AI consulting proposal, targeting a February start. When I asked what data they actually had in their ERP and in what shape, the answer was: ‘The accounting module is running. I think inventory is in there somewhere.’ That sentence contains the entire story of the disappointment many Turkish SMEs will accumulate over the next two years. My argument is straightforward: buying generative AI looks far easier than building a data inventory — and that ordering is wrong. Pouring fuel into a car without an engine does not make the car move. It just makes a mess.Let us be concrete about what a large language model actually does in a corporate context. At its core, it performs three functions: it generates documents, answers questions, and extracts patterns from data. All three depend entirely on data that is clean, accessible, and trustworthy. However sophisticated the model’s parameters, if the enterprise data fed into it is inconsistent and fragmented, the output will be too — the ‘garbage in, garbage out’ principle did not become obsolete with generative AI. Looking honestly at the data reality of most Turkish SMEs, the picture is familiar: customer records sitting across three systems in conflicting formats; product names spelled one way in the ERP, another in export documents, and abbreviated differently in accounting entries; sales figures living simultaneously in Excel, in a CRM, and in accounting software, rarely fully reconciled. Pumping AI into this environment does not reduce errors. It accelerates them.Return to the Konya food producer — framed here as a representative scenario, not an identified firm. A company at that scale (312 employees, five production lines) is almost certainly wrestling with at least three of these problems at once: supplier codes that are not standardised in the ERP, so the same supplier appears under four different names; raw material lot tracking that happens on paper forms and gets transcribed into Excel days later; customer complaints scattered across the personal email accounts of the sales team. Give that company the most capable LLM available. The outcome: the model produces inconsistent supplier recommendations because the consolidation was never done; it cannot answer a food safety traceability query because lot data has gaps; customer experience analysis is meaningless because the complaint data is unstructured and incomplete. The model is not at fault. The data is.So what does ‘Data First’ actually mean, without staying abstract? It has three layers. The first layer is a data inventory: a document listing every data source the company holds, who owns it, how often it is updated, and what its quality condition is. A well-maintained spreadsheet is enough to start — no enterprise data catalogue software required at this stage. The second layer is data quality measurement: for a chosen pilot domain, measure record completeness, deduplication rate, and format consistency. In a customer master dataset, what percentage of records have a duplicate? How complete is the address field? Starting an AI project without these measurements is like adding floors to a building before testing whether the foundations hold. The third layer is access architecture: a defined rule for which data can be read by whom, through which tool. This third layer is particularly critical in Turkey in 2023, because KVKK — Turkey’s personal data protection law — does not disappear when you call your project an AI initiative. Sending customer names and order histories to an external LLM API constitutes personal data processing under KVKK, and without a data processing inventory and a valid legal basis, that transfer carries real legal exposure. Most Turkish SMEs have not yet constructed that legal basis.The tension between KVKK and generative AI barely appears in Western narratives about AI adoption. European companies worked through the equivalent GDPR questions from 2018 onward; Turkey arrives at this intersection later, with different institutional dynamics and a smaller pool of specialist legal advice. The first question a Turkish SME should be asking about ChatGPT integration is not ‘What data will I send to the model?’ but ‘What data am I legally permitted to send?’ That question cannot be answered without a data inventory and a documented access architecture. At a textile exporter in Istanbul last year, I watched exactly this collision: the IT consultant had the API integration scoped in five minutes; the legal team froze the project because no data processing agreement existed with the service provider. The bottleneck was not technical readiness. It was data governance readiness. Any AI investment that ignores this friction will eventually hit the same wall — the only variable is how much money will have been spent before it does.Here is a practical starting frame for the Monday morning question. First step: run a half-day workshop with representatives from every department that touches data — finance, sales, operations, procurement. List every data source in the company: ERP modules, Excel files, email archives, paper forms still in use. Note the owner and the update frequency for each. This is your first draft data inventory; it costs nothing and takes one afternoon. Second step: choose the single domain where AI would generate the most obvious value for your business — customer complaints, supplier performance, product descriptions, or something else specific to your sector. Measure the quality of data in that domain only, using three criteria: completeness, consistency, and deduplication. Third step: before sending any data from that domain to an external model, run the plan past a KVKK-aware legal adviser. Does your data processing inventory cover this use case? Do you have a valid legal basis? These three steps do not require a large budget, a specialised IT team, or a technology vendor. They require honest assessment and disciplined sequencing. The company that completes them before buying software will be generating real value six months from now. The company that skips them will have accumulated expensive frustration by the same point. Excitement about ChatGPT is legitimate — the technology is genuinely capable. But capability at the model level and readiness at the data level are two entirely separate questions, and confusing them is the most avoidable mistake a Turkish SME can make right now.
This article was originally published in Turkish by Gökhan MERCANOĞLU on January 23, 2023. The English edition has been reviewed and edited by the author.