AI-Powered ERP, CRM and BI Selection: The Criteria That Actually Matter

The vendor presentation ended, the screen went dark — and half the evaluation team said ‘impressive.’ I have heard that sentence in dozens of meetings over the past two years. The real question is whether the smooth forecasting engine on that demo screen will perform the same way with your raw data, your broken processes and your actual IT capacity. Most of the time, it will not. When selecting AI-powered ERP, CRM or BI, the decisive factor is not the presence of artificial intelligence but the architecture it runs on, the data it consumes and whether the organisation can know, at any moment, who is accountable for its outputs. This article offers a practical framework for navigating those questions — one that applies both at the procurement table and during pilot design.Let me state the thesis plainly: in 2026, ‘has AI features’ is marketing noise, not a selection criterion. The real criterion is the architecture behind that AI layer and whether it matches the organisation’s data maturity. Consider a packaging materials manufacturer in Kayseri with 312 employees — a representative but realistic profile. The firm received bids from three ERP vendors, all of whom offered ‘AI-powered demand forecasting.’ The first used only the last 24 months of sales data stored in its own database. The second used a RAG (Retrieval-Augmented Generation) architecture that pulled in external raw material price indices alongside internal transaction records. The third described itself as ‘large-language-model-powered’ but could not produce a written specification of which model it used, even during the pilot phase. All three carried the same ‘AI’ label. Without selection criteria, distinguishing between them was impossible.The first evaluation dimension is AI architecture and data connectivity. Ask the vendor directly and technically: ‘Which data does the AI model consume, and how does it keep that data current?’ In RAG-based systems, corporate documents, contracts, inventory history and external price feeds can all be connected to the model — a meaningfully different capability than an LLM running in isolation on its training data alone. But RAG is not magic either. It requires clean underlying data. If stock movements are entered irregularly into the ERP, or if a significant share of records contain manual corrections, a RAG-based forecasting engine will amplify that noise back to you with apparent confidence. How do you measure this? Before the pilot begins, request a data quality score: what percentage of records in the last 12 months required manual intervention? If that figure exceeds twenty percent, fix the data management process first and activate the AI layer afterwards. That sequencing is not optional; it is the difference between a useful system and an expensive one.The second dimension covers AI agent boundaries and accountability mapping. Most ERP and CRM vendors now include ‘autonomous agent’ capabilities — components that can suggest purchase orders, redraw customer segments or update supplier performance scores. Three questions must be answered during procurement. First: does the agent recommend or execute? Recommendation systems carry far lower governance risk than execution systems; an agent that generates a purchase order is an entirely different liability profile. Second: under the EU AI Act, what is the risk classification of this system? Turkish companies that export to the EU or work with EU clients face direct compliance and insurance implications from that classification — and as of 2025 the Act is in phased enforcement. Third: when something goes wrong, where does accountability rest? If vendor contracts do not explicitly address ‘operational losses arising from AI-generated decisions,’ that clause must be negotiated before signature. In a representative scenario at an automotive tier-two supplier in Bursa, an agent-based procurement module stalled at pilot evaluation precisely because this question remained unanswered. The vendor said ‘the system recommends, the human approves’ — but neither the user role nor the approval time window was defined in the system configuration. The pilot was suspended for six weeks, not because of contract renegotiation but because of functional ambiguity that should have been surfaced in the selection process.The third dimension is the SLM option and the data sovereignty calculation. Large cloud providers’ LLM-based ERP add-ons are commercially attractive, but when production secrets and KVKK obligations are involved, what a cloud API query actually transmits becomes a material concern. SLMs — small language models operating in the one-to-thirteen billion parameter range and deployable on local hardware or private cloud — have become a credible alternative since 2025. In an evaluation conducted by a textile exporter in Gaziantep with 295 employees (representative profile), an SLM deployed on-premises made monthly AI infrastructure costs roughly 4.5 times more predictable in Turkish lira terms compared to a token-based cloud billing model, because every currency fluctuation had been feeding directly into the lira equivalent of token charges. The question to ask any vendor is straightforward: does your system support local model deployment, or does it only run within your own cloud environment? If the answer is the latter, that is a dependency cost and it must be entered numerically into your selection matrix, not absorbed as a vague risk note.The fourth dimension is the reliability boundary of AI interpretation in BI. Natural language querying and automated insight features in BI tools genuinely reduce the query burden on analysts — that is a real benefit. But these tools have a limit vendors do not put on slides: the model interprets how data looks, not what data means. In a pilot conducted by a retail chain in Izmir (representative case), the system automatically labelled a sales decline as a ‘seasonal effect.’ The actual cause was a stock gap created by a logistics disruption. The model could not have known this because supply chain data had not been integrated into the BI layer. The lesson is that AI interpretation accuracy scales with the breadth of data integration, not with the sophistication of the model alone. During pilot evaluation of any automated insight feature, deliberately feed the system a known anomaly and observe how it is labelled. That is how you test reliability — not during a vendor demo with curated data.The practical selection framework reduces to five steps. First, ask vendors to document their AI architecture in writing — RAG, fine-tuned model, SLM, or cloud API dependency. A verbal answer is insufficient; request a technical data sheet or architecture diagram. Second, run the pilot on your actual messy data, not the clean sample set the vendor prepared. Third, request written EU AI Act risk classification from the vendor; this document is mandatory for both legal and operational due diligence. Fourth, separate agent capabilities into ‘recommendation’ and ‘execution’; every execution-capable component must have a defined accountability clause in the contract. Fifth, before resolving the data sovereignty question, project token-based costs in Turkish lira terms across a 24-month currency scenario. When the Kayseri packaging manufacturer applied these five steps systematically, the initially favoured vendor was eliminated because no architecture documentation existed. The second candidate survived because an SLM deployment path and a contractual accountability clause were already prepared. Without the framework, neither distinction would have been visible from the demo screen — and an expensive mistake would have been signed with a confident handshake.

This article was originally published in Turkish by Gökhan MERCANOĞLU on June 1, 2026. The English edition has been reviewed and edited by the author.


When reporting infrastructure succeeds, it does not merely put more information on a screen; it gives management clearer decisions. Silos decrease, responsibility becomes visible, and measurable progress starts. Therefore, the issue is not tool selection but rebuilding operating discipline through technology.


Gökhan Mercanoğlu
ERP ve Kurumsal Yazılım