Big Data and Data Lakes: What the New Infrastructure Logic Tells Business Leaders

Picture the IT manager of a mid-size manufacturing company sitting in front of three screens. One shows structured sales records flowing out of the ERP system. Another holds a folder of supplier spreadsheets in five different formats. A third displays a growing archive of machine log files that nobody has touched in months. Getting all of this into the company data warehouse means cleaning, transforming, and forcing each source into a predefined schema before a single analyst can run a query. That process, in many organisations, takes longer than collecting the data in the first place. This operational bottleneck is where the big data conversation actually begins.

The data lake concept offers a direct answer to this bottleneck. Traditional data warehouse architecture operates on a ‘schema-on-write’ principle: before data enters the system, its structure must be defined, tables must be created, and relationships must be mapped. A data lake inverts this logic with ‘schema-on-read.’ Raw data lands in a central storage layer without transformation; structure is applied only at the moment of analysis, shaped by the specific question being asked. Distributed file systems such as HDFS, the storage backbone of the Hadoop ecosystem, make this technically feasible at scale. In principle, the IT team no longer needs to fully understand a data source before storing it.

The flexibility promise translates into a concrete operational advantage. Adding a new data source to a traditional warehouse typically means a schema change request, a rewrite of the ETL pipeline, a testing cycle, and a deployment window. This can stretch across weeks. In a data lake, the same task means pointing a new data stream at the storage layer and stopping there. The analytics team shapes the data later, according to their own requirements. Large e-commerce and media companies have moved toward this model precisely because of this speed advantage — the architecture reduces the time between ‘we have data’ and ‘we can ask questions about it.’

That flexibility, however, carries a serious governance cost. The rigid schema of a data warehouse was also a quality guarantee: data entering the system had to conform to defined rules. In a data lake, that guarantee disappears. As raw files accumulate, it becomes progressively harder to determine which file version is current, what a particular log field actually means, or who created a given dataset. The industry term for the end state of a poorly governed data lake is entirely apt: a ‘data swamp.’ Storage fills up, costs rise, and analysts find themselves unable to trust the data enough to act on it. Projects that launch without metadata management, a data catalogue, and access control policies almost always arrive at this destination.

A total cost of ownership analysis complicates the picture further. The initial investment in a data lake architecture looks attractive compared to commercial data warehouse licences — the open-source Hadoop ecosystem is the primary reason for this gap. But when operating costs enter the calculation, the balance shifts. Running distributed systems and writing queries in MapReduce or Hive requires an engineering profile that remains scarce in most markets. That skills gap converts into external consulting fees or extended internal training cycles. For a mid-size company, these hidden cost items can erode the initial licence savings faster than the business case projected.

Practical implementation experience points toward a conclusion that the vendor conversation rarely emphasises: data lake and data warehouse are not competing architectures. They are complementary ones. Raw, exploratory, and unstructured analysis belongs in the lake. Operational reporting, KPI tracking, and board-level dashboards belong in the warehouse. Hybrid architectures that combine both consistently outperform either in isolation. The problem is not the technology — it is the tendency to position the data lake as a complete replacement rather than an extension. Every project that takes this shortcut eventually confronts the swamp.

For an SMB decision-maker, the practical question is this: has the existing data warehouse genuinely reached its architectural limits, or are the ETL processes simply poorly designed? In most cases the answer is the latter, and the right investment is process quality rather than a new infrastructure paradigm. If the organisation genuinely needs to ingest and analyse high volumes of data arriving in varied formats from multiple sources, a data lake is a serious option worth evaluating. But that decision must come with a parallel commitment to budget and headcount for metadata governance, data ownership policies, and access controls. Buying the technology is the easy part. Building the management layer around it is the real investment.

This article was originally written in Turkish by Gökhan MERCANOĞLU on March 19, 2012 and has been automatically translated into English and other languages using machine translation.


If financial bi is approached only as an efficiency agenda, it remains incomplete. Customer experience, employee behavior, financial impact, and operational resilience must be evaluated together. Corporate technology changes not a single department, but the way the whole business operates.


Gökhan Mercanoğlu
İş Zekâsı ve Raporlama