Big Data and the Data Lake: Is Storing Everything Really Smart?

The IT manager of a mid-sized logistics company recently put it plainly: ‘Storage is so cheap now that we keep everything — just in case.’ That single sentence captures the data strategy of a surprising number of businesses today. Hadoop-based data lake solutions are spreading, cloud storage costs per unit continue to drop, and in this environment the instinct to ‘store everything’ is growing stronger. What many managers fail to see is how expensive that instinct can become.

The data lake concept is genuinely appealing in principle: ingest structured and unstructured data in raw form into a central repository, process it when you need it. Unlike a traditional data warehouse, you do not need to transform data before loading it. This flexibility is a real advantage for organizations dealing with high-volume, heterogeneous data sets — transaction logs, machine-generated records, customer behavior streams. But the architecture is a tool, not a strategy. And like any tool used without a clear purpose, it creates problems rather than solving them.

When you run a proper total cost of ownership (TCO) analysis, the picture shifts. Raw storage costs look low on paper, but managing, cataloguing, securing, and keeping that data accessible over time adds significant cost layers. Every gigabyte flowing into a data lake is a potential security exposure, a compliance obligation, and a source of noise. Financial records covered under mandatory e-Invoice and e-Ledger regulations must be retained for defined periods — that is a legal requirement with no room for debate. But applying the same logic to operational log files or temporary reporting data is not intuition; it requires a deliberate policy.

Data quality is the issue that gets overlooked most often. Loading data without a defined purpose leads to what practitioners call a ‘data swamp’: the repository grows, the contents become opaque, and analysts spend increasing amounts of time locating the right data rather than working with it. The core promise of big data — faster insight — reverses itself. As the swamp deepens, analytical capacity shrinks. A retail company that has been accumulating three years of sales data without regular cleansing will find that any customer segmentation analysis built on that data carries a reliability problem that no processing power can fix.

From an ROI perspective, a data lake project generates value only when two conditions are met simultaneously: the intended use of the data is defined before ingestion, and the analytical capacity to process that data already exists within the organization. If either condition is missing, the storage investment produces cost, not value. A significant share of mid-sized Turkish companies are still in the process of building that analytical capacity — recruiting data analysts, establishing reporting processes, learning to ask the right questions of their data. Investing in data lake infrastructure before that foundation is in place is like constructing a building without a load-bearing structure.

The practical difficulty is that a data lifecycle policy is an organizational decision, not a technical one. Determining which data to retain, which to archive, and which to delete cannot be resolved by the IT department alone. Finance, legal, and operations teams all need a seat at the table. Achieving that coordination takes time — particularly in Turkish corporate environments where hierarchical structures are still strong and cross-functional decision-making moves slowly. Cloud storage scalability makes it easy to defer this conversation, and that deferral is precisely the greatest risk in data lake projects today.

If you are a manager evaluating a data lake investment, answer three questions before adding a single gigabyte to the repository: Who will use this data, and to make which specific decision? Does the technical and analytical capacity to process this data exist in-house? What do we actually lose if we do not store it? If you cannot answer those questions clearly, the priority is not building a data lake — it is designing the analytical process that will generate value from the data you already have. Storing everything because storage is cheap is about as efficient a strategy as filling your warehouse because shelf space is affordable.

This article was originally written in Turkish by Gökhan MERCANOĞLU on April 7, 2014 and has been automatically translated into English and other languages using machine translation.


analytical modeling should be designed not to record the company’s past, but to strengthen its future decisions. The right architecture creates visibility, speed, control, and learning capacity. Otherwise, data is collected and reports multiply, while decision quality remains unchanged.


Gökhan Mercanoğlu
Büyük Veri ve Veri Bilimi