Most HR teams treat employee data the way households treat a spare room: whatever might be useful later gets kept, and nothing gets thrown out. Resume PDFs from candidates hired in 2019, biometric attendance logs, exit-interview recordings, engagement-survey exports, screen-monitoring archives, and background-check files all accumulate in the same systems, year after year, because deleting them feels riskier than storing them.
That instinct made sense when storage was close to free. In 2026, it no longer is, and the hidden bill is arriving in three forms at once: rising infrastructure costs, analytics that get worse as the pile grows, and a privacy exposure that expands with every record retained past its purpose.
The operational case for holding less employee data has become concrete, and it starts with a number a founder put on the record.
Storage is No Longer Free, and the Cost Curve Just Turned
The assumption underneath data-hoarding is that keeping a record costs nothing. That assumption broke in 2026. Memory and storage prices have climbed sharply, and the executives running India’s software firms are saying so publicly rather than absorbing it quietly.
Zoho founder Sridhar Vembu, in comments to India Today, tied the squeeze directly to component costs, citing a report that memory prices had risen 500% in twelve months to roughly ten times their lowest level. His words, posted on X, were blunt about the pressure.
“Memory prices, along with AI token prices, have made business very difficult. We have held back from raising prices, but it is becoming hard.”
Vembu’s warning was aimed at the software industry broadly, but the logic lands squarely on HR. Every employee record an organisation keeps sits on infrastructure whose unit cost is climbing, and the volume of that data is growing at the same time.
Payroll history, performance archives, video-interview files, and monitoring logs are not weightless. They occupy paid capacity, they get backed up in duplicate for disaster recovery, and they get replicated across environments for compliance. When the per-gigabyte cost was falling every year, the growth was invisible on the balance sheet. Now the cost curve has turned, and the data that HR never decided to delete has a running meter attached to it.
The teams that have moved beyond spreadsheets into full people analytics platforms feel this first, because those systems ingest and retain the most. Scale multiplies the exposure. A firm keeping five years of granular activity data on 50,000 employees is carrying a very different storage liability from one keeping two.
More Data Makes Analytics Worse, Not Better
The belief that more data automatically produces better insight is the second false assumption, and it fails in practice more often than it holds. Beyond a point, additional employee data does not sharpen analytics. It dilutes the signal, slows the queries, and multiplies the ways an analysis can mislead.
Indian firms are already struggling to extract value from the data they hold. A Deloitte analysis reported that only 8% of Indian organisations have reached maturity in turning people’s data into decisions. Adding more raw data to an already immature capability does not fix the gap. It widens it, because the team now has more noise to sort through before finding anything usable.
The specific ways excess data degrades analysis are worth separating out.
| Problem | What Happens | Effect on HR Decisions |
| Stale data | Records from roles, policies, or org structures that no longer exist stay in the dataset | Trends get pulled toward conditions that ended years ago |
| Duplicate records | The same employee appears under multiple IDs across merged systems | Headcount, attrition, and cost figures come out wrong |
| Irrelevant fields | Data collected “just in case” with no defined use | Models weight noise as if it were signal |
| Volume drag | Query and report times grow as tables balloon | Analysis slows, and teams default to gut feel instead |
A dataset curated to what matters beats a larger one clogged with history that no longer describes the workforce. This is the counterintuitive part of the operational case: the discipline of deleting improves the analytics function rather than starving it. Cleaner inputs produce faster, more trustworthy outputs, which is the entire point of building an analytics capability in the first place.
The India Layer: Fragmented Systems Multiply the Mess
Indian enterprises rarely run one clean HR system. Mergers, GCC expansions, and a decade of stacking point solutions on top of legacy payroll leave most large employers with employee data scattered across half a dozen platforms that do not fully reconcile. Each system holds its own partial copy, and the copies disagree.
That fragmentation is where duplicate records and stale fields breed fastest. When the same person exists as three slightly different records across an old HRMS, a newer analytics tool, and a standalone attendance system, no query returns a clean answer. The organisations getting value from data are the ones consolidating and pruning, not the ones hoarding across silos and hoping a dashboard will paper over the inconsistency.
Every Retained Record is a Record That Can Leak
The third cost is the one that turns a storage decision into a security incident. Data an organisation no longer needs, but still holds, is pure downside. It cannot generate value, and it can still be stolen. The clearest recent illustration involves the country’s largest IT employer.
In a BSE filing dated August 10, 2026, Tata Consultancy Services addressed threat-intelligence alerts after posts on X claimed employee records had been listed on the hacking forum BreachForums, allegedly pulled from the company’s Azure tenant using compromised credentials via password spraying and MFA fatigue. TCS said it found no credible evidence of a breach of its systems or customer environments, and that its safeguards against both attack methods had been in place for more than two years and remained effective.
The TCS employee-data claim recorded the company’s position that the referenced information, if authentic, appeared to be more than four years old and limited to basic employee details such as names, IDs, job titles, phone numbers, and addresses.
The analytical point sits in that detail: the data was allegedly four years old. Whatever its origin, information that far past its operational usefulness was still floating in a form that could be listed for sale. That is the retention problem stated precisely.
A record kept long after its purpose ends stops being an asset and becomes a liability sitting on a shelf, waiting to either cost storage budget or show up in a threat-intelligence feed. Organisations serious about breach-proofing employee records start by shrinking the surface, since the safest record is the one that was deleted on schedule and no longer exists to be stolen.
Retention Now Carries a Legal Meter Too
India’s data-protection regime has removed the option of indefinite retention by default. The Digital Personal Data Protection Act, 2023 sets a purpose-limitation principle: personal data should be kept only as long as the purpose it was collected for remains alive. Once that purpose ends, continued storage is not neutral. It is processing that has to be justified.
The DPDP Act’s storage-limitation provisions push employers toward deleting employee data when its purpose is served, and the framework attaches significant financial penalties to non-compliance. Retention that once looked like harmless caution now reads as a compliance exposure with a statutory number behind it. The four-year-old records in the TCS episode illustrate exactly the category the law is designed to shrink.
Deciding What to Keep and What to Cut
The practical response is a retention discipline that treats every data category as a decision rather than a default. The governing question for each field is simple: what live purpose does keeping this serve, and what is the cost of holding it against that purpose?
A workable approach sorts employee data by whether it still earns its storage.
- Keep with a defined clock: Payroll, tax, and statutory records carry legal retention periods. Store them, and set a deletion date tied to the legal requirement rather than to “forever.”
- Keep only in aggregate: Engagement scores, productivity trends, and workforce patterns rarely need to persist at named-individual level. Anonymise, roll up, and discard the raw.
- Cut on a schedule: Rejected-candidate files, expired monitoring logs, and superseded performance drafts have a natural end date. Automate their deletion instead of leaving it to nobody.
- Never collect in the first place: The cheapest record to store and secure is the one that was never captured. Data gathered “just in case” fails both the cost test and the DPDP purpose test.
Anchoring this to a periodic HR audit and compliance review is what keeps the discipline from decaying. A retention policy written once and never enforced drifts straight back to hoarding within a year or two.
In the End…
The cheapest quarter to cut employee data is this one, before storage contracts renew at 2026 prices and before the next threat-intelligence alert names records you forgot you were keeping. The action is an inventory, not a policy memo.
Pull a full list of every employee-data category your systems hold, and against each one write three things: the live purpose it serves, the legal retention period that applies, and the date it should be deleted. Anything with no live purpose and no legal hold gets a deletion schedule. Anything needed only as a trend gets aggregated and stripped of its raw identifiers. Anything collected “just in case” gets switched off at the source.
The organisations that do this will spend less on storage, get cleaner answers from their analytics, and hand any future attacker a smaller pile to steal. The spare room was never free. It just took a turning cost curve to send the bill.

