While synthetic data offers HR an ideal digital twin for innovation, it does not escape the GDPR, and requires strict governance.
The acceleration of digital transformation is pushing human resources departments towards an untenable paradox. On the one hand, People Analytics, predictive recruitment and recruitment modelsartificial intelligence promise to optimize talent management, map skills and anticipate employee departures. On the other hand, the manipulation of this data, which is among the most sensitive in the company such as salaries or evaluations, comes up against an extremely strict regulatory and ethical framework. It is in this context of permanent tension that a technology presented by its promoters as the miracle cure of the modern era emerges: synthetic data. Artificially generated by mathematical models, this information is intended to faithfully imitate certain statistical properties and correlations of real data, without belonging to existing physical individuals. For many decision-makers, the promise is attractive since it offers an ultra-realistic simulation environment for innovation, while theoretically freeing itself from the constraints of the General Data Protection Regulation. However, the technical and legal reality turns out to be much more subtle. Between strategic opportunities, risks of model degradation and strict anonymization criteria, analysis of a double-edged technology.
The legal mirage: Are we really escaping the GDPR?
Faced with these use cases, the strong commercial argument remains legal immunity. The equation seems simple at first glance because the GDPR protects individuals. Since synthetic data is created by an algorithm, it does not correspond to any real individual and should therefore logically escape regulation. This statement, however, deserves to be strongly qualified because the border between the pure virtual and personal data remains porous. To understand this clearly, it is necessary to distinguish anonymization, which irreversibly destroys the link with the individual and falls outside the scope of the GDPR, from pseudonymization, which simply replaces direct identifiers and remains subject to all legal obligations.
According to European doctrine and the CNIL, the key criterion for validating true anonymization is not the absence of a first or last name, but the reasonable impossibility of re-identifying an individual, taking into account all the technical and financial means available. However, if artificial intelligence does overlearning, it risks reproducing combinations of variables so precise that they make it possible, by simple cross-checking, to unmask a real employee. For example, a virtual profile displaying the role of financial director, forty-two years old, based in Lyon and recruited in 2023, becomes immediately identifiable if there is only one person corresponding to these criteria in the company. From then on, the file is reclassified as personal data and the employer is exposed to a major risk of non-compliance. In addition, the initial model training phase, which uses real employee data to initialize the system, constitutes in itself a processing operation that must be legally documented.
However, a pragmatic approach should be adopted since not all synthetic data presents the same level of risk. Some uses are inherently very secure, particularly when it comes to conducting purely technical infrastructure testing, quality control, or working in isolated development environments. In these specific scenarios, the fine mathematical fidelity does not matter, which makes it possible to inject sufficient voluntary statistical noise to cancel out any risk of reidentification without harming the work of the developers.
The three scientific and ethical blind spots
Beyond the purely legal sphere, recent scientific literature highlights risks often underestimated by operational management during the advanced use of these technologies. The first major danger lies in the phenomenon of model collapse, highlighted by researchers at the University of Oxford. When systems are repeatedly trained on data they have generated themselves, their performance degrades irreversibly. The algorithm gets tired, forgets rare cases and ends up over-representing the average. For human resources, this means that if we feed recruitment tools with synthetic profiles in an iterative manner, the system will gradually eliminate singularity, complex life paths and minorities. The long-term risk is to automate profile cloning, destroying cognitive diversity within teams.
The second pitfall concerns the illusion of statistical fidelity. Research shows that synthetic data is not always a perfect mirror of reality. Comparative analyzes carried out on complex socio-economic bases reveal that synthesized versions tend to artificially inflate short-term volatility and distort wage inequalities, particularly on low incomes. Relying on these imperfect simulations to manage a remuneration or talent retention policy can lead to major strategic errors of assessment, by making decisions based on a distortion of market reality.
Finally, the uncontrolled proliferation of artificial content presents a risk of contamination of company databases. If virtual profiles or erroneous predictive scores mix with real records in information systems without perfect traceability, the overall quality of the information collapses. Under European regulations, an employee has the right to understand the logic behind an automated decision that impacts them, such as refusing internal mobility. If the model suffers from a lack of explainability due to erroneous or machine-invented synthetic data, the employer will be unable to justify its decision, opening the way to litigation.
Towards an operational governance charter for HR departments
Synthetic data represents an undeniable technological lever for the future of talent management, but it must be governed by strict and responsible governance to become a safe asset. To move from theory to practice, human resources departments must deploy concrete operational solutions.
It is first essential to adopt metadata marking tools in order to isolate synthetic flows in computer systems, which guarantees compliance with the principle of accuracy by avoiding any pollution of real files. Before sharing a dataset with a third party, the Data Protection Officer should also require mathematical robustness testing using recognized indicators of differential privacy to scientifically validate resistance to re-identification attacks. Finally, synthetic data must be used exclusively for simulation or mass pre-selection. In an internal mobility process, the final evaluation and the decision to assign a position must remain supremely human to promote the nuance and empathy that no machine can copy.