Data Steward - Classification & Disclosure Risk
GenerativX
last date
Open Access
Location/Place/Mode
India
Eligibility
Entry-level candidates with strong SQL & hands-on data profiling skills; understanding of data governance, quality, privacy or audit; statistical understanding of cardinality, uniqueness & distributions; strong written English and analytical thinking; ability to work with system/data owners to investigate data. Nice to have: K-anonymity, l-diversity, generalization, suppression, GDPR, Python, entity recognition, data catalogues, audit/risk/fraud investigation or research ethics experience.

Opportunity
The Rising Imperative of Data Stewardship in India's AI Revolution
Over the past twenty-four months, the intersection of artificial intelligence and data protection law has shifted from a niche compliance concern to a boardroom-level priority across the globe, and India is no exception. With the enactment of the Digital Personal Data Protection Act, 2023, and the subsequent tightening of cross-border data flow regulations, Indian enterprises are scrambling to build internal capabilities that can classify, tag, and defuse the risks hidden inside massive datasets. The emergence of generative AI platforms has only amplified the stakes: models trained on undisclosed personal data can inadvertently leak direct identifiers, quasi-identifiers, and sensitive attributes. This is precisely the backdrop against which GenerativX, a forward-looking AI industry player, floated the role of Data Steward – Classification & Disclosure Risk. Although the specific listing may have closed, the architectural shift it represents is permanent.
While the position is technically an entry-level contract job, its strategic weight cannot be overstated. A data steward in this context acts as the moral compass and technical sentinel for a privacy-preserving data platform. They are the ones who stare into the abyss of unfamiliar tables and columns, profiling cardinality and distributions to surface what the organization did not even know it collected. In a journalistic sense, this hiring trend signals a broader realignment: legal compliance is no longer the sole domain of lawyers but a multidisciplinary craft requiring SQL fluency, statistical intuition, and a working knowledge of GDPR-style frameworks. For a country like India, where the data economy is projected to contribute hundreds of billions to GDP, the steward function is nothing short of infrastructural.
"We're Hiring: Data Steward — Data Classification & Disclosure Risk. Are you someone who enjoys investigating data, finding hidden risks, and understanding what data really contains?" – GenerativX Job Posting
Inside GenerativX's Hunt for a Data Steward — Classification & Disclosure Risk
GenerativX, operating at the cutting edge of artificial intelligence, advertised this role targeted at its India wing, though the LinkedIn title initially referenced The Dalles, OR, reflecting the global nature of the team. The position is contract-based, entry-level, and squarely within the Information Technology job function. The core mandate is to own the data inventory across sources, tables, and columns, and to classify direct identifiers, quasi-identifiers, and sensitive data with surgical precision. The successful candidate becomes the backbone of a trusted, privacy-preserving data platform.
The selected candidate will profile unfamiliar datasets to identify hidden personal data, conduct re-identification and disclosure-risk testing, and maintain classification catalogues as schemas evolve. Crucially, the steward will work shoulder-to-shoulder with the Data Protection Officer, documenting classification decisions and their rationale—a task that blends legal documentation with technical metadata management. This fusion of duties explains why the role sits at the crossroads of law, computer science, and audit.
Core Responsibilities That Define the Role
- Own the data inventory across heterogeneous sources, tables, and columns.
- Classify direct identifiers, quasi-identifiers, and sensitive personal data.
- Profile unfamiliar datasets to pinpoint hidden personal information.
- Conduct re-identification and disclosure-risk testing using statistical methods.
- Maintain living classification catalogues as database schemas change.
- Collaborate closely with the Data Protection Officer on governance matters.
- Document classification decisions and the analytical rationale behind them.
- Investigate data alongside system and data owners to resolve ambiguity.
Why This Contract Role is a Launchpad for Legal-Tech Careers
For a fresh graduate or early-career professional, a contract position at an AI-native firm offers an unconventional but high-velocity pathway into the legal-technology ecosystem. Unlike a traditional paralegal role, this job places you at the data pipeline's origin, where privacy engineering meets regulatory compliance. The exposure to concepts such as k-anonymity, l-diversity, generalization, and suppression equips you with a vocabulary that senior privacy counsel respect. Moreover, because the employment type is contract, it allows for rapid portfolio building: you can cite concrete disclosure-risk assessments in future CVs aimed at multinational law firms, Big Four risk advisory practices, or in-house legal teams at fintech unicorns.
Career benefits extend beyond resume lines. Working under a Data Protection Officer provides mentorship in applied privacy law. You learn to translate legal text into SQL queries—a rare skill that commands premium compensation. The contract nature also means you can pivot after six months into a full-time governance analyst role or pursue certifications like CIPP/E or the proposed Indian privacy credential, having already accumulated demonstrable field experience.
"Strong SQL & hands-on data profiling, understanding of data governance, quality, privacy or audit, statistical understanding of cardinality, uniqueness & distributions." – Candidate Profile Sought
Strategic Preparation: Building a CV That Wins the Room
Aspiring applicants should treat the application as a competitor's breakdown: prestige lies not in brand name alone but in demonstrable skill adjacency. To stand out, craft a one-page CV that leads with technical proof. Show a GitHub repo where you profiled a public dataset and flagged quasi-identifiers. Highlight any audit or research ethics experience. Even if you lack Python, emphasize SQL queries that uncovered duplicate records or outlier distributions.
Step-by-step preparation guide for similar openings:
- Step 1: Enroll in a free SQL course and practice joins, group by, and cardinality functions.
- Step 2: Study the DPDP Act 2023 and map its definitions to data fields in a sample database.
- Step 3: Learn the math of k-anonymity; simulate a small re-identification attack on a CSV.
- Step 4: Write a LinkedIn article explaining disclosure risk in simple terms to showcase communication.
- Step 5: Connect with talent leads like Meenakshi Bhardwaj, referencing their hiring domain specifically.
Networking is equally vital. The posting was championed by Meenakshi Bhardwaj, Senior Talent Acquisition Lead with global hiring remit. A direct message that references the specific challenge of re-identification testing—perhaps linking to a short write-up on LinkedIn—can bypass the black hole of online applications. Since the role is no longer accepting applications at the time of writing, proactive networking remains the only viable channel to express interest for future cohorts.
Navigating the Broader Impact on India's Data Economy
The journalistic angle reveals a quiet revolution: data stewards are becoming the unsung guardians of citizen privacy. As India digitizes healthcare, agriculture, and judiciary records, the risk of disclosure through linkage attacks grows. A single quasi-identifier like zip code plus birthdate can re-identify individuals in anonymized datasets. By embedding stewards inside AI firms, the industry pre-empts regulatory penalties and builds public trust. For law students and legal analysts, this is a signal to acquire "quantitative literacy"—a skill set that multiplies career trajectory value.
In summary, while the GenerativX opening may have closed, the template it provides is evergreen. The convergence of AI and privacy law will continue to spawn similar roles. Preparing today by mastering SQL, governance frameworks, and statistical disclosure control ensures you are first in line when the next posting appears. The data steward is not just a job title; it is a frontline defense for the rights of the data subject in the algorithmic age.
Frequently Asked Questions
Q1: What exactly does a Data Steward – Classification & Disclosure Risk do at an AI company?
A1: They inventory data sources, classify personal and sensitive fields, run re-identification tests, and maintain catalogs to ensure the firm's AI training data complies with privacy norms and internal risk thresholds.
Q2: Is this role suitable for a law graduate without coding background?
A2: The role prefers strong SQL and data profiling; however, a law graduate with audit, research ethics, or GDPR knowledge can bridge the gap by learning basic SQL and highlighting privacy governance coursework and analytical writing samples.
Q3: Why is contract employment beneficial for early-career legal-tech aspirants?
A3: Contracts offer rapid exposure to real-world data protection challenges, build a portfolio of classification projects, and often convert to full-time roles or serve as springboards to larger governance positions in multinational firms.
Q4: How can one connect with the hiring lead if applications are closed?
A4: Prospective candidates should reach out to Meenakshi Bhardwaj on LinkedIn with a concise note demonstrating understanding of disclosure-risk testing and sharing any relevant data governance artifacts or personal projects.