CSI Analytics - Synthetic Data Technical Lead
Job description
At EY, you’ll have the chance to build a career as unique as you are, with the global scale, support, inclusive culture and technology to become the best version of you. And we’re counting on your unique voice and perspective to help EY become even better, too. Join us and build an exceptional experience for yourself, and a better working world for all.
CSI Analytics – Synthetic Data Technical Lead
Job Details
Position Title: CSI Analytics – Synthetic Data Lead
Job Purpose
The Synthetic Data Technical Lead (Assistant Director) will be responsible for establishing, leading, and scaling the Synthetic Data Service capability across the organization. This role will act as the technical authority, service owner, and people leader for Synthetic Data initiatives, ensuring the successful adoption of synthetic data solutions whilst maintaining the highest standards of privacy, security, governance, and data quality.
The successful candidate will combine deep synthetic data expertise with strong leadership and stakeholder management skills. They will be expected to define strategy, operating models, governance frameworks, and delivery standards whilst remaining actively involved in the design, implementation, validation, and continuous improvement of synthetic data solutions.
This is a highly hands-on leadership role. The successful candidate will lead a team of Synthetic Data Technical Analysts whilst working directly with Data, AI, Engineering, Security, Privacy, Risk, and Business stakeholders to deliver innovative synthetic data solutions that enable testing, analytics, AI development, data sharing, performance testing, integration testing, and business innovation.
The role will initially focus on building and operationalizing Synthetic Data generation capabilities using reusable blueprints, Python, SQL, Machine Learning, statistical modelling techniques, and Azure-based engineering patterns. Over time, the role will help evolve the capability toward reusable frameworks, API-driven generation services, and AI-assisted self-service synthetic data generation.
Key Role and Responsibilities
This Synthetic Data Technical Lead will be responsible for delivering scalable, secure, and high quality synthetics data solutions to meet exceptional client and internal stakeholder expectations. Standard work includes but is not limited to:
Synthetic Data Strategy & Service Leadership
- Define and execute the Synthetic Data Service strategy, roadmap, operating model, governance framework, and adoption approach.
- Establish service offerings, delivery methodologies, engagement models, support processes, and success measures.
- Drive enterprise adoption of Synthetic Data by identifying opportunities, demonstrating value, and educating stakeholders on appropriate use cases and benefits.
- Act as the primary point of contact and trusted advisor for Synthetic Data initiatives across business and technology teams.
- Continuously evaluate emerging market capabilities, industry trends, and new technologies to enhance the Synthetic Data service.
Technical Leadership & Solution Delivery
Act as the technical authority for Synthetic Data generation methodologies, data modelling, machine learning approaches, and engineering standards.
- Lead the design and delivery of Synthetic Data solutions using Python, SQL, statistical modelling, rule-based generation, simulation techniques, supervised and unsupervised machine learning models.
- Define reusable generation blueprints, business rule frameworks, schema patterns, validation controls, and automation standards.
- Lead the modelling of facts, dimensions, relationships, distributions, hierarchies, and business logic required to generate realistic Synthetic Data.
- Review and approve generation logic, Python code, validation methodologies, machine learning approaches, and delivery artefacts.
- Drive the evolution of the capability from manually engineered datasets toward reusable frameworks, Azure-based automation, API enablement, and future AI-assisted generation services.
Governance, Privacy & Risk Management
- Establish and maintain governance processes that ensure Synthetic Data solutions meet organizational, legal, regulatory, and security requirements.
- Define standards and controls covering privacy protection, disclosure risk, data quality, model governance, auditability, and responsible AI practices.
- Lead the assessment and validation of Synthetic Data outputs to ensure appropriate levels of data utility, statistical fidelity, business realism, and privacy protection.
- Work closely with Data Governance, Information Security, Risk, Legal, and Privacy teams to obtain approvals and ensure compliance.
- Ensure all Synthetic Data solutions are fully documented, traceable, auditable, and aligned to enterprise governance standards.
Service Operations & Platform Evolution
- Own the operational performance and continuous improvement of the Synthetic Data service.
- Establish service metrics, KPIs, reporting, support processes, and operational controls.
- Manage demand intake, prioritization, resource planning, and delivery governance across multiple engagements.
- Monitor service adoption, customer satisfaction, delivery quality, and business outcomes.
- Support the evolution of the capability from a managed service model toward broader platform ownership, automation, and self-service capabilities.
People Leadership & Capability Development
- Lead, mentor, and develop a team of Synthetic Data Technical Analysts.
- Foster a culture of technical excellence, innovation, collaboration, and continuous learning.
- Support recruitment, performance management, career development, and capability building within the team.
- Provide coaching on Synthetic Data methodologies, technical delivery, stakeholder engagement, and governance practices.
- Promote knowledge sharing and establish a community of practice for Synthetic Data across the organization.
The role will also require the periodic allocation of additional time on the job to support multiple demands, and escalating issues or to accommodate teams or staff in other time zones.
Qualifications, Experience and Skills
Qualifications
- Bachelor’s degree in Computer Science, Data Science, Engineering, Information Systems, Mathematics, Statistics, or a related discipline.
- Master's degree or relevant specialist certifications are advantageous.
Experience
- Synthetic Data Engineering Expertise (Mandatory) – Minimum 7+ Years Practical Experience
- Candidates must demonstrate significant practical experience designing, building, validating, and governing Synthetic Data solutions within enterprise environments.
The successful candidate should have experience in:
- Python and SQL development.
- Data modelling, schema design, and relationship modelling.
- Statistical modelling and distribution analysis.
- Data profiling and data quality assessment.
- Synthetic Data generation methodologies.
- Business rule engineering and rule-based generation.
- Machine Learning and AI concepts.
- Synthetic Data validation and testing.Data privacy, governance, and disclosure risk management.
- Azure-based data engineering and automation solutions.
Technical Skills/Competencies
Expert Level
- Python
- SQL
- Synthetic Data Generation Techniques
- Data Modelling
- Statistical Modelling
- Data Quality & Validation
- Data Privacy & Governance
- Technical Architecture
- Stakeholder Management
Advanced Level
- Machine Learning
- Azure Data Services
- Data Engineering
- API Integration
- DevOps & Automation
- Metadata Management
- Data Product Management
Candidates must demonstrate practical Synthetic Data engineering experience using Python, SQL, statistical modelling, machine learning, business-rule implementation, data modelling, and validation techniques. Experience with commercial Synthetic Data platforms is advantageous but not mandatory.
Behavioral Competencies
- Teamwork
- Integrity
- Empathy & care
- Ownership
- Accountability
- Analytical thinking and problem solving
- Coherent communication
- Time management and organization
- Extreme attention to detail
Additional Candidate Validation Requirement:
Synthetic Data Experience Validation (Mandatory)
Candidates must be able to demonstrate at least two enterprise-scale Synthetic Data implementations where they personally:
- Designed and built Synthetic Data generation solutions.
- Developed Python, SQL, and Machine Learning-based generation logic.
- Defined business rules, schemas, relationships, and distributions.
- Established validation approaches and quality controls.
- Worked directly with Data Governance, Privacy, Security, Risk, and Business stakeholders.
- Delivered Synthetic Data solutions into operational or production use cases.
Candidates should be prepared to discuss:
- Synthetic Data generation methodologies used.
- Python and Machine Learning approaches applied.
- Business rule implementation techniques.
- Validation and quality assessment approaches.
- Data modelling decisions.
- Key challenges encountered and lessons learned.
- Outcomes achieved and business value delivered.
Experience limited to vendor demonstrations, academic research, training courses, proof-of-concepts, or theoretical knowledge will not be considered sufficient.
EY | Building a better working world
EY exists to build a better working world, helping to create long-term value for clients, people and society and build trust in the capital markets.
Enabled by data and technology, diverse EY teams in over 150 countries provide trust through assurance and help clients grow, transform and operate.
Working across assurance, consulting, law, strategy, tax and transactions, EY teams ask better questions to find new answers for the complex issues facing our world today.