CSI Analytics - Synthetic Data Technical Analyst
Job description
At EY, you’ll have the chance to build a career as unique as you are, with the global scale, support, inclusive culture and technology to become the best version of you. And we’re counting on your unique voice and perspective to help EY become even better, too. Join us and build an exceptional experience for yourself, and a better working world for all.
CSI Analytics – Synthetic Data Technical Analyst
Job Details
Position Title : CSI Analytics – Synthetic Data Technical Analyst
Job Purpose
The Synthetic Data Technical Analyst (Senior Associate) will be responsible for supporting the delivery, implementation, validation, and operation of Synthetic Data solutions across the organization. Working as part of the Synthetic Data Service team, the role will focus on generating, testing, validating, and maintaining Synthetic Data assets that support analytics, AI development, testing, performance testing, integration testing, business innovation, and secure data sharing.
The successful candidate will possess strong technical capabilities across Python development, SQL, Data Engineering, Machine Learning, Statistical Analysis, and Data Modelling disciplines. The role requires a hands-on technical practitioner responsible for analysing source datasets, implementing generation logic, validating outputs, and supporting the development of reusable Synthetic Data generation frameworks.
Working under the guidance of the Synthetic Data Technical Lead, the analyst will contribute to the delivery of Synthetic Data engagements from requirements gathering through implementation, validation, documentation, and support.
Key Role and Responsibilities
This Synthetic Data Technical Analyst will be responsible for delivering scalable, secure, and high quality synthetic data solutions to meet exceptional client and internal stakeholder expectations. Standard work includes but is not limited to:
Synthetic Data Solution Delivery
- Analyse source datasets, schemas, relationships, dimensions, distributions, and business rules required for Synthetic Data generation.
- Develop and maintain Python, SQL, and automation scripts supporting Synthetic Data generation and validation.
- Implement rule-based, statistical, simulation, and machine-learning-driven generation techniques.
- Support the creation and maintenance of reusable generation blueprints and framework components.
- Generate and validate Synthetic Data datasets for testing, analytics, AI, and business innovation use cases.
Synthetic Data Engineering & Validation
- Support modelling of facts, dimensions, hierarchies, relationships, and business rules within Synthetic Data solutions.
- Execute data profiling, statistical analysis, distribution testing, and validation activities.
- Validate generated datasets to ensure appropriate levels of data utility, statistical fidelity, business realism, privacy protection, and business-rule compliance.
- Troubleshoot generation issues, data quality problems, and validation exceptions.
- Support development of reusable validation frameworks and automation assets.
Governance, Documentation & Compliance
- Support governance processes to ensure Synthetic Data solutions comply with privacy, security, and enterprise standards.
- Assist with disclosure risk assessments, validation reporting, and governance documentation.
- Maintain technical documentation, delivery artefacts, audit trails, and supporting evidence.
- Ensure all activities align with organizational data governance and compliance requirements.
Service Support & Continuous Improvement
- Support day-to-day operation of the Synthetic Data Service.
- Assist with user support requests, issue investigations, and delivery activities.
- Contribute to reusable assets, templates, accelerators, and best practices.
- Participate in knowledge-sharing initiatives and continuous improvement activities.
- Stay current with emerging Synthetic Data technologies, methodologies, and industry trends.
The role may require periodic allocation of additional time to support critical deliveries, operational activities, or collaboration with global teams operating across multiple time zones.
Qualifications, Experience and Skills
Qualifications
- Bachelor's degree in Computer Science, Data Science, Engineering, Information Systems, Mathematics, Statistics, or a related discipline.
- Relevant Certifications in data Engineering, Analytics, Cloud Technologies, AI, or Synthetic Data Technologies are advantageous.
Experience
- Synthetic Data Engineering Expertise (Mandatory) – Minimum 3+ Years Practical Experience
- Candidates must demonstrate practical experience supporting the implementation, operation, or delivery of Synthetic Data, Test Data Management, Data Engineering, Data Analytics, Machine Learning, Data Simulation, or Data Modelling solutions within enterprise environments.
The successful candidate should have experience in:
- Python development.
- SQL development.
- Data profiling and analysis.
- Data quality assessment and validation.
- Statistical analysis and distributions.
- Business rule implementation.
- Synthetic Data generation techniques.
- Data modelling and schema analysis.
- Machine Learning fundamentals.
Data & Analytics Experience
- Experience working within Data Engineering, Data Analytics, Data Management, AI, or related technical disciplines.
- Experience working with structured and unstructured datasets.
- Experience supporting data transformation, integration, and automation activities.
- Strong understanding of data quality and data governance principles.
Technical Skills/Competencies
Advanced Level
- Python
- SQL
- Data Profiling
- Data Quality & Validation
- Synthetic Data Generation Techniques
- Statistical Analysis
- Data Modelling
Working Level
- Machine Learning
- Azure Data Services
- Data Engineering
- APIs & Automation
- Git & DevOps
- Data Governance
- Metadata Management
Candidates must demonstrate practical Synthetic Data engineering experience using Python, SQL, statistical modelling, machine learning, business-rule implementation, data modelling, and validation techniques. Experience with commercial Synthetic Data platforms is advantageous but not mandatory.
Behavioral Competencies
- Teamwork
- Integrity
- Empathy & care
- Ownership
- Accountability
Analytical thinking and problem solving
- Coherent communication
- Time management and organization
- Extreme attention to detail
Additional Candidate Validation Requirement:
Synthetic Data Experience Validation (Mandatory)
Candidates must be able to demonstrate at least two enterprise-scale Synthetic Data implementations where they personally:
- Built or maintained Python and SQL generation logic.
- Analyzed source schemas, relationships, dimensions, and business rules.
- Generated or validated Synthetic Data outputs.
- Performed data quality, statistical, and validation activities.
- Supported delivery of Synthetic Data, Test Data, Analytics, or AI solutions into operational use cases.
Candidates should be prepared to discuss:
- Synthetic Data generation techniques used.
- Python and SQL development activities performed.
- Validation and quality assessment approaches.
- Business rule implementation.
- Technical challenges encountered and lessons learned.
Experience limited to vendor demonstrations, academic research, training courses, proof-of-concepts, or theoretical knowledge will not be considered sufficient.
EY | Building a better working world
EY exists to build a better working world, helping to create long-term value for clients, people and society and build trust in the capital markets.
Enabled by data and technology, diverse EY teams in over 150 countries provide trust through assurance and help clients grow, transform and operate.
Working across assurance, consulting, law, strategy, tax and transactions, EY teams ask better questions to find new answers for the complex issues facing our world today.