EY - GDS Consulting - AIA - Gen AI - Manager
Job description
At EY, you’ll have the chance to build a career as unique as you are, with the global scale, support, inclusive culture and technology to become the best version of you. And we’re counting on your unique voice and perspective to help EY become even better, too. Join us and build an exceptional experience for yourself, and a better working world for all.
Position: Manager
Department: Technology Consulting, AI&A
Title: Enterprise Observability – AI Data Engineer – Manager
Educational qualification: BTech/Masters/PhD
The opportunity
We are seeking an experienced AI Data Engineer Manager with 8–12+ years of professional experience and deep expertise in AI engineering, enterprise observability, telemetry, traceability, and operational monitoring of AI systems. The ideal candidate should be passionate about building reliable, scalable, and governable AI platforms with strong experience in Agentic AI implementations, observability architectures, monitoring frameworks, and enterprise AI operations.
The individual will lead the design, development, implementation, and optimization of observability solutions for AI, Generative AI, Agentic AI, and Agentic RAG systems while collaborating with cross-functional teams, clients, and stakeholders to improve reliability, transparency, performance, and operational readiness across enterprise AI platforms.
Important note: This role requires professionally delivered AI experience in business environments. Personal projects, certifications, hackathons, and tutorial-based work may support an application, but do not replace the required core experience.
Application guidance: Candidates should be able to clearly explain at least two relevant AI implementations, including the business use case, their direct contribution, the technical approach, and the outcome delivered.
Your key responsibilities
- Lead the design and implementation of enterprise-scale observability frameworks for AI, GenAI, Agentic AI, and Agentic RAG systems.
- Define enterprise standards for telemetry, tracing, monitoring, logging, and operational governance of AI solutions.
- Design and implement traceability frameworks to track agent reasoning, tool usage, retrieval paths, model interactions, and workflow execution.
- Build end-to-end observability solutions for AI agents, retrieval systems, APIs, and distributed AI applications.
- Design and implement telemetry collection pipelines to monitor model performance, latency, cost, quality, reliability, and user interactions.
- Develop monitoring and evaluation frameworks for Agentic AI workflows using LangGraph, LangChain, MCP integrations, and Azure AI services.
- Build production-grade observability services, APIs, and monitoring components using Python and FastAPI.
- Design operational dashboards, alerting frameworks, and real-time monitoring solutions for AI workloads.
- Establish evaluation, benchmarking, experimentation, and continuous improvement processes for enterprise AI systems.
- Implement traceability and governance controls to support auditing, compliance, security, and Responsible AI requirements.
- Deploy and manage scalable AI observability solutions on Azure Kubernetes Service (AKS).
- Integrate telemetry data from AI workloads, enterprise applications, APIs, databases, and distributed systems.
- Design data pipelines for collection, processing, and enrichment of observability and monitoring data.
- Collaborate with AI engineers, platform teams, data engineers, and business stakeholders to improve operational visibility and system reliability.
- Create accelerators, reusable observability frameworks, monitoring standards, and operational best practices.
- Stay current with advancements in Agentic AI, observability, telemetry platforms, distributed tracing, AI evaluation frameworks, and emerging technologies.
Skills and Attributes:
Professional Experience
- 8–12+ years of total experience, including 4+ years of directly relevant experience in AI engineering, observability, telemetry, distributed systems monitoring, or AI operations.
- 2+ years of people, technical, or delivery leadership experience.
- Proven experience delivering enterprise-scale monitoring, observability, or operational intelligence platforms.
- Strong depth in at least 3 of the following areas, with direct ownership in at least 2:
- AI observability
- Agentic AI implementations
- Monitoring and telemetry
- Traceability and governance
- Enterprise AI platforms
- Distributed systems monitoring
- AI evaluation frameworks
- Cloud-native architecture
- AI operations and reliability engineering
- Data engineering and operational analytics
- Strong ability to translate operational and governance requirements into scalable technical solutions.
- Comfortable working across engineering, platform, business, risk, and governance stakeholders.
Educational Background
- Bachelor’s/master’s degree in computer science, Data Science, Engineering, or related field.
Technical Skills
- Deep expertise in Agentic AI, Agentic AI Implementation, and Agentic RAG Systems.
- Strong understanding of observability principles, telemetry collection, traceability frameworks, monitoring architectures, and operational analytics.
- Experience establishing end-to-end observability for AI applications, AI agents, retrieval systems, and enterprise AI platforms.
- Experience working with distributed tracing, telemetry pipelines, logging frameworks, and monitoring solutions.
- Strong understanding of AI evaluation, reliability engineering, performance measurement, quality monitoring, and operational excellence.
- Hands-on experience implementing observability frameworks using LangChain, LangGraph, and Model Context Protocol (MCP).
- Experience monitoring agent workflows, tool chains, retrieval paths, model interactions, and multi-agent systems.
- Strong experience with Microsoft Azure AI Platform and Azure AI services.
- Experience designing Azure-native monitoring and telemetry architectures.
- Experience deploying and managing AI workloads on Azure Kubernetes Service (AKS).
- Experience designing scalable Azure architectures for enterprise AI platforms.
- Hands-on experience developing monitoring services, telemetry collectors, and operational APIs using Python and FastAPI.
- Experience integrating telemetry data from distributed applications, AI services, APIs, databases, and cloud platforms.
- Familiarity with MQTT, event-driven architectures, messaging systems, and real-time telemetry frameworks.
- Experience using AI observability and evaluation platforms including Fiddler and similar monitoring technologies.
- Strong proficiency in Python programming for monitoring, data processing, log analysis, telemetry enrichment, and operational analytics.
- Experience implementing alerting frameworks, anomaly detection solutions, dashboarding platforms, and operational reporting.
- Familiarity with MLOps, LLMOps, and AI lifecycle monitoring practices.
- Experience implementing Responsible AI controls, governance frameworks, compliance monitoring, and auditability solutions.
- Strong understanding of security, reliability, scalability, performance optimization, and operational support for enterprise AI systems.
Soft Skills
- Excellent problem-solving skills and ability to connect AI capabilities to business value.
- Strong communication and presentation skills.
Other Responsibilities:
- Lead end-to-end delivery of AI observability and monitoring initiatives, ensuring reliable and scalable AI operations.
- Manage and mentor AI engineers, platform engineers, and observability specialists, fostering a culture of operational excellence and continuous improvement.
- Partner with business, risk, governance, and technology stakeholders to establish enterprise AI monitoring and governance capabilities.
- Ensure AI systems meet reliability, explainability, traceability, security, compliance, and operational requirements.
- Define and govern enterprise standards for AI telemetry, monitoring, evaluation, alerting, and operational reporting.
- Communicate operational health, risks, performance trends, and observability insights to clients and leadership teams.
- Support proposal development, AI governance initiatives, AI platform modernization efforts, and operational transformation programs.
- Drive innovation through reusable observability frameworks, monitoring accelerators, engineering standards, and best practices.
- Contribute to practice growth through thought leadership, capability building, talent development, coaching, and technical mentorship.
Why Join Us
- Be at the forefront of AI-driven innovation across multiple client sectors.
- Work with global clients to drive real business impact.
- Collaborate with a team of AI experts, analytics leaders, and industry specialists in a highly entrepreneurial environment.
What we offer
EY Global Delivery Services (GDS) is a dynamic and truly global delivery network. We work across six locations – Argentina, China, India, the Philippines, Poland and the UK – and with teams from all EY service lines, geographies and sectors, playing a vital role in the delivery of the EY growth strategy. From accountants to coders to advisory consultants, we offer a wide variety of fulfilling career opportunities that span all business disciplines. In GDS, you will collaborate with EY teams on exciting projects and work with well-known brands from across the globe. We’ll introduce you to an ever-expanding ecosystem of people, learning, skills and insights that will stay with you throughout your career.
- Continuous learning: You’ll develop the mindset and skills to navigate whatever comes next.
- Success as defined by you: We’ll provide the tools and flexibility, so you can make a meaningful impact, your way.
- Transformative leadership: We’ll give you the insights, coaching and confidence to be the leader the world needs.
- Diverse and inclusive culture: You’ll be embraced for who you are and empowered to use your voice to help others find theirs.
EY | Building a better working world
EY exists to build a better working world, helping to create long-term value for clients, people and society and build trust in the capital markets.
Enabled by data and technology, diverse EY teams in over 150 countries provide trust through assurance and help clients grow, transform and operate.
Working across assurance, consulting, law, strategy, tax and transactions, EY teams ask better questions to find new answers for the complex issues facing our world today.