Head of EFT Platform I&P
Description
About this Role
Wells Fargo is seeking a Technology Director .
This role will lead the Enterprise Functions Technology (EFT) platform teams in I&P. This leader will own platform reliability, production operations, resiliency, observability, automation, capacity management, operational risk reduction, and AI-enabled operations across one of the firm's most critical payment and money movement platforms.
The role is responsible for transforming traditional production support organizations into a modern engineering-led operating model focused on Site Reliability Engineering (SRE), AIOps, platform observability, automation, resiliency, and operational excellence. The leader will partner closely with Application Development, Platform Engineering, Infrastructure Services, Cybersecurity, Architecture, Product Management, and Enterprise Operations teams to improve reliability, reduce operational toil, accelerate recovery, strengthen resiliency, and improve customer outcomes.
The position leads globally distributed teams responsible for production operations, reliability engineering, platform lifecycle management, performance engineering, capacity planning, vulnerability remediation, and AI operations.
##
In this role, you will
- Manage a team of engineering managers and engineering leads
- Focus on delivering commitments aligned to enterprise strategic priorities
- Build support for strategies with business and technology leaders
- Guide development of actionable roadmaps and plans
- Identify opportunities and strategies for continuous improvement of software engineering practices
- Provide oversight to software craftsmanship, security, availability, resilience, and scalability of solutions developed by the teams or third party providers
- Identify financial management and strategic resourcing
- Set risk management guidelines and partner with stakeholders to implement key risk initiatives
- Develop strategies for hiring engineering talent
- Lead implementation of projects and encourage engineering innovation
- Collaborate and influence all levels of professionals including more experienced managers
- Lead team to achieve objectives
- Interface with external agencies, regulatory bodies or industry forums
- Manage allocation of people and financial resources for Technology Strategic Leadership
- Develop and guide a culture of talent development to meet business objectives and strategy
##
Success Measures
- Improve platform availability and resiliency.
- Reduce customer-impacting incidents.
- Improve MTTR and operational responsiveness.
- Increase automation and reduce operational toil.
- Improve engineering productivity through AI-enabled operations.
- Strengthen operational risk posture and resiliency compliance.
- Build a high-performing leadership team and succession pipeline.
##
Reliability Engineering & Production Operations
- Lead end-to-end production operations for the EFT platform.
- Own platform service availability, operational health, recovery objectives, customer experience metrics, and reliability outcomes.
- Establish and execute a reliability strategy aligned to platform availability, resiliency, and customer experience objectives.
- Establish and maintain service level objectives (SLOs), recovery time objectives (RTOs), recovery point objectives (RPOs), and operational performance targets.
- Lead major incident management, problem management, root cause analysis, and continuous service improvement initiatives.
- Improve platform resiliency and reduce customer-impacting incidents through proactive engineering and automation.
##
Observability, AIOps & Automation
- Establish comprehensive observability capabilities across applications, infrastructure, middleware, cloud platforms, and AI services.
- Drive adoption of metrics, logs, traces, service maps, dependency visualization, and intelligent monitoring.
- Implement AI-powered operational capabilities including anomaly detection, predictive alerting, automated triage, event correlation, and self-healing automation.
- Reduce operational toil through engineering-led automation and platform tooling.
- Drive engineering productivity through AI-assisted operations, intelligent automation, automated diagnostics, operational knowledge management, and developer experience improvements.
##
Platform Lifecycle & Risk Management
- Own platform operational readiness, infrastructure change governance, middleware lifecycle management, patching strategies, certificate management, and vulnerability remediation execution.
- Ensure compliance with enterprise security, resiliency, audit, and regulatory requirements.
- Drive configuration management, drift remediation, and platform hygiene initiatives.
- Partner with application engineering teams to establish self-service operational capabilities, golden paths, platform standards, and reusable engineering services.
##
Performance Engineering & Capacity Management
- Lead performance engineering, workload characterization, forecasting, and capacity planning activities.
- Establish proactive demand and capacity management practices.
- Optimize infrastructure utilization and operational efficiency while maintaining required resiliency standards.
- Partner with engineering teams to improve scalability, efficiency, and cost effectiveness.
- Drive technology efficiency initiatives through infrastructure optimization, automation, cloud cost governance, workload modernization, and operational simplification.
##
AI Operations & Emerging Technology
- Lead AI platform operations supporting GenAI, machine learning, model serving, inference platforms, ModelOps, and emerging Agentic AI solutions.
- Establish operational practices for model monitoring, model drift detection, observability, lifecycle management, and governance.
- Develop capabilities supporting Agentic AI operations and Agentic Reliability Engineering.
- Evaluate and implement emerging AI technologies to improve operational effectiveness and engineering productivity.
##
Leadership & Talent Development
- Lead and develop a globally distributed organization consisting of managers, engineers, reliability engineers, platform specialists, and operations professionals.
- Build a culture of engineering excellence, accountability, continuous improvement, and operational ownership.
- Drive organizational transformation from traditional support functions to an engineering-first operating model.
- Develop succession plans, career paths, and technical growth opportunities for team members.
##
Stakeholder Management
- Serve as the executive point of contact for platform reliability, operational risk, and production stability.
- Influence senior technology leaders and partner organizations to drive reliability and operational transformation initiatives.
- Present platform health, reliability trends, risks, investments, and strategic priorities to executive leadership.
- Influence executive technology strategy and investment decisions through data-driven recommendations, operational insights, reliability trends, and risk assessments.
##
Required Qualifications
- 10+ years of Technology Strategic Leadership experience, or equivalent demonstrated through one or a combination of the following: work experience, training, military experience, education
- 4+ years of management or leadership experience
- 10+ years of technology experience with significant leadership responsibility across software engineering, platform engineering, infrastructure engineering, reliability engineering, or production operations.
- 10+ years of experience leading managers and large-scale global technology organizations.
- Experience leading mission-critical production environments with high availability and resiliency requirements.
- Deep knowledge of production support, incident management, problem management, change management, and operati