职位描述
As the Principal SRE / DevOps Engineer, you will play a foundational role in architecting the future of our platform reliability and operational ecosystem, serving as a technical lead and strategist as we build a robust, scalable, and highly supportable infrastructure to support our clients and products.
This is a hands-on technical position with some project management and leadership responsibilities. In this role, you will work closely with our Software Product, Engineering, and Operations teams to document and evangelize a vision for platform reliability, observability, and automation. You will guide the team to break the high-level vision into well-defined milestones. You will assist in establishing and will champion and participate in best practices in workload and reliability management, including providing visibility to stakeholders and executives for progress toward our vision.
What You'll Do
At Omnidian we believe in trust and autonomy. How you create an impact is ultimately up to you. Here is an outline of the balance of responsibilities:
Architecture Strategy and Vision (35%)
- Formulate the multi-year technical roadmap for our platform reliability, infrastructure, and operational excellence, identifying where intelligent automation can significantly remove engineering and operational bottlenecks (e.g., routine toil, deployment friction, scaling overhead, and incident response).
- Obtain executive and stakeholder buy-in for a documented vision for our SRE / DevOps deliverables. Communicate changes to and/or progress against the vision to executive and technical audiences.
- Translate the technical vision to actionable and trackable work plans, including meaningful milestones. Work with Software Product, Go-To-Market, and other partner teams and stakeholders to align on appropriate milestones and timelines.
- Translate business and reliability requirements to right-sized technical specifications; mentor team members to create technical documentation and clear success criteria (including SLIs/SLOs and error budgets).
- Maintain a security-first mindset, ensuring all infrastructure, automation frameworks, and CI/CD pipelines are robustly defended against operational and security vulnerabilities.
- Provide strong technical leadership and mentorship across teams, establishing clear, practical departmental standards for the safe and ethical use of automation and AI coding/ops assistants.
- Develop an appropriate reliability and testing strategy (leveraging automation and AI where effective for chaos testing, integration validation, and coverage), ensuring LOE estimates include right-sized reliability and observability work.
SRE / DevOps Engineering (35%)
- Design, implement, and maintain production-grade Kubernetes platforms, cluster management, and related orchestration.
- Build and evolve Infrastructure as Code and configuration management using Ansible and Terraform for consistent, repeatable environments.
- Own and continuously improve CI/CD pipelines, deployment strategies, and release automation to enable safe, frequent, and reliable software delivery.
- Implement and refine observability stacks centered on Grafana (along with metrics, logging, and tracing systems) to provide actionable visibility into system health, performance, and reliability.
- Create and maintain automation for operational toil reduction, self-healing systems, and infrastructure provisioning as needed.
- Use AI coding and ops assistants daily for scripting, infrastructure code, refactoring, pipeline improvements, and reliability testing.
Project Management (25%)
- Follow established agile ceremonies, including best practice metrics that are reviewed with the team to support continuous improvement (DORA metrics, SLO attainment, error budget burn, incident metrics, cycle time, technical debt, etc.).
- Create and deliver on visible project plans, break down milestones into distinct work, plan and assign work, and ensure timely and accurate delivery.
- Clearly and proactively communicate progress against the plan including status, blockers, dependencies, and risks.
- Clear blockers to timely or accurate delivery; escalate to leadership as appropriate.
- Identify and proactively communicate work needed from outside the team, obtain commitment, and follow through to ensure dependencies will be delivered to plan.
- Set appropriate documentation expectations for the team (runbooks, architecture decision records, post-incident reviews); ensure this is included in LOE estimates.
Continuous Improvement (5%)
- Identify, align stakeholders on, and implement improvements to our processes, codebases, and architecture.
Who You Are
- You are an exceptional SRE / DevOps engineer who is passionate about reliability, observability, clean automation, robust infrastructure architecture, and production stability. You view generative AI as a pragmatic utility to accelerate engineering and operational velocity.
- You understand that technology best practices and patterns are always shifting, and you love evaluating whether emerging frameworks, tools, or practices should evolve our roadmap.
- You have experience delivering flexible, highly available platforms that provide meaningful reliability and operational insights.
- You possess a security-oriented mindset and deeply understand the security boundaries required when exposing infrastructure and automation tooling.
- You treat everyone with empathy and respect.
- You are a strong communicator, effectively clarifying technical needs, reliability trade-offs, and operational impacts for cross-functional and non-technical partners.
- You excel at balancing rapid responses to real business and operational needs with foundational, long-term architectural and reliability stability.
Experience You'll Need
- 14+ years of experience building, operating, and optimizing large-scale distributed systems, cloud infrastructure, and production platforms.
- Deep hands-on expertise with Kubernetes (cluster architecture, networking, scaling, security, and day-2 operations).
- Strong experience with Ansible (or equivalent configuration management) and Terraform (or equivalent Infrastructure as Code tools) for infrastructure automation, consistency, provisioning, and managing cloud and infrastructure resources.
- Proven ownership of modern CI/CD pipelines, deployment strategies, and release engineering practices.
- Extensive experience designing and operating observability platforms, with strong proficiency in Grafana (dashboards, alerting, and integration with metrics/logging/tracing systems).
- Solid scripting and automation skills (Python, Bash, or equivalent) plus Infrastructure as Code practices.
- Demonstrated ability to establish and drive SRE practices including SLIs/SLOs, error budgets, incident management, and toil reduction.
- Experience using advanced coding/ops assistants extensively for day-to-day automation, infrastructure code, testing, and the ability to establish team standards around their safe and effective use.
- Experience designing systems with appropriate human oversight for automated or agentic operational workflows.
Experience Thats a Plus
- 2+ years of hands-on experience integrating generative AI/LLM components or advanced automation frameworks into production infrastructure or operational systems.
- Familiarity with advanced Kubernetes operators, service meshes, or multi-cluster management.
- Experience implementing tracing frameworks and advanced observability for complex distributed systems.
- Knowledge of event-driven architectures, time-series data, or high-cardinality metrics environments.
- Exposure to IoT telemetry or similar high-volume data streams.
- Solar industry experience
Work-Life & Culture
- We offer a competitive total compensation package that includes monthly health insurance premiums, bonuses and long-term stock options for every employee
- We love to lift each other up through company-wide slack channels such as #puppiesandpets, #omnidian-wellness, #praiseandbooms and #sustainablefuture
- We are a passionate, mission driven team that believes in collaboration, mutual respect and trust. For examples, come Discover our Story!
Grow With Us
- We mentor and invest in our employees and prioritize them for future opportunities. Check out our Instagram reels to see a few career journey examples
- Internal candidates: Check out our advice on Internal Transfer: Job Application Process
- We're a fast-growing startup, which means we're constantly reinventing processes, adding new products, and asking people to use all of their skills and talents. That means there's gonna be a lot of opportunities for you to grow, which also means you will likely be stretched in ways you've never experienced in a job before. If you are resilient, determined, and not afraid of a big challenge, come apply.
$71,000 - $95,000 a year
Our Philosophy on Pay: The range listed above reflects the starting base salary for most new hires in this role. However, your financial growth doesn't stop at the offer letter. We invest deeply in our people and reward results; annual merit increases can bring your total compensation above this initial range. Our goal is for you to continue growing within the pay band as you sharpen your skills and drive meaningful outcomes.
- Committed to Parity: We eliminate the "negotiation tax" by placing candidates within the band based on verified professional experience and skill set, not bartering ability, and by giving our best offer up front. This is a core part of our mission to ensure gender pay equity and fairness across the board.
- Comprehensive Health Coverage: Your well-being is our priority. We cover 100% of health insurance monthly premiums for employees and 50% for your dependents.
- Performance Bonus: Exceptional work deserves exceptional rewards. You'll be eligible for bonuses that reflect your contributions to our collective success.
- Equity Stake: We offer stock options so you can share in the long-term value you help create as we shape the future together.
- Continuous Growth: We are a learning organization. To support your evolution, we provide up to $500 in annual learning reimbursement for courses, certifications, or conferences.
Privacy
California-based candidates: To understand more about the data we collect and process as part of your application, please view our California Job Candidate Privacy Policy. https://www.omnidian.com/privacy-policy-ca-candidates/
Diversity and Inclusion
We strongly believe that diversity of experience, perspectives, and background will lead to a better environment for our employees and a better product for our customers. We are committed to the principle of equal employment opportunity for all employees and to providing employees with a work environment free of discrimination and harassment. We value diversity and inclusion and are committed to ensuring our hiring and retention practices, as well as our office culture, reflects this value.
We will ensure that individuals with disabilities are provided reasonable accommodation to participate in the job application or interview process, to perform essential job functions, and to receive other benefits and privileges of employment. Please contact us to request accommodation.
Omnidian is an equal opportunity employer. We are committed to diversity in the workplace. We make employment decisions on the basis of merit and business need. We hire without consideration to age, ancestry, citizenship, disability, gender expression, gender identity, marital status, national origin, political activity or affiliation, race, religion, sexual orientation, veteran status, or any other basis protected by law.
We invite you to be part of our mission to create a workplace that is inclusive and welcoming to all.
We may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses and identifying potential inconsistencies or verification signals in application materials based on available information. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are ultimately made by humans. If you would like more information about how your data is processed, please contact us.
薪资与福利
职位详情
关于 Omnidian
公司概况
业务介绍
Omnidian是一家技术驱动的服务公司,专注于住宅和商业市场分布式太阳能及电池储能系统的性能保障和资产管理。公司由太阳能行业资深人士创立,利用自主研发的技术提供持续监控和实时诊断,识别性能不足及根本原因,确保能源产出达到最佳水平(来源:omnidian.com)。其服务包括现金返还的能源保障、涵盖硬件和软件组件的全面维护,以及全国范围的服务协调,使Omnidian不仅是软件提供商,更是全方位服务合作伙伴(来源:builtinseattle.com)。公司服务对象涵盖业主、商业企业、太阳能开发商、融资机构、独立电力生产商、企业、房地产投资者及资产管理者,秉持“无忧太阳能”理念,结合主动监控、技术故障排查及维修协调(来源:startupintros.com)。Omnidian管理的能源产能超过2吉瓦,彰显其在清洁能源资产管理领域的重要地位(来源:omnidian.com)。
项目与业绩
Omnidian拥有丰富的分布式太阳能和电池储能资产管理及性能保障经验。其自主技术平台实现持续监控和诊断,快速识别并解决性能不足问题,帮助资产所有者最大化清洁能源投资回报(来源:solarplaza.com)。公司将技术与运营相结合,通过全国网络提供维护保障和维修协调,支持超过2吉瓦的能源产能(来源:omnidian.com)。Omnidian收购了澳大利亚最大的太阳能服务网络Solar Service Guys,标志着其国际业务拓展和服务能力提升的重要里程碑(来源:geekwire.com)。
最新动态
2025年4月,Omnidian完成由B Capital Group领投的8700万美元融资,Activate Capital、Liberty Mutual Investments、National Grid Partners、WIND Ventures、Marunouchi Innovation Partners、BNP Paribas Solar Impulse Venture Fund、Citi Impact Fund及Alumni Ventures等知名投资者参与,体现了投资者对公司成长前景的高度信心(来源:geekwire.com)。此前,2023年10月公司完成了2500万美元的风险投资轮,显示持续的资金支持以推动扩张计划(来源:startupintros.com)。公司计划利用新资金扩大维护服务,拓展澳大利亚等国际市场,并探索电动汽车充电基础设施及商业能源储存等相关清洁能源服务(来源:geekwire.com)。
工作环境
Omnidian招聘涵盖客户支持、技术支持及太阳能专家等多个岗位,体现公司技术与服务运营的结合(来源:omnidian.com)。职位职责包括一级和二级技术支持、故障排查及客户服务,要求具备太阳能运营、现场服务、电池储能及升级管理经验(来源:jobs.bluebearcap.com)。西雅图办公室文化强调深思熟虑的协作、勇气、可衡量的成果及尊重每个人,符合其作为认证B型企业对社会和环境责任的承诺(来源:omnidian.com)(来源:bcorporation.net)。客户服务时间为周一至周五太平洋时间上午7点至下午5点,计划未来扩展至周末,体现对响应支持的承诺(来源:jobs.bluebearcap.com)。
您的 LinkedIn 人脉
在 LinkedIn 上查看您在 Omnidian 的联系人,申请时善用您的人脉。
查看人脉