Explore internships, research positions, jobs, scholarships and fellowships from top organizations across India and worldwide.

Graphcore
VerifiedAbout us Graphcore is one of the world's leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world's most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore's teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore brings together deep expertise to solve complex problems and deliver meaningful progress in AI compute. Job Summary Reporting to the Head of Data & Transformation, the Staff Data Engineer – Agentic Applications is a senior individual contributor responsible for designing, building and operating Python-based web applications, agentic workflows and supporting data services for Graphcore’s business teams. The role combines strong data engineering foundations with hands-on application development, using AI agents, trusted data and business system integrations to improve decision-making and automate operational workflows. Depending on the use case, you will either deliver applications end to end, using AI-assisted coding tools to build and maintain the front end, or own the backend and enable business colleagues to create and look after their own front ends. In both approaches, you will help establish clear ownership, secure interfaces and appropriate engineering controls so applications remain reliable, maintainable and supportable. Working closely with engineers, analysts and business stakeholders, you will take solutions from initial requirements through to production, contribute to the evolution of the underlying data platform and provide technical leadership through design reviews, code reviews and mentoring. The Team The Data & Analytics team enables better decision-making across Graphcore by building trusted data foundations, scalable platforms and high-quality data products. The team works across a broad range of business and technical domains, partnering with colleagues throughout the company to improve access to reliable information, strengthen operational insight and support efficient, data-informed ways of working. Within this team, the Staff Data Engineer – Agentic Applications turns these foundations into practical business applications and automated workflows. The role also enables colleagues outside engineering to participate in application development, supported by reusable services, clear guidance and appropriate governance. Responsibilities and Duties Build business-facing applications. Design, develop and operate production-ready Python web applications that help business teams access trusted data, complete tasks and improve operational workflows. Develop agentic capabilities. Build applications and workflows that connect AI agents to approved tools, APIs, datasets and business systems, with clear boundaries on what agents can access and do. Translate business needs into working solutions. Partner with analysts, engineers and business stakeholders to understand user needs, identify appropriate opportunities for automation and take applications from prototype through to supported production services. Deliver end-to-end applications where appropriate. Use AI-assisted coding tools to build and maintain front ends alongside Python backend services, taking responsibility for the quality, testing, security and maintainability of generated code. Enable business-owned front ends. Where business teams develop and maintain their own interfaces, own the backend and provide documented APIs, reusable templates, examples and practical coaching so colleagues can work effectively with AI-assisted coding tools. Establish clear ownership and support arrangements. Agree responsibilities for application changes, testing, releases and ongoing support, including appropriate review and deployment controls for business-maintained front ends. Build secure, reusable backend services. Design APIs and application services that enforce authentication, authorisation, input validation and business rules, keeping secrets and sensitive operations within controlled backend systems. Engineer reliable agentic workflows. Implement appropriate evaluation, monitoring, audit trails and safeguards, including constrained tool permissions, human approval for consequential actions and safe handling of failures or unreliable model outputs. Maintain strong data foundations. Design, build and enhance Python-based batch and streaming pipelines, trusted datasets and reusable data models that support applications, analytics, reporting and operational workloads. Own key platform components. Take ownership of relevant backend and data-platform services, using AWS services including S3, Lambda, Aurora PostgreSQL, Athena, Glue and Redshift to deliver secure, resilient and cost-effective solutions. Enable safe, repeatable delivery. Build and maintain orchestration, CI/CD workflows, automated testing, deployment processes, monitoring and operational support for applications and data workflows. Improve resilience and performance. Apply robust error handling, idempotent processing, retry and recovery mechanisms, and backfill or replay capabilities to improve reliability, scalability and operational performance. Apply engineering and governance standards. Contribute to standards for data quality, documentation, observability, security and access management, including least-privilege access, database permissions and secure secrets handling. Provide technical leadership. Review designs and code, mentor engineers and guide business application builders, helping raise engineering quality across both manually written and AI-generated software. Improve shared capabilities. Identify opportunities to simplify delivery, reuse application patterns and enhance platform capabilities, contributing to technical roadmaps and engineering practices across the Data & Analytics team. Candidate Profile Essential Strong experience designing, building and operating production-grade data pipelines and platforms using Python. Experience building and supporting Python web applications, backend services and APIs that integrate data and business systems. Practical experience developing applications or workflows that use large language models or AI agents, including connecting them to tools, APIs and governed data sources. Experience using AI-assisted coding tools effectively, with the ability to understand, review, test, debug and maintain generated code rather than relying on generated output without validation. Ability to either build usable web front ends with AI-assisted tools or enable business colleagues to develop and maintain front ends against well-defined backend services. Sufficient understanding of web development to review integrations and troubleshoot application behaviour is required; deep specialist front-end expertise is not essential. Strong hands-on experience with modern orchestration, automated testing, CI/CD, deployment and monitoring practices in production environments. Experience building cloud-based solutions using AWS services, including data storage, processing and query technologies. Strong understanding of data modelling, schema design, data quality and performance optimisation across relational and analytical systems. Experience working with both batch and streaming data pipelines, including operational support, troubleshooting and designing systems that recover gracefully from failures. Strong understanding of security, access control and governance for cloud-based data platforms and applications, including authentication, authorisation, IAM, database permissions and secure secrets management. Practical understanding of the reliability and security considerations of agentic applications, including permission boundaries, evaluation, auditability and appropriate human oversight. Experience providing technical leadership as a senior individual contributor through design reviews, code reviews, engineering standards and mentoring. Ability to collaborate effectively with technical and non-technical stakeholders, translate business requirements into practical, scalable solutions and explain technical concepts clearly. Ability to coach business colleagues in AI-assisted application development and establish clear boundaries between business-owned interfaces and engineering-owned services. Desirable Experience with Python web application frameworks such as Streamlit, Flask or FastAPI. Working knowledge of JavaScript or TypeScript and a modern front-end framework. Experience with agent orchestration, retrieval over business data, or automated evaluation of AI-enabled applications. Experience with Prefect or a similar workflow orchestration platform. Experience with streaming or data collection technologies. Experience with PostgreSQL, Redshift, ClickHouse or similar database and data warehouse technologies. Experience with Infrastructure as Code and reusable application deployment patterns. Familiarity with dbt and analytics engineering practices. Experience improving observability, operational monitoring, application performance and cloud cost optimisation. Experience contributing to technical roadmaps and platform improvements within a collaborative engineering team. Experience working within a fast-moving technology or engineering environment. Benefits In addition to a competitive salary, Graphcore offers flexible working, a generous annual leave policy, private medical insurance and health cash plan, a dental plan, pension (matched up to 5%), life assurance and income protection. We have a generous parental leave policy and an employee assistance programme (which includes health, mental wellbeing, and bereavement support). We offer a range of healthy food and snacks at our central Bristol office and have our own barista bar! We welcome people of different backgrounds and experiences; we’re committed to building an inclusive work environment that makes Graphcore a great home for everyone. We offer an equal opportunity process and understand that there are visible and invisible differences in all of us. We can provide a flexible approach to interview and encourage you to chat to us if you require any reasonable adjustments. Applicants for this position must hold the right to work in the UK. Unfortunately at this time, we are unable to provide visa sponsorship or support for visa applications
Graphcore
VerifiedAbout us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Responsibilities and Duties We are seeking a highly skilled System Tests & Diagnostics Engineer to develop, extend, and integrate specialized silicon validation and diagnostics tools for next-generation AI SoCs. Unlike traditional validation roles focused on executing test plans, this position is responsible for developing the diagnostic software and stress tools that expose hardware failures, characterize silicon behavior, and improve platform observability throughout bring-up and validation. You will work closely with Arm engineers to understand and extend existing diagnostics technologies while developing Graphcore-specific capabilities for future AI hardware. Role Summary You will work with existing Arm-developed diagnostics technologies and extend them to support Graphcore's next-generation AI silicon. You will be responsible for developing system-level diagnostics and stress tools that integrate with an existing framework to detect data integrity, computational correctness, performance, and reliability issues across CPUs, AI accelerators, memory, storage, PCIe, firmware, BMC, and other platform components. Examples include silent data corruption (SDC) tests, power transient stress tools, and platform diagnostics, with opportunities to develop new diagnostics as future hardware capabilities evolve. This role requires close collaboration with hardware architects, firmware engineers, validation engineers, and diagnostics teams to understand new hardware capabilities and translate them into reusable diagnostic software. Responsibilities Learn and extend existing Arm diagnostics tools to support Graphcore’s AI platform. Develop specialized diagnostics and stress-testing software for CPU, memory, interconnect, PCIe, and other system components. Build reusable validation utilities that integrate into the existing validation framework. Develop tools that expose intermittent hardware failures including silent data corruption, timing-related issues, and hardware instability. Design configurable diagnostic applications supporting multiple silicon revisions, platforms, and execution environments. Develop automation for executing diagnostics across server and rack-scale validation environments. Design configuration-driven diagnostics supporting scalable parameterization and workload variation. Collaborate with hardware architects to understand new hardware capabilities and identify missing diagnostics. Work closely with firmware and validation teams to improve hardware observability. Debug low-level hardware, firmware, operating system, and driver interactions. Analyze diagnostic output and improve fault isolation methodologies. Develop software that enables engineers to reproduce and investigate difficult hardware failures. Contribute reusable libraries and utilities that improve engineering productivity across multiple validation teams. Candidate Profile Essential Bachelor's or Master's degree in Computer Science, Computer Engineering, Electrical Engineering, or a related technical discipline. 10+ years of experience developing hardware diagnostics, system validation software, firmware validation tools, or low-level systems software. Strong software development skills in Python and C/C++. Strong Linux systems experience. Strong understanding of modern server platform architecture, including how CPU, memory, PCIe, storage, networking, firmware, operating system, and device drivers interact within a complete system. Experience debugging hardware/software interactions across multiple layers of the platform. Experience developing diagnostics, stress testing, or hardware validation software. Experience with server platforms, embedded systems, or SoC validation. Strong analytical and debugging skills. Experience collaborating across hardware, firmware, software, and validation teams. Excellent communication and problem-solving skills. Desirable: Preferred Qualifications Arm-based server platforms AI accelerators or high-performance computing systems Silent Data Corruption (SDC) testing Hardware stress testing Power transient analysis Hardware fault injection methodologies Firmware validation Silicon bring-up Performance characterization Hardware telemetry and observability BMC firmware or server management RAS (Reliability, Availability, Serviceability) technologies Validate across subsystems: CPU scaling and cache behavior Memory (DDR/HBM) bandwidth, latency, and NUMA effects Interconnect contention under multi-core load PCIe/I-O throughput, latency, and multi-device scenarios High-speed I/O validation What Success Looks Like Successful candidates will: Extend existing diagnostics technologies to support new Graphcore hardware. Develop new diagnostics tools for emerging hardware capabilities. Build reusable software components that integrate seamlessly with the existing validation framework. Improve hardware observability and root-cause isolation. Detect and characterize intermittent hardware failures that traditional validation techniques cannot expose. Deliver configurable diagnostics that scale across multiple platforms and silicon revisions. Work effectively across hardware, firmware, validation, and diagnostics teams to continuously improve platform quality. Why Join Graphcore? This role provides a unique opportunity to work at the intersection of hardware architecture, diagnostics, and software engineering while helping define how next-generation AI systems are validated. Rather than maintaining an existing validation framework, you will develop the specialized diagnostics and stress tools that make that framework valuable. Your work will directly influence hardware quality, silicon bring-up, and long-term platform reliability for future AI infrastructure. USA Benefits In addition to a competitive salary, Graphcore offers flexible working and a comprehensive benefits package designed to support your health, wellbeing and financial future. Our benefits include medical, dental and vision coverage, Flexible Spending Accounts (FSAs), Health Savings Accounts (HSAs), disability and life insurance, a 401(k) retirement plan, commuter benefits, wellness services and an Employee Assistance Programme (EAP). We welcome people of different backgrounds and experiences; we're committed to building an inclusive work environment that makes Graphcore a great home for everyone. We offer an equal opportunity process and understand that there are visible and invisible differences in all of us. We can provide a flexible approach to interview and encourage you to chat to us if you require any reasonable adjustments.
Tenstorrent
VerifiedTenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. With AI redefining the computing paradigm, solutions must evolve to unify innovations in software models, compilers, platforms, networking, and semiconductors. Our diverse team of technologists have developed a high performance RISC-V CPU from scratch, and share a passion for AI and a deep desire to build the best AI platform possible. We value collaboration, curiosity, and a commitment to solving hard problems. We are growing our team and looking for contributors of all seniorities. We’re looking for a hands-on Staff Software Engineer with a platform infrastructure / Site Reliability Engineering (SRE) background to help build and operate the systems behind Tenstorrent’s AI cloud. This team develops and runs both on-prem and public cloud infrastructure and services that support high-performance AI workloads. The role spans backend development, service integrations, infrastructure-as-code, CI/CD, and reliability engineering, with room to go deeper in a specialty while working across the stack. This role is hybrid, based out of Austin, TX; Santa Clara, CA; or Toronto, ON. We welcome candidates at various experience levels for this role. During the interview process, candidates will be assessed for the appropriate level, and offers will align with that level, which may differ from the one in this posting. Who You Are Strong software engineer with a track record of designing and delivering complex production services and applications. Fluent in Python, with Go and/or TypeScript as a plus. Experienced with Kubernetes and comfortable working across bare metal, virtual machines, and modern cloud infrastructure; AWS experience is a plus. Familiar with observability and production operations across hardware, system, and application layers, including monitoring, alerting, and telemetry. Bonus points for depth in network automation, storage platforms, and/or cloud or application security. What We Need Build and improve backend services and platform integrations that power Tenstorrent’s on-prem and public AI cloud environments. Raise the bar on infrastructure automation through CI/CD and Infrastructure-as-Code practices using tools such as Ansible and Terraform. Strengthen platform reliability with better observability, monitoring, alerting, and operational response across the stack. Partner effectively with end-users, peers, and domain experts to turn complex infrastructure needs into dependable systems. Bring technical leadership, strong engineering judgment, and a mindset of continuous learning that helps the team scale its capabilities. What You Will Learn How AI cloud infrastructure comes together end-to-end, from data center hardware and fleet operations to user-facing services and interfaces. What it takes to run reliable, high-performance AI workloads across both on-prem and public cloud environments. How hardware, software, networking, storage, and operations intersect inside a real-world AI platform. How to expand into adjacent parts of the stack without getting boxed into a narrow silo. How to apply the latest AI tools in day-to-day engineering work while building practical systems that matter. Compensation for all engineers at Tenstorrent ranges from $100k - $500k including base and variable compensation targets. Experience, skills, education, background and location all impact the actual offer made. Tenstorrent offers a highly competitive compensation package and benefits, and we are an equal opportunity employer. This offer of employment is contingent upon the applicant being eligible to access U.S. export-controlled technology. Due to U.S. export laws, including those codified in the U.S. Export Administration Regulations (EAR), the Company is required to ensure compliance with these laws when transferring technology to nationals of certain countries (such as EAR Country Groups D:1, E1, and E2). These requirements apply to persons located in the U.S. and all countries outside the U.S. As the position offered will have direct and/or indirect access to information, systems, or technologies subject to these laws, the offer may be contingent upon your citizenship/permanent residency status or ability to obtain prior license approval from the U.S. Commerce Department or applicable federal agency. If employment is not possible due to U.S. export laws, any offer of employment will be rescinded.
Graphcore
VerifiedManufacturing Test Engineer – Server Hardware Graphcore is a globally recognised leader in Artificial Intelligence computing systems. The company designs advanced semiconductors and data centre hardware that provide the specialised processing power needed to drive AI innovation, while delivering the efficiency required to support its broader adoption. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. We are opening a new AI Engineering Campus in Austin, which will play a central role in Graphcore's work building the future of AI computing. Role Overview We are seeking an experienced Manufacturing Test Engineer to support high-volume server manufacturing from board-level test through system-level production test. This role will work closely with an ODM manufacturing partner to define, implement, validate, and optimize the manufacturing test strategy for L6 board-level products, including ICT, MDA, and Board Functional Test, as well as support L10 system-level manufacturing test. The ideal candidate has strong experience in server hardware manufacturing, Linux-based test environments, diagnostic test coverage, fixture requirements, yield improvement, and root cause corrective action processes. This role requires both technical depth and hands-on manufacturing execution experience, with the ability to drive best practices across test development, factory readiness, quality planning, and ongoing production support. Key Responsibilities Manufacturing Test Strategy and Planning Work with ODM partners to define and execute the manufacturing test strategy for L6 board-level production. Develop and review test plans covering: In-Circuit Test, or ICT Manufacturing Defect Analyzer, or MDA Board Functional Test Diagnostic coverage requirements Manufacturing line test flow Failure detection and containment strategy Ensure test plans provide appropriate coverage for board-level defects, assembly issues, component-level failures, and functional performance requirements. Partner with hardware engineering, design validation, diagnostics, operations, quality, and ODM teams to align manufacturing test coverage with product risk areas. L6 Board-Level Test Development and Deployment Define requirements for board-level test stations, fixtures, test software, diagnostic content, and production test infrastructure. Support the development, validation, and release of board functional tests into the ODM manufacturing environment. Review ICT and MDA coverage reports and drive improvements to ensure adequate manufacturing defect detection. Define pass/fail criteria, test limits, data collection requirements, retest rules, and failure handling processes. Support bring-up, debug, and qualification of manufacturing test processes during NPI and production ramp. L10 System-Level Manufacturing Test Support development and deployment of L10 system-level manufacturing test processes. Port board functional tests into the L10 manufacturing environment where appropriate. Ensure L10 test coverage validates system-level integration, board functionality, firmware readiness, thermal behavior, power behavior, I/O functionality, and platform health. Work with ODM and internal engineering teams to ensure test execution is scalable, repeatable, and suitable for high-volume server production. Manufacturing Line and Fixture Requirements Specify manufacturing line requirements for test station configuration, test sequencing, data capture, networking, tooling, and operator workflow. Define requirements for test fixtures, cabling, adapters, load boards, debug interfaces, power delivery, signal access, and fixture maintenance. Ensure fixtures and test stations meet manufacturing requirements for reliability, repeatability, safety, ease of use, throughput, and serviceability. Drive fixture validation, correlation, preventive maintenance planning, and readiness for production ramp. Quality Planning, Yield, and Continuous Improvement Create and maintain an overall manufacturing quality plan focused on yield, defect containment, test coverage, and production readiness. Monitor manufacturing test yield, first-pass yield, failure pareto trends, retest rates, false failures, and escape risks. Lead technical investigations into manufacturing test failures, quality excursions, and yield loss. Drive structured root cause analysis and corrective action with ODM partners and internal stakeholders. Define and track corrective actions, containment plans, and long-term process improvements. Establish best practices for test development, test deployment, fixture readiness, failure analysis, data review, and manufacturing quality control. Cross-Functional Collaboration Serve as the primary technical interface between internal teams and ODM manufacturing test teams. Collaborate with hardware engineering, diagnostics, firmware, software, quality, supply chain, and operations teams. Support NPI builds, pilot builds, production ramp, and sustaining manufacturing activities. Communicate test readiness, risks, yield issues, corrective actions, and manufacturing quality status to program stakeholders. Travel to ODM manufacturing sites as needed to support build readiness, test deployment, debug, and ramp activities. DIFFERENTIATORS 8+ years of experience in manufacturing test engineering, hardware test engineering, or production test development. Experience supporting high-volume server manufacturing or similar complex compute, networking, storage, or data center hardware products. Strong understanding of board-level manufacturing test processes, including: ICT MDA Board Functional Test Diagnostic test execution Manufacturing defect detection Experience working directly with ODM, CM, or JDM manufacturing partners. Hands-on experience with Linux-based test environments, including test execution, scripting, log collection, and failure triage. Familiarity with server hardware architecture, including CPUs, memory, storage, networking, BMCs, firmware, power subsystems, and high-speed interfaces. Experience defining test fixture requirements and supporting fixture bring-up, validation, and production readiness. Strong understanding of manufacturing quality metrics, including first-pass yield, retest rate, defect paretos, failure analysis, and corrective action. Demonstrated ability to drive root cause analysis and corrective action across engineering and manufacturing teams. Ability to review test logs, identify failure signatures, isolate issues, and determine whether failures are related to hardware, firmware, software, test process, or fixture design. Strong written and verbal communication skills with the ability to clearly communicate technical issues, risks, and action plans. Preferred Qualifications Experience with L6 board-level and L10 system-level manufacturing processes. Experience porting board-level functional tests into system-level manufacturing environments. Familiarity with server diagnostics, BMC interfaces, BIOS/UEFI, firmware update flows, hardware health checks, and system stress testing. Experience with Python, Bash, or other scripting languages used in manufacturing test automation. Knowledge of manufacturing data systems, test result databases, yield dashboards, and factory analytics. Experience with high-volume NPI, EVT/DVT/PVT, pilot builds, and production ramp. Familiarity with test coverage analysis, DFT/DFM principles, and manufacturing escape prevention. Experience working with global manufacturing teams and offshore ODM sites. Key Success Measures Complete and production-ready L6 test plan covering ICT, MDA, and Board Functional Test. Successful deployment of board functional test into L6 and L10 manufacturing environments. Clearly defined manufacturing line, station, and fixture requirements. Strong diagnostic coverage aligned to product risk and manufacturing defect modes. Stable test processes with low false-failure rates and scalable execution time. Improved first-pass yield and reduced manufacturing defect escapes. Timely root cause identification and corrective action closure for yield and quality issues. Adoption of manufacturing test and quality best practices across ODM production lines. We welcome people of different backgrounds and experiences and are committed to building an inclusive work environment that makes Graphcore a great home for everyone. We are an equal opportunity employer and want to build a work environment where everyone is happy, productive and respectful so they can do their best work. If you have a disability or additional need that requires accommodation, just let us know. USA Benefits In addition to a competitive salary, Graphcore offers flexible working and a comprehensive benefits package designed to support your health, wellbeing and financial future. Our benefits include medical, dental and vision coverage, Flexible Spending Accounts (FSAs), Health Savings Accounts (HSAs), disability and life insurance, a 401(k) retirement plan, commuter benefits, wellness services and an Employee Assistance Programme (EAP). We welcome people of different backgrounds and experiences; we're committed to building an inclusive work environment that makes Graphcore a great home for everyone. We offer an equal opportunity process and understand that there are visible and invisible differences in all of us. We can provide a flexible approach to interview and encourage you to chat to us if you require any reasonable adjustments.
Graphcore
VerifiedAbout us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation. Job Summary We are looking for an experienced Staff Engineer to join our System Management team to help develop the critical interfaces used by internal and external customers to manage system state. Positioned between the hardware and customer workload, we are responsible for building the substrate upon which all other teams build upon. As part of the System Management team, you’ll collaborate across teams to make complex systems work seamlessly and reliably. While your primary focus will be on engineering excellence and system-level development, we value individuals who are versatile and hands-on — willing to contribute wherever needed, providing direct support and troubleshooting as necessary. The ideal candidate brings deep experience with HPC and cloud-native infrastructure, thrives in fast-paced, loosely scoped environments, and leads with action, decisiveness, and teamwork. The Team The System Management team, part of the Software Platform group, helps build Graphcore products into large-scale AI solutions for our customers. The System Management Team is responsible for building the interfaces between hardware and AI software and frameworks as well as providing interfaces for public/private cloud use. This takes the form of a rack management solution that abstracts complex hardware administrative duties. As one of the first teams to work with pre-release hardware and software it’s vital you are comfortable with unproven components and capable of problem solving solutions no matter what. Responsibilities and Duties Ownership of software engineering efforts across the full SDLC, including implementation, automated testing, integration, and production readiness for the rack management solution. Ownership of critical infrastructure with the need to drive issues to resolution while collaborating effectively across teams. Configure and test new Graphcore AI hardware and systems using Continuous Deployment and Infrastructure-as-code in internal and external datacentres. Work with our Datacenter Operations Engineers to maintain and operate the fleet of AI systems at peak performance. Drive corrective actions for systems that are not operating correctly, working with DC operations and Graphcore Engineering as required. Candidate Profile Essential: Bachelor's degree or equivalent practical experience in a relevant subject. Experience with RESTful API development. Experience building, deploying, and operating containerized workloads using Kubernetes and container runtimes such as Docker or Podman. Experience with managing production Kubernetes clusters and workloads. Programming experience with Go. Hands-on experience deploying and operating infrastructure using Infrastructure-as-Code, source code version control, and CI/CD automation tools (e.g. Terraform/OpenTofu, Ansible, GitLab, GitHub Actions, Git version control). Experience with Redfish for datacenter hardware management, telemetry, provisioning, and control. Experience specifying, scoping, estimating and detailing work plans in an AGILE and SCRUM framework, including priorities, risks, issues, impacts and constraints. Strong Linux systems engineering experience, including administration, automation, and scripting with Bash and Python. Desirable Experience with AI coding assistants (Codex, Claude, etc). Experience with Kubernetes operator development (Custom resources). Experience with High Performance Computing (HPC) environments using SLURM or similar batch workload solutions. Experience with virtualized deployments and the technologies they rely on (e.g. Open vSwitch, KVM, QEMU). Experience with distributed object, block, and file storage (e.g., Ceph). Experience in end-to-end deployment automation and CI of containerized services. Complete automation of pipelines for build, test, deploy, manage, alert, destroy, rebuild. Experience with solutions for monitoring and observability (e.g. Grafana, Prometheus, OpenSearch/ElasticSearch, Loki, Mimir, OpenTelemetry, Fluentd, Kafka). Experience with managed switch configuration (e.g. EOS, SONiC, DNOS). Experience with PyTorch for AI workloads. Solid understanding of cloud and infrastructure technologies, including APIs, virtualization, networking, block storage, resource management, and monitoring systems. In addition to a competitive salary, Graphcore offers flexible working and a comprehensive benefits package designed to support your health, wellbeing and financial future. Our benefits include medical, dental and vision coverage, Flexible Spending Accounts (FSAs), Health Savings Accounts (HSAs), disability and life insurance, a 401(k) retirement plan, commuter benefits, wellness services and an Employee Assistance Programme (EAP). We welcome people of different backgrounds and experiences; we're committed to building an inclusive work environment that makes Graphcore a great home for everyone. We offer an equal opportunity process and understand that there are visible and invisible differences in all of us. We can provide a flexible approach to interview and encourage you to chat to us if you require any reasonable adjustments.
Graphcore
VerifiedAbout Graphcore At Graphcore, we’re building the future of AI compute.We’re a team of semiconductor, software and AI experts, with deep experience in creating the complete AI compute stack - from silicon and software to infrastructure at datacenter scale. As part of the SoftBank Group, backed by significant long-term investment, we are delivering key technology into the fast-growing SoftBank AI ecosystem.To meet the vast and exciting AI opportunity, Graphcore is expanding its teams around the world.We are bringing together the brightest minds to solve the toughest problems, in a place where everyone has the opportunity to make an impact on the company, our products and the future of artificial intelligence. Job Summary We are looking for an experiences Staff Engineer to join our Cloud Platform Team and help develop and deploy clouds and services. Working closely with our colleagues in Software Platform, Datacentre Operations and Product Development teams, you will deploy services on our fleet of cutting-edge AI systems. As part of our Software Platform organisation, you will be involved in the cloud integration, validation, performance benchmarking, optimisation, and development of our high-performance AI solutions. These include in-house AI systems alongside off-the-shelf high-performance servers, switches and storage solutions. This is a hand-on technical role requiring a solid background in the use of cloud infrastructure, deployment using Infrastructure-as-Code, observability, high-performance networking and storage systems. You may have been working in an IT organisation, a datacentre, a cloud provider or as a developer of orchestration or cloud services. The Software Platform team at Graphcore We build Graphcore products into large-scale AI solutions for our customers and the Cloud Platform Team is responsible for providing such systems to both internal users via private clouds and customers via our own public clouds. Often the internal systems will be using and developing pre-release hardware and software, so it’s vital you are comfortable with unproven components. Responsibilities and Duties Operate and extend existing OpenStack-based cloud services and contribute to the deployment and development of new ones. Develop and operate end-user services on our clouds and support internal users in their use. You will turn end-user and product requirements into deployed services. Help to build automation to collect and analyse metrics and other observability data from the cloud services to support clear identification and reporting of any issues. Work with users to provide information of any product-related issues to Engineering and QA departments. Work with our Datacentre Operations Engineers to maintain and operate the fleet of AI systems at peak performance in our private clouds. Configure and test new Graphcore AI hardware and systems using Continuous Deployment and Infrastructure-as-code in internal and external datacentres. Drive corrective actions for systems that are not operating correctly, working with DC operations and Graphcore Engineering as required. Work with external vendors of off-the-shelf switches, servers and storage solutions to specify, benchmark and integrate 3rd party products into our Cloud Reference Design. Skills and Experience [ALL REQUIRED] Bachelor's degree or equivalent practical experience in a relevant subject. Solid infrastructure or IT experience with a proven track record of delivering technical output as an individual contributor. Experience managing or operating on-premises or private-cloud environments. Experience specifying, scoping, estimating and detailing work plans in an AGILE and SCRUM framework, including priorities, risks, issues, impacts and constraints. Strong proven Linux scripting ability (bash and python required). Strong proven Linux system administration (Ubuntu, RHEL and variants). Experience with a version control system (preferably Git) and using it to manage system configuration or automation. Experience with Continuous Integration or testing pipelines using GitLab, GitHub or similar. Hands-on experience deploying services into public or private clouds using Infrastructure-as-Code (IAC). A solid understanding of the technologies underpinning cloud services (APIs, virtualisation of CPUs, IO, systems), virtual networks, block storage, resource management and monitoring. Experience with IAC automation tools (e.g. Terraform/OpenTofu, Ansible, Packer). Experience with container deployment and management tools (e.g. docker, podman, apptainer). Experience with solutions for monitoring and observability. e.g. Grafana, Prometheus, OpenSearch/ElasticSearch, Loki, Mimir, OpenTelemetry, Fluentd ,Kafka Good communication and presentation skills, and experience dealing with end-users of IT or cloud services. An ability to work independently on critical infrastructure without oversight, and with a focus on end-user availability. Desirable but not required: Experience with OpenStack deployments or the technologies they rely on (e.g. Ceph, Open vSwitch, KVM, QEMU ). Experience with High Performance Computing (HPC) environments using SLURM or similar batch workload solutions. Strong skillset and experience in end-to-end deployment automation and CI of containerised services. Complete automation of pipelines for build, test, deploy, manage, alert, destroy, rebuild. Experience with managing production Kubernetes clusters and workloads. Experience with workload queue management systems (SLURM, LSF, Kueue). Experience with managed switch configuration (e.g. EOS, SONiC, DNOS). Programming experience with Python3 utilising classes and inheritance. Programming experience with Go. Benefits In addition to a competitive salary, Graphcore offers flexible working, a generous annual leave policy, private medical insurance and health cash plan, a dental plan, pension (matched up to 5%), life assurance and income protection. We have a generous parental leave policy and an employee assistance programme (which includes health, mental wellbeing, and bereavement support). We offer a range of healthy food and snacks at our central Bristol office and have our own barista bar! We welcome people of different backgrounds and experiences; we’re committed to building an inclusive work environment that makes Graphcore a great home for everyone. We offer an equal opportunity process and understand that there are visible and invisible differences in all of us. We can provide a flexible approach to interview and encourage you to chat to us if you require any reasonable adjustments. Sponsorship Applicants for this position must hold the right to work in the UK. Unfortunately at this time, we are unable to provide visa sponsorship or support for visa applications.
Graphcore
VerifiedAbout us We are looking for a disciplined and dynamic Systems Engineer with focus on server CPU based system to join our growing compute rack validation team. Candidate we are seeking should have demonstrated work-experience in leading server rack and blade hardware systems deployment, hardware installation, and inventory management activities in the Austin, TX area. As a diligent leader in Systems Engineering, you will drive multiple aspects of post-silicon validation throughout the life cycle of the program. In this high visibility position, you will be part of a technical team chartered to innovate and improve system bring-up and enablement capabilities, as well as silicon and system validation to deliver the highest quality, industry leading technologies to market. Your technical leadership skills, systems engineering and hardware bring-up, validation and debug expertise will be necessary towards product development, definition, root cause and resolution. Your agility and collaborative approach will be essential to work within System Validation & other engineering teams (System Architects, SoC and Rack FW etc). The ideal candidate will be driving key areas around at-scale system validation including ARM based server and rack level systems bring-up (nodes and rack level systems). Candidate will be immersed in challenging system enablement work, ramp-up post-silicon capabilities in engineering lab environments, validation tests execution/triage. The candidate will be leading contributor towards state-of-the-art HW bring-up and lab capabilities for Grapchore’s system engineering. The candidate should be able to work in a global environment while maintaining a synergetic culture. Primary Responsibilities: Install, configure, commission (and decommission if needed) blade servers, chassis, switches, and supporting infrastructure. Lead rack and stack activities, including mounting equipment, cable management, and labelling. Execute hardware upgrades, replacements, and troubleshooting of server and network components. Maintain accurate asset records within DCIM platforms and inventory management systems. Conduct physical audits and reconcile inventory discrepancies. Track hardware movements, deployments, and decommissions through established change management processes. Document installation procedures, rack layouts, cabling diagrams, and inventory updates. Support data center migration, expansion, and refresh projects. Collaborate with engineering, operations, logistics, and project management teams. Adhere to all data center safety, security, and operational standards. Develop, setup and scale key methodologies for at-scale test execution, lab HW and system SW capabilities as well as system visibilities and debug tools necessary for successful system (HW/SW/FW) bring-up and system validation at blade and rack level for AI compute rack. Ability to work independently in a production ready environment, and a commitment tomaintaining accurate inventory and asset records. Triage issues found during server rack validation bring-up, Post-Silicon Validation, and production phases of the program. Ensure issues are solved on time with quality. Lead test execution of key domains within AI compute solutions like CPU, GPU, memory, HBM, IO etc. Drive technical innovation to improve capabilities across system validation, including tools, script development, technical and procedural methodology enhancement, and various internal and cross-functional technical initiatives. Qualifications: Strong analytical/problem-solving skills and pronounced attention to details Experience in Blade server installation and maintenance (Cisco UCS, HPE Synergy, Dell MX, or similar). Rack and stack deployments in enterprise or hyperscale environments. Copper and fiber cabling installation and management. DCIM and asset management platforms. Strong understanding of server, storage, and networking hardware. Experience performing inventory audits and maintaining asset accuracy. Ability to read rack elevation diagrams, cabling schematics, and deployment documentation. Familiarity with ticketing and change management systems. Exposure to Linux (ubuntu) OS bootable images and system firmware basics for image building, provisioning and firmware flashing. Exposure to automation testing, to enable execution of hardware acceptance tests, best-known-config testing etc. Exposure to python script development and execution. Proven experience in understanding, defining and enabling storage (storage rack), networking capabilities (network rack, DNS, DHCP etc) in a lab environment to help add end-to-end validation and debug capabilities for rack and blade validation. Excellent communication and coordination skills. Detailed oriented, highly organized, able to prioritize, and juggle multiple work streams to tight deadlines. Technical leadership: capable of championing new tools, methods, and capabilities to drive platform validation improvements in schedule, quality, or coverage. Experience working with data center technical staff, 3rd party vendors, ODMs etc throughout the life cycle of server system product development. Must be a self-starter, and able to independently drive tasks to completion Preferred Qualifications: Masters or PhD in Electrical Engineering, Computer Engineering or a related field. 10+ years of work experience demonstrating working on complex systems engineering challenges to validate and debug HW-FW-SW challenges in a server compute rack or data center blade environment. Experience designing and deploying modern AI/ML rack scale systems Knowledge of industry standards and best practices for hardware development Familiarity with emerging technologies in AI and Data Center infrastructure. Comfortable meeting, engaging and collaborating with ODM partners and staffing vendors across the globe. USA Benefits In addition to a competitive salary, Graphcore offers flexible working and a comprehensive benefits package designed to support your health, wellbeing and financial future. Our benefits include medical, dental and vision coverage, Flexible Spending Accounts (FSAs), Health Savings Accounts (HSAs), disability and life insurance, a 401(k) retirement plan, commuter benefits, wellness services and an Employee Assistance Programme (EAP). We welcome people of different backgrounds and experiences; we're committed to building an inclusive work environment that makes Graphcore a great home for everyone. We offer an equal opportunity process and understand that there are visible and invisible differences in all of us. We can provide a flexible approach to interview and encourage you to chat to us if you require any reasonable adjustments.
Graphcore
VerifiedWe are looking for a disciplined and dynamic, Lead System Engineer – compute blade and rack Validation to join our growing compute rack validation team. As a diligent leader in Systems Engineering, you will drive multiple aspects of validation throughout the life cycle of the program. In this high visibility position, you will be part of a leading team to innovate and improve system bring-up and enablement abilities, as well as silicon and system validation to deliver the highest quality, industry leading technologies to market. Your technical leadership skills, validation and debug expertise will be necessary towards product development, definition, root cause and resolution. Your agility and collaborative approach will be essential to work within System Validation & other engineering teams (System Architects, SoC and Rack FW etc). The technical leader will be driving keys areas of system validation including leading first silicon & system bring-up (nodes and rack level systems) - rack level systems and blades will be based of ARM server architecture. Candidate will be immersed in challenging system enablement work, system validation (end-to-end) methodology, tests development and execution as well as triage/debug of critical issues to meet critical program milestones at POR quality. The candidate will also be a key contributor to state-of-the-art HW and lab capabilities for Grapchore’s system engineering. The candidate should be able to work in a global environment while maintaining a synergetic culture. Primary Responsibilities: Lead the systemenablement (including first silicon and other FW components) to ensure system capabilities are brought up as per plan of record and system architecture spec. Drive organization wide methodology for Firmware integration and best known configuration (HW/FW/SW) usage model by leading the release of deployment ready solutions. Develop key methodologies, lab HW and system SW capabilitiesas well as system visibilities and debug tools necessary for successful system (HW/SW/FW) bring-up and system validation at blade and rack level for AI compute rack. Triageissues found during server rack validation bring-up, Post-Silicon Validation, and production phases of the program. Ensure issues are solved on time with quality. Develop and own test plans,lead test case development and execution of key domains within AI compute solutions like CPU, GPU, memory, HBM, IO etc. Drive technical innovation to improve capabilities acrosssystem validation, including tool, script development, technical and procedural methodology enhancement, and various internal and cross-functional technical initiatives. Qualifications: Strong analytical/problem-solving skills and pronounced attention to details Extensive experience in validation roles involvingfirst silicon and system bring up, OS, FW, Silicon, and HW Understanding of PC industry standardbuses and their software stack, such as PCIe, CXL. Proven experience in understanding,defining and enabling storage (storage rack), networking capabilities (network rack, DNS, DHCP etc) in a lab environment to help add end-to-end validation and debug capabilities for rack and blade validation. Strong knowledge ofARM CPU or X86 architecture, SoC design, memory, RAS & power management as well as HW/SW based tools, and revision control systems. Extensive knowledge of system architecture, technical debug, and validation strategy Good understanding and experience in platform/ system level debug, Operating System, DeviceDrivers and System BIOS interactions. Excellent communication and coordination skills. Detailed oriented, highly organized, able to prioritize, and juggle multiple workstreams to tight deadlines. Technical leadership: capable of championing new tools, methods, and capabilities to drive platform validation improvements in schedule, quality, or coverage. Deep experience with Linux and Windows Operating Systems, Hypervisors (VMware, KVM, Hyper-V, etc.), and development and certification processes of these environments Experience in technical program management. A thorough understanding of datacenter industry technologies and their software stack. Must be a self-starter, and able to independently drive tasks to completion Preferred Qualifications: Mastersor PhD in Electrical Engineering, Computer Engineering or a related 14+ years of work experience demonstrating working on complex systems engineering challenges to validate and debug HW-FW-SW challenges in a server compute rack or data center blade environment. Experience designing and deploying modern AI/ML rack scale systems Knowledge of industry standards and best practices for hardwaredevelopment Familiarity with emerging technologies in AI andData Center Comfortable meeting,engaging and collaborating with ODM partners across the globe. USA Benefits In addition to a competitive salary, Graphcore offers flexible working and a comprehensive benefits package designed to support your health, wellbeing and financial future. Our benefits include medical, dental and vision coverage, Flexible Spending Accounts (FSAs), Health Savings Accounts (HSAs), disability and life insurance, a 401(k) retirement plan, commuter benefits, wellness services and an Employee Assistance Programme (EAP). We welcome people of different backgrounds and experiences; we're committed to building an inclusive work environment that makes Graphcore a great home for everyone. We offer an equal opportunity process and understand that there are visible and invisible differences in all of us. We can provide a flexible approach to interview and encourage you to chat to us if you require any reasonable adjustments.
Graphcore
VerifiedSalary Range: PLN 260,400 - 352,200 + Benefits + Equity Subject to alignment to the responsibilities and duties of the role. Location: Gdańsk - Hybrid Working Policy - 2-3 Days per week in office About Graphcore At Graphcore, we’re building the future of AI compute.We’re a team of semiconductor, software and AI experts, with deep experience in creating the complete AI compute stack - from silicon and software to infrastructure at datacenter scale.As part of the SoftBank Group, backed by significant long-term investment, we are delivering key technology into the fast-growing SoftBank AI ecosystem.To meet the vast and exciting AI opportunity, Graphcore is expanding its teams around the world.We are bringing together the brightest minds to solve the toughest problems, in a place where everyone has the opportunity to make an impact on the company, our products and the future of artificial intelligence. Job Summary As a Senior QA Engineer within the Management & Observability team, you will be responsible for validating Graphcore's end-to-end telemetry and observability platform. Working closely with Telemetry and Observability engineers, you will design, develop and automate comprehensive test strategies covering telemetry generation, data collection, processing, storage, visualization and alerting. Your work will ensure that Graphcore's observability solutions are reliable, scalable and production-ready for both internal engineering teams and customers. You will contribute throughout the software development lifecycle by defining quality standards, building automated test frameworks and validating distributed systems operating at scale. Responsibilities and Duties Define and implement the end-to-end quality strategy for Graphcore's telemetry and observability platform. Design, develop and maintain automated functional, integration, system and regression tests covering the complete telemetry lifecycle - from telemetry generation to dashboards, APIs and alerting. Develop automated validation frameworks for telemetry pipelines, data quality, metrics, logs, traces and time-series data. Design realistic test environments capable of validating large-scale distributed deployments and production-like workloads. Work closely with software engineers throughout design and implementation to ensure testability, reliability and quality are built into every component. Validate performance, scalability, resilience and fault recovery of telemetry and observability solutions under realistic operating conditions. Integrate automated testing into CI/CD pipelines and continuously improve test coverage, execution time and release quality. Investigate defects through root-cause analysis, working with engineering teams to resolve complex system-level issues. Develop quality metrics, test reports and release readiness criteria to support engineering and product decisions. Contribute to continuous improvement of testing methodologies, automation frameworks and engineering best practices. Skills and Experience BSc or MSc degree in Computer Science, Computer Engineering or equivalent practical experience. 5–8 years of experience in Software QA, Test Automation or Software Engineering. Experience designing automated test frameworks for distributed systems. Experience testing cloud-native or infrastructure software running on Linux. Experience with Python programming. Experience building automated integration and system tests. Familiarity with CI/CD platforms and automated testing pipelines. Experience with Kubernetes, Docker and containerized environments. Understanding of distributed systems, networking and API testing. Experience validating REST and gRPC APIs. Strong debugging and root-cause analysis skills. Excellent written and verbal communication skills. Desirable: Experience testing observability platforms based on Prometheus, Grafana, OpenTelemetry, ClickHouse, Kafka or Elastic Stack. Experience validating telemetry pipelines and large-scale time-series data. Experience with performance, scalability and resiliency testing. Familiarity with Infrastructure as Code technologies such as Terraform or Ansible. Experience testing AI infrastructure, HPC platforms or cloud infrastructure. Experience with one additional programming language such as Go or C++. Knowledge of modern observability practices including metrics, logs and distributed tracing. Benefits In addition to a competitive salary, annual leave policy, medical and dental health plans, a gym card and employee pension (matched up to 4%). We review our benefits on a yearly basis to ensure we offer a valuable and rewarding benefits programme to our employees. We welcome people of different backgrounds and experiences; we’re committed to building an inclusive work environment that makes Graphcore a great home for everyone. We offer an equal opportunity process and understand that there are visible and invisible differences in all of us. We can provide a flexible approach to interview and encourage you to chat to us if you require any reasonable adjustments. Sponsorship Applicants for this position must hold the right to work in the Poland. Unfortunately at this time, we are unable to provide visa sponsorship or support for visa applications.
Graphcore
VerifiedAbout Us Graphcore is a globally recognised leader in Artificial Intelligence computing systems. The company designs advanced semiconductors and data centre hardware that provide the specialised processing power needed to drive AI innovation, while delivering the efficiency required to support its broader adoption. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. We are opening a new AI Engineering Campus in Austin, which will play a central role in Graphcore's work building the future of AI computing. Job Overview: Graphcore is seeking a Principal Engineer, Power Engineering to aid in defining the architecture of the power delivery components from grid to chip. You will work closely with multiple cross-functional teams, including hardware engineers, firmware developers, and data center operations, to ensure the infrastructure we build meets customer requirements as well as attaining standards of performance, reliability, and scalability. This role will be based in Austin or San Jose and could potentially be performed onsite or remotely within the US. Responsibilities: Provide thought leadership, planning, and design for power delivery from data center power drop through delivery to systems & chips Establish power roadmap. Work with vendors to ensure their planned future technologies are pulled into our system engineering roadmap AND ensure that our (system engineering) required technological innovation is being driven into their (power vendor) roadmap. Work to improve reliability and efficiency of power engineering component over time. Rack-level power components. Competency with specification of hardware and firmware for rack-level power shelves. Competency with design for resiliency. Set power policies in the data center for: power buffer/oversubscription, scheduling of jobs based on power, and monitoring (power quality faults) / accommodation faults. Resolve telemetry and data needed to support power policy algorithms. Ensure robustness of IT power in the presence of grid disturbances and make sure that IT power loads are compatible with grid power requirements (“good grid power citizen”) Board and system-level power regulation components. Understand and able to apply current VRM, TLDR, and POL technologies. Help to develop new technologies to increase efficiency. Understand manufacturing capabilities and limitations of new technology. Drive Provide guidance to the power equipment manufacturers, ODM and engineering teams on the required testing and validation procedures for power hardware and firmware components. Develop and test hardware prototypes as needed to support the exploration and validation of new concepts in power delivery. Tackle and resolve sophisticated power hardware and firmware issues as needed. Ensure compliance with industry standards and regulations. Required Skills and Experience : Bachelor’s or Master’s or PhD degree in Electrical Engineering with a specialized knowledge of Power Engineering. demonstrated ability in power system and delivery architecture, design, and development. Experience with rack-level hardware design, including servers, storage, networking, and power distribution. Proficiency in power hardware and firmware development and debugging. Excellent problem-solving and analytical skills. Strong communication and teamwork skills. Ability to work in a fast-paced, multifaceted environment. "Nice to Have" Skills and Experience: Familiarity with new technologies in AI and data center infrastructure. Comfortable meeting with and engaging directly with customers (internal and external) during requirements gathering and solution development Benefits In addition to a competitive salary, Graphcore offers a competitive benefits package. We welcome people of different backgrounds and experiences; we’re committed to building an inclusive work environment that makes Graphcore a great home for everyone. We offer an equal opportunity process and understand that there are visible and invisible differences in all of us. We can provide a flexible approach to interview and encourage you to chat to us if you require any reasonable adjustments.
Graphcore
VerifiedAbout us Graphcore is one of the world’s leading innovators in artificial intelligence compute. We are developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and support the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of a family of companies responsible for some of the world’s most transformative technologies. Together, we share a bold vision to enable advanced artificial intelligence and ensure its benefits are accessible to everyone. Graphcore brings together AI researchers, silicon designers, software engineers and systems architects to solve complex technical challenges and deliver innovative computing solutions. Job Summary The Principal Electrical Engineer will be a technical authority within Data Center Engineering, leading the architecture and delivery of safe, resilient and scalable electrical infrastructure for high-density AI computing environments. Working with internal teams, data center developers, utilities, consultants and equipment partners, this role will guide projects from early technical studies through design, construction, commissioning, operation and lifecycle improvement. The successful candidate must reside in, or be willing to relocate to, Austin, Texas. Approximately 10% travel may be required. The Team The Data Center Engineering team is responsible for defining and enabling the infrastructure needed to deploy and operate Graphcore’s computing systems at scale. The team works across electrical, mechanical, thermal, controls, systems and operational disciplines, collaborating with external engineering and construction partners to deliver reliable, efficient and maintainable data center environments. Responsibilities and Duties Act as the technical authority for electrical engineering across data center infrastructure projects, from the utility or on-site power source through to the IT rack. Lead electrical architecture development through conceptual studies, detailed design, construction, commissioning, operation and lifecycle improvement. Develop electrical standards, specifications, reference designs, qualification plans and acceptance criteria for power distribution and conversion systems. Define electrical protection and safety approaches for AC and DC systems, including fault detection, interruption, isolation, grounding, stored energy, insulation monitoring, lockout and safe maintenance. Develop resilient power solutions for highly dynamic AI and high-performance computing loads, including ride-through, short-duration energy buffering, backup generation, energy storage and utility interconnection. Evaluate established and emerging power conversion, protection, control and energy-storage technologies, making clear recommendations based on safety, reliability, efficiency and maintainability. Lead technical due diligence, design reviews, failure-mode assessments, factory and site acceptance testing, site inspections and commissioning activities. Direct engineering consultants and technology partners, ensuring that designs, calculations, specifications and construction documentation meet project requirements. Develop and manage owner’s project requirements and review consultant deliverables, submittals, requests for information and proposed technical changes. Collaborate with construction partners during the delivery of laboratory and data hall spaces, supporting site reviews, technical issue resolution and project risk management. Partner with teams across IT, silicon, mechanical, thermal, controls, operations and sustainability to resolve cross-disciplinary design challenges. Support operational readiness, incident investigation and root-cause analysis, identifying practical improvements to reliability, safety and system performance. Engage effectively with authorities, utilities and relevant industry bodies to support compliant project delivery. Maintain current knowledge of developments in AI and high-performance computing infrastructure, electrical equipment, engineering standards, codes and regional utility requirements. Mentor engineers and provide technical guidance across projects without direct people-management responsibility. Candidate Profile Essential Bachelor’s or master’s degree in Electrical Engineering, a closely related discipline, or equivalent relevant experience. Substantial experience delivering mission-critical electrical infrastructure in hyperscale data centers, colocation facilities, semiconductor facilities, utilities or comparable high-availability environments. Demonstrated ability to lead a significant technical area, make independent engineering decisions and influence outcomes across multiple disciplines and external partners. Deep knowledge of medium- and low-voltage AC systems and high-power DC distribution systems. Strong understanding of power electronics and conversion technologies, including rectifiers, inverters, isolated and non-isolated DC/DC conversion, controls and galvanic isolation. Strong knowledge of electrical fault behaviour and protection, including grounding, insulation monitoring, selective coordination, interruption, arc hazards, stored energy and safe maintenance. Broad technical knowledge of utility and medium-voltage systems, transformers, switchgear, uninterruptible power supplies, generators, battery energy storage, power distribution units, busway, monitoring and controls. Experience designing electrical interfaces for high-density computing or other highly dynamic mission-critical loads. Proficiency with ETAP or SKM PowerTools, together with experience using an electromagnetic-transient or power-electronics simulation environment such as PSCAD, EMTP-RV, PLECS, PSIM or MATLAB/Simulink. Strong familiarity with applicable North American and international electrical requirements, including relevant NFPA, IEEE, IEC, UL, CSA and local standards. Experience directing consultants, reviewing technical deliverables and supporting construction, testing and commissioning activities. Excellent communication and influencing skills, with the ability to explain complex engineering decisions clearly to technical and non-technical stakeholders. Sound engineering judgement and the ability to lead cross-functional work in a fast-moving environment with competing priorities. Willingness to travel approximately 10% as required. Desirable Professional Engineer licence or equivalent international professional certification. Experience with behind-the-meter generation, microgrids or other on-site energy systems. Experience developing, testing or deploying high-power DC data center architectures. Hands-on experience with emerging transformer, converter or electrical-protection technologies. Experience integrating battery energy storage, supercapacitors, fuel cells, renewable generation or demand-response capabilities. Experience designing Tier III or Tier IV data centers, or infrastructure with comparable availability requirements. Participation in relevant industry standards groups or technical forums. Experience with reliability engineering, FMEA or FMECA, functional safety, controls, telemetry or operational-technology cybersecurity. Familiarity with engineering and design tools such as Visio, Bluebeam, AutoCAD or Revit. In addition to a competitive salary, Graphcore offers flexible working and a comprehensive benefits package designed to support your health, wellbeing and financial future. Our benefits include medical, dental and vision coverage, Flexible Spending Accounts (FSAs), Health Savings Accounts (HSAs), disability and life insurance, a 401(k) retirement plan, commuter benefits, wellness services and an Employee Assistance Programme (EAP). We welcome people of different backgrounds and experiences; we're committed to building an inclusive work environment that makes Graphcore a great home for everyone. We offer an equal opportunity process and understand that there are visible and invisible differences in all of us. We can provide a flexible approach to interview and encourage you to chat to us if you require any reasonable adjustments.
Graphcore
VerifiedAbout Graphcore At Graphcore, we’re building the future of AI compute.We’re a team of semiconductor, software and AI experts, with deep experience in creating the complete AI compute stack - from silicon and software to infrastructure at datacenter scale. As part of the SoftBank Group, backed by significant long-term investment, we are delivering key technology into the fast-growing SoftBank AI ecosystem.To meet the vast and exciting AI opportunity, Graphcore is expanding its teams around the world.We are bringing together the brightest minds to solve the toughest problems, in a place where everyone has the opportunity to make an impact on the company, our products and the future of artificial intelligence. Job Summary We are looking for an experienced Principal Engineer to join our Cloud Platform Team and help develop and deploy clouds and services. Working closely with our colleagues in Software Platform, Datacentre Operations and Product Development teams, you will deploy services on our fleet of cutting-edge AI systems. As part of our Software Platform organisation, you will be involved in the cloud integration, validation, performance benchmarking, optimisation, and development of our high-performance AI solutions. These include in-house AI systems alongside off-the-shelf high-performance servers, switches and storage solutions. This is a hand-on technical role requiring a solid background in the use of cloud infrastructure, deployment using Infrastructure-as-Code, observability, high-performance networking and storage systems. You may have been working in an IT organisation, a datacentre, a cloud provider or as a developer of orchestration or cloud services. The Software Platform team at Graphcore We build Graphcore products into large-scale AI solutions for our customers and the Cloud Platform Team is responsible for providing such systems to both internal users via private clouds and customers via our own public clouds. Often the internal systems will be using and developing pre-release hardware and software, so it’s vital you are comfortable with unproven components. Responsibilities and Duties Operate and extend existing OpenStack-based cloud services and contribute to the deployment and development of new ones. You will be responsible for significant technical initiatives and projects, mentoring a small team of more junior engineers in best engineering practices. Develop and operate end-user services on our clouds and support internal users in their use. You will turn end-user and product requirements into deployed services. Help to build automation to collect and analyse metrics and other observability data from the cloud services to support clear identification and reporting of any issues. Work with users to provide information of any product-related issues to Engineering and QA departments. Work with our Datacentre Operations Engineers to maintain and operate the fleet of AI systems at peak performance in our private clouds. Configure and test new Graphcore AI hardware and systems using Continuous Deployment and Infrastructure-as-code in internal and external datacentres. Drive corrective actions for systems that are not operating correctly, working with DC operations and Graphcore Engineering as required. Work with external vendors of off-the-shelf switches, servers and storage solutions to specify, benchmark and integrate 3rd party products into our Cloud Reference Design. Skills and Experience [ALL REQUIRED] Bachelor's degree or equivalent practical experience in a relevant subject. Solid infrastructure or IT experience with a proven track record of delivering technical output as an individual contributor. Experience managing or operating on-premises or private-cloud environments. Experience specifying, scoping, estimating and detailing work plans in an AGILE and SCRUM framework, including priorities, risks, issues, impacts and constraints. Expert-level, proven Linux scripting ability (bash and python required). Expert-level, proven Linux system administration (Ubuntu, RHEL and variants). Experience with a version control system (preferably Git) and using it to manage system configuration or automation. Experience with Continuous Integration or testing pipelines using GitLab, GitHub or similar. Hands-on experience deploying services into public or private clouds using Infrastructure-as-Code (IAC). A solid understanding of the technologies underpinning cloud services (APIs, virtualisation of CPUs, IO, systems), virtual networks, block storage, resource management and monitoring. Expert with IAC automation tools (e.g. Terraform/OpenTofu, Ansible, Packer). Experience with container deployment and management tools (e.g. docker, podman, apptainer). Experience with solutions for monitoring and observability. e.g. Grafana, Prometheus, OpenSearch/ElasticSearch, Loki, Mimir, OpenTelemetry, Fluentd ,Kafka Excellent communication and presentation skills, and experience dealing with end-users of IT or cloud services. An ability to work independently and lead others on critical infrastructure without oversight, and with a focus on end-user availability. Desirable but not required: Experience with OpenStack deployments or the technologies they rely on (e.g. Ceph, Open vSwitch, KVM, QEMU ). Experience with High Performance Computing (HPC) environments using SLURM or similar batch workload solutions. Strong skillset and experience in end-to-end deployment automation and CI of containerised services. Complete automation of pipelines for build, test, deploy, manage, alert, destroy, rebuild. Experience with managing production Kubernetes clusters and workloads. Experience with workload queue management systems (SLURM, LSF, Kueue). Experience with managed switch configuration (e.g. EOS, SONiC, DNOS). Programming experience with Python3 utilising classes and inheritance. Programming experience with Go. Benefits In addition to a competitive salary, Graphcore offers flexible working, a generous annual leave policy, private medical insurance and health cash plan, a dental plan, pension (matched up to 5%), life assurance and income protection. We have a generous parental leave policy and an employee assistance programme (which includes health, mental wellbeing, and bereavement support). We offer a range of healthy food and snacks at our central Bristol office and have our own barista bar! We welcome people of different backgrounds and experiences; we’re committed to building an inclusive work environment that makes Graphcore a great home for everyone. We offer an equal opportunity process and understand that there are visible and invisible differences in all of us. We can provide a flexible approach to interview and encourage you to chat to us if you require any reasonable adjustments. Sponsorship Applicants for this position must hold the right to work in the UK. Unfortunately at this time, we are unable to provide visa sponsorship or support for visa applications.
Graphcore
VerifiedAt Graphcore, we’re building the future of AI compute. We’re a team of semiconductor, software and AI experts, with deep experience in creating the complete AI compute stack - from silicon and software to infrastructure at datacenter scale. As part of the SoftBank Group, backed by significant long-term investment, we are delivering key technology into the fast-growing SoftBank AI ecosystem.To meet the vast and exciting AI opportunity, Graphcore is expanding its teams around the world.We are bringing together the brightest minds to solve the toughest problems, in a place where everyone has the opportunity to make an impact on the company, our products and the future of artificial intelligence. Job Summary We're looking for a Linux Engineering Lead to help shape and operate the Linux platforms that underpin our engineering environments. This is a hands-on technical leadership role. You'll lead a small team of engineers while driving automation, reliability and operational excellence across the Linux infrastructure used by Graphcore's engineering organisation. You will be responsible for leading incident response, driving operational improvements, and setting standards for how Linux systems are managed and supported across the organization. While the role includes leadership responsibilities, it will initially require a hands-on approach, including direct involvement in troubleshooting, system support, and automation efforts, while building team capability and scaling processes. Collaborating intimately with engineering groups, platform engineers, and infrastructure experts, you will guarantee systems stay stable, efficient, and consistent with changing business and product delivery requirements. The Team You’ll be joining a multi-disciplinary team with strong technical skills and a very supportive culture. We work closely together, regularly share knowledge, and your skills will make a direct impact on our business. It’s an exciting and pivotal moment for us right now, with plenty of new projects ahead. If you're looking to solve interesting problems and see your work deliver real-world results, this is the team for you! Responsibilities and Duties Lead and mentor a team of Linux engineers Own the reliability, performance and scalability of engineering Linux environments Drive adoption of Infrastructure-as-Code and GitOps practices Build automation that reduces operational overhead and improves consistency Lead major incident response and root cause analysis activities Partner with software, platform and infrastructure teams to support evolving engineering requirements Establish standards, tooling and processes that enable systems to scale efficiently Improve observability, monitoring and operational visibility across the estate Help define the future direction of Linux platform engineering at Graphcore Candidate Profile Essential Significant experience administering Linux environments at scale Strong troubleshooting skills across systems, networking, storage and applications Experience leading engineers, projects or operational initiatives Strong automation and scripting skills (Python, Bash or similar) Experience with Infrastructure-as-Code and configuration management tools such as Ansible, Terraform or Puppet Experience working with Git-based workflows and CI/CD pipelines Experience managing production incidents and driving operational improvements Excellent communication and stakeholder management skills Desirable Experience supporting AI, HPC or large-scale engineering environments Experience with observability platforms and monitoring systems Experience working alongside platform engineering, SRE or DevOps teams Knowledge of identity and access management Experience building or scaling operational processes We welcome people of different backgrounds and experiences; we’re committed to building an inclusive work environment that makes Graphcore a great home for everyone. We offer an equal opportunity process and understand that there are visible and invisible differences in all of us. We can provide a flexible approach to interview and encourage you to chat to us if you require any reasonable adjustments.
Graphcore
VerifiedAt Graphcore, we’re building the future of AI compute. We’re a team of semiconductor, software and AI experts, with deep experience in creating the complete AI compute stack - from silicon and software to infrastructure at datacenter scale. As part of the SoftBank Group, backed by significant long-term investment, we are delivering key technology into the fast-growing SoftBank AI ecosystem.To meet the vast and exciting AI opportunity, Graphcore is expanding its teams around the world.We are bringing together the brightest minds to solve the toughest problems, in a place where everyone has the opportunity to make an impact on the company, our products and the future of artificial intelligence. Job Summary Our Engineering Labs are where new silicon, systems, and platforms are brought to life, tested, and scaled. We’re looking for an experienced Engineering Lab Infrastructure Lead to help us build and operate the technical foundations that support this work. You will lead the Engineering Lab Support function, ensuring our labs, infrastructure, and services remain reliable, scalable, and effective for the engineers developing Graphcore’s next generation technologies. Initially, this is a highly hands-on role. You’ll work directly with engineering teams, supporting lab environments, managing infrastructure, automating workflows, and solving complex technical problems. As the function grows, you’ll help build and lead a small team while defining the processes, standards, and service model that will support the organisation long term. You’ll work closely with silicon, hardware, systems, and software engineering teams, making a direct impact on the speed and effectiveness of product development. The Team You’ll be joining a multidisciplinary team with strong technical skills and a very encouraging culture. We work closely together and regularly share knowledge, and your skills will make a direct impact on our business. It’s an exciting and pivotal moment for us right now, with plenty of new projects ahead. If you're looking to solve interesting problems and see your work deliver real-world results, this is the team for you. Responsibilities and Duties Leading the development of Engineering Lab infrastructure and support services Acting as the technical escalation point for complex infrastructure and lab issues Managing Linux-based servers and engineering environments Supporting hardware bring-up, validation, and testing activities Designing and improving operational processes, tooling, and automation Maintaining and improving network, storage, and compute infrastructure within engineering labs Managing infrastructure through configuration management and Infrastructure-as-Code practices Building strong relationships with engineering teams and understanding their evolving requirements Developing knowledge bases, documentation, and operational standards Recruiting, mentoring, and leading a small team of Lab Infrastructure Engineers as the function grows Essential Strong Linux systems administration experience across Debian and/or Red Hat environments Experience supporting engineering, research, laboratory, HPC, or data-centre environments Solid networking knowledge including routing, VLANs, VPNs, and troubleshooting complex connectivity issues Experience managing physical infrastructure including servers, rack-mounted equipment, BMCs, firmware, and out-of-band management Experience with configuration management and automation tools such as Ansible, Puppet, or similar Familiarity with authentication and access-management systems such as LDAP, RADIUS, or Active Directory integrations Strong troubleshooting skills with a structured and methodical approach to problem solving Excellent communication skills and a customer-focused mindset Desirable Container technologies such as Docker, containerd, or Kubernetes Monitoring and observability platforms such as Prometheus, Grafana, Zabbix, OpenTelemetry, or similar Python scripting and automation CI/CD tooling including GitLab or GitHub Actions Experience supporting hardware development, silicon validation, embedded systems, or electronics laboratories Performance analysis and troubleshooting across compute, storage, and network infrastructure Web infrastructure technologies including NGINX, HAProxy, or load balancing platforms We welcome people of different backgrounds and experiences; we’re committed to building an inclusive work environment that makes Graphcore a great home for everyone. We offer an equal opportunity process and understand that there are visible and invisible differences in all of us. We can provide a flexible approach to interview and encourage you to chat to us if you require any reasonable adjustments.
Graphcore
VerifiedAbout us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore fosters continuous learning and innovation. Job Summary Reporting into the Systems Engineering organisation, the Distinguished Engineer, End-to-End Security Architect will define and lead the security architecture for Graphcore’s inference service platform. This role is responsible for establishing a comprehensive security strategy spanning platform, infrastructure, networking, service operations, customer assurance, and compliance readiness. Working across multiple engineering and operational functions, the successful candidate will provide technical leadership, drive security requirements, and ensure the platform delivers robust protection, resilience, and trust for customers. The Team You will work closely with teams across security architecture, infrastructure engineering, networking, site reliability engineering, platform software, firmware, data centre operations, compliance, legal, customer engineering, and customer security. The team collaborates across the business to deliver secure, reliable, and scalable AI infrastructure and services while supporting customer assurance, regulatory requirements, and operational excellence. Responsibilities and Duties Own the end-to-end security architecture for the inference service platform, covering infrastructure, networking, APIs, operational controls, monitoring, and customer assurance. Define security principles, threat models, trust boundaries, tenant isolation requirements, and architectural standards for the service. Establish security requirements for deployment environments, including physical security, operational controls, access management, and asset protection. Define platform security requirements across hardware, firmware, secure boot, attestation, software integrity, and lifecycle management. Lead the design of network and service isolation controls, secure communications, segmentation strategies, and administrative access protections. Own security architecture for authentication, authorisation, secrets management, encryption, key management, and data protection controls. Define privileged access management approaches, audit requirements, access review processes, and emergency access procedures. Establish logging, monitoring, telemetry, incident response, and security assurance requirements across the service lifecycle. Partner with engineering and operations teams to ensure security requirements are effectively implemented, maintained, and validated. Assess security posture against customer, contractual, regulatory, and internal requirements, managing risk-based decisions where required. Support customer security reviews, audits, penetration testing activities, security questionnaires, and technical assurance discussions. Provide technical leadership, architectural guidance, mentoring, and design review expertise across multiple teams and disciplines. Candidate Profile Essential Advanced degree in Computer Science, Computer Engineering, Cybersecurity, Electrical Engineering, or a related technical discipline, or equivalent practical experience. Significant experience in security architecture, platform security, cloud security, infrastructure security, or large-scale service security. Demonstrated experience defining and owning security architecture for customer-facing platforms, infrastructure services, or large-scale production environments. Deep understanding of threat modelling, zero-trust principles, tenant isolation, privileged access management, cryptographic controls, and secure operations. Strong knowledge of platform security technologies including trusted execution mechanisms, secure boot, attestation, firmware integrity, and supply-chain security concepts. Experience defining security requirements for data centre deployments, operational environments, and physical security controls. Expertise in network security architecture, segmentation, management-plane protection, secure communications, monitoring, and access control. Experience with key management, secrets management, certificate lifecycle management, and encryption technologies. Experience securing APIs, deployment pipelines, service control planes, and operational tooling. Strong understanding of logging, monitoring, incident response, security operations, and evidence management practices. Experience supporting customer security reviews, audits, penetration testing activities, and executive-level security discussions. Excellent communication and stakeholder management skills, with the ability to influence technical and non-technical audiences. Proven ability to lead through influence across engineering, security, operations, compliance, and customer-facing teams. Desirable Experience securing AI/ML inference platforms, model-serving infrastructure, accelerator-based systems, or confidential AI workloads. Experience with trusted execution environments, platform attestation technologies, secure firmware development, and workload identity solutions. Experience with confidential computing, secure enclaves, measured boot technologies, and remote attestation. Experience supporting highly regulated environments, critical infrastructure, or security-sensitive customer deployments. Familiarity with industry security and compliance frameworks such as SOC 2, ISO 27001, ISO 27017, ISO 27018, PCI DSS, FedRAMP, or NIST standards. Knowledge of hardware platform security, high-performance networking, storage security, and accelerator ecosystem technologies. Experience with supply-chain risk management, secure manufacturing practices, asset lifecycle controls, and secure disposal processes. Understanding of emerging cryptographic technologies and future security trends. Experience working with hyperscalers, enterprise customers, AI organisations, or managed infrastructure providers. In addition to a comprehensive benefits package, Graphcore offers flexible working arrangements designed to support your health, wellbeing and financial future. Our benefits include medical, dental and vision coverage, Flexible Spending Accounts (FSAs), Health Savings Accounts (HSAs), disability and life insurance, a 401(k) retirement plan, commuter benefits, wellness services and an Employee Assistance Programme (EAP). We welcome people of different backgrounds and experiences and are committed to building an inclusive work environment where everyone can thrive. We offer an equal opportunity recruitment process and can provide a flexible approach to interviews. Please let us know if you require any reasonable adjustments.
Graphcore
VerifiedAbout Graphcore Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation. Job Summary We are looking for a versatile Desktop & Engineering Support Engineer to be the main IT contact in our Taipei office. This role involves hands-on support for Windows, macOS, and Linux systems. It offers an outstanding mix of daily desktop support and close work with engineering teams to resolve issues and improve workflows. The ideal candidate will have extensive technical knowledge, strong problem-solving abilities, and the independence to handle both general IT tasks and engineering-specific needs. The Team You’ll be joining a multidisciplinary team with outstanding technical skills and a very encouraging culture. We work closely together and regularly share knowledge, ensuring that your contributions will have a direct impact on our business. It’s an exciting and pivotal moment for us right now, with plenty of new projects ahead. If you're looking to solve interesting problems and see your work deliver real-world results, this is the team for you. Responsibilities and Duties Serve as the main IT liaison for the region, delivering hands-on assistance across Windows, macOS, and Linux desktops and laptops Deliver day-to-day desktop support, including hardware setup, software installation, patching, and troubleshooting Diagnose and resolve complex technical issues across operating systems, networks, and applications Work together with engineering groups to troubleshoot, improve, and document development workflows and toolchains Support specialized engineering software, compilers, IDEs, and version control systems, ensuring reliable operation Manage user accounts, access, and permissions across corporate systems Proactively monitor system health, apply updates, and ensure endpoint security compliance Advance issues to global IT or vendor support where necessary, ensuring timely resolution Train and support end-users on IT guidelines and new tools Contribute to continuous improvement by identifying recurring issues and recommending long-term solutions Candidate Profile Essential: Technical Skills Proficiency in operating systems: Windows (11), macOS, Linux Desktop/laptop hardware troubleshooting, peripheral setup, printers, monitors, docking stations Networking fundamentals: TCP/IP, DHCP, DNS, Wi-Fi troubleshooting, VPN support Application support: Microsoft 365 suite, collaboration tools (Zoom, Slack) Active Directory & user management: Password resets, group policy basics, and user account administration Security awareness: Antivirus, encryption, and multi-factor authentication support Familiarity with cloud services (AWS, Oracle) Experience partnering with engineering teams to identify issues and refine workflows Proven ability to implement and manage change requests Datacentre/comms room management Essential travel between office and 3rd party factory locations to support on-site infrastructure Experience as smart hands for remote functional teams Soft Skills Strong communication skills to explain technical issues to both technical and non-technical users Ability to prioritize and manage multiple tickets under pressure Problem-solving and analytical thinking for root cause analysis Collaboration with cross-functional teams Desirable Microsoft certified (cloud/server/client) Experience with imaging & deployment tools (SCCM or others) Understanding of remote deployment tools (Ansible, Puppet) Network infrastructure configuration/management (Switches/Firewalls/Routers) Microsoft AD administration Experience with network storage technologies Experience with VCS Proficiency in documentation and issue-tracking software suites We welcome people of different backgrounds and experiences; we’re committed to building an inclusive work environment that makes Graphcore a great home for everyone. We offer an equal opportunity process and understand that there are visible and invisible differences in all of us. We can provide a flexible approach to interview and encourage you to chat to us if you require any reasonable adjustments.
Graphcore
VerifiedAbout Graphcore How often do you get the chance to build a technology that transforms the future of humanity? At Graphcore, we’re building the future of AI compute.We’re a team of semiconductor, software and AI experts, with deep experience in creating the complete AI compute stack - from silicon and software to infrastructure at datacenter scale. As part of the SoftBank Group, backed by significant long-term investment, we are delivering key technology into the fast-growing SoftBank AI ecosystem.To meet the vast and exciting AI opportunity, Graphcore is expanding its teams around the world.We are bringing together the brightest minds to solve the toughest problems, in a place where everyone has the opportunity to make an impact on the company, our products and the future of artificial intelligence. Job Summary We are looking for a Senior Network Engineer to join our Cloud Platform Team and help develop and deploy clouds and services. Working closely with our colleagues in Software Platform, Datacentre Operations and Product Development teams, you will deploy services onour fleet of cutting-edge AI systems. As part of our Software Platform organisation, you will be involved in the cloud integration, validation, performance benchmarking, optimisation, and development of our high-performance AI solutions. These include in-house AI systems alongside off-the-shelf high-performance servers, switches and storage solutions. This is a hand-on technical role requiring a solid background in the use of cloud infrastructure, deployment using Infrastructure-as-Code, observability, high-performance networkingand storage systems. You may have been working in an IT organisation, a datacentre, a cloud provider or as a developer of orchestration or cloud services. The Software Platform team at Graphcore We build Graphcore products into large-scale AI solutions for our customers and the Cloud Platform Team is responsible for providing such systems to both internal users via private clouds and customers via our own public clouds. Often the internal systems will be using and developing pre-release hardware and software, so it’s vital you are comfortable with unproven components. Responsibilities and Duties Develop and operate high-performance ethernet infrastructure on our private clouds and support internal users in their use. You will turn end-user and product requirements into deployed services. Help to build automation to collect and analyse metrics and other data from the network infrastructure to support clear identification and reporting of any issues. Work with users to provide information of any product-related issues to Engineering and QA departments. Work with our Datacentre Operations Engineers to maintain, tune and operate the fleet of AI systems at peak performance in our private clouds. Work with external vendors of off-the-shelf switches, servers and storage solutions to integrate 3rd party products into our Cloud Reference Design, with a focus on network performance, automation and resilience. Skills and Experience (all required) Bachelor's degree or equivalent practical experience in a relevant subject. Significant hands-on experience with 1 or more vendor’s high-end (100Gb/s+) ethernet switch solutions. Experience managing on-premises or private-cloud environments. Solid software engineering or IT experience with a proven track record of delivering technical output as an individual contributor. Experience working in an AGILE and SCRUM framework, including understanding of priorities, risks, issues, impacts and constraints. Strong proven Linux scripting ability (bash, python, awk, sed). Strong proven Linux system administration (Ubuntu, RHEL and variants). Experience with a version control system (preferably Git) and using it to manage system configuration or automation. Experience with Continuous Integration or testing pipelines using GitLab, GitHub or similar. A solid hands-on understanding of the technologies underpinning cloud services (APIs, virtualisation of CPUs, IO, systems) and how they relate to high-performance networking. Experience with IAC automation tools (Terraform/OpenTofu, Ansible). Experience with container deployment and management tools (e.g. docker). Experience with solutions for monitoring and observability. e.g. Grafana, Prometheus, OpenSearch/ElasticSearch, Loki. Good communication and presentation skills, and experience dealing with end-users of IT services. An ability to work independently on critical infrastructure with minimal oversight, and with a focus on end-user availability. Desirable but not required: Experience with Openstack cloud platforms. Experience with High Performance Computing (HPC) environments using SLURM or similar batch workload solutions. Experience with hardware offloading on RDMA-capable NICs and how that integrates with virtual networking on Open V-switch, KVM/QEMU. Experience with managing production Kubernetes clusters and workloads with an automation tool such as ArgoCD. Benefits In addition to a competitive salary, Graphcore offers flexible working, a generous annual leave policy, private medical insurance and health cash plan, a dental plan, pension (matched up to 5%), life assurance and income protection. We have a generous parental leave policy and an employee assistance programme (which includes health, mental wellbeing, and bereavement support). We offer a range of healthy food and snacks at our central Bristol office and have our own barista bar! We welcome people of different backgrounds and experiences; we’re committed to building an inclusive work environment that makes Graphcore a great home for everyone. We offer an equal opportunity process and understand that there are visible and invisible differences in all of us. We can provide a flexible approach to interview and encourage you to chat to us if you require any reasonable adjustments. Sponsorship Applicants for this position must hold the right to work in the UK. Unfortunately at this time, we are unable to provide visa sponsorship or support for visa applications.
Graphcore
VerifiedAbout Graphcore At Graphcore, we’re building the future of AI compute.We’re a team of semiconductor, software and AI experts, with deep experience in creating the complete AI compute stack - from silicon and software to infrastructure at datacenter scale. As part of the SoftBank Group, backed by significant long-term investment, we are delivering key technology into the fast-growing SoftBank AI ecosystem.To meet the vast and exciting AI opportunity, Graphcore is expanding its teams around the world.We are bringing together the brightest minds to solve the toughest problems, in a place where everyone has the opportunity to make an impact on the company, our products and the future of artificial intelligence. Job Summary We are looking for a Senior Engineer to join our Cloud Platform Team and help develop and deploy clouds and services. Working closely with our colleagues in Software Platform, Datacentre Operations and Product Development teams, you will deploy services on our fleet of cutting-edge AI systems. As part of our Software Platform organisation, you will be involved in the cloud integration, validation, performance benchmarking, optimisation, and development of our high-performance AI solutions. These include in-house AI systems alongside off-the-shelf high-performance servers, switches and storage solutions. This is a hand-on technical role requiring a solid background in the use of cloud infrastructure, deployment using Infrastructure-as-Code, observability, high-performance networking and storage systems. You may have been working in an IT organisation, a datacentre, a cloud provider or as a developer of orchestration or cloud services. The Software Platform team at Graphcore We build Graphcore products into large-scale AI solutions for our customers and the Cloud Platform Team is responsible for providing such systems to both internal users via private clouds and customers via our own public clouds. Often the internal systems will be using and developing pre-release hardware and software, so it’s vital you are comfortable with unproven components. Responsibilities and Duties Develop and operate Kubernetes-managed end-user services on our private clouds and support internal users in their use. You will turn end-user and product requirements into deployed services. Work with our Datacentre Operations Engineers to maintain and operate the fleet of AI systems at peak performance in our private clouds. Configure and test new Graphcore AI hardware and systems using Continuous Deployment and Infrastructure-as-code in internal and external datacentres. Skills and Experience (all required) Bachelor's degree or equivalent practical experience in a relevant subject. Experience with managing production Kubernetes clusters and workloads with a continuous delivery tool such as ArgoCD. Solid software engineering or IT experience with a proven track record of delivering technical output as an individual contributor. Experience working in an AGILE and SCRUM framework, including understanding of priorities, risks, issues, impacts and constraints. Strong proven Linux scripting ability (bash, python, awk, sed). Strong proven Linux system administration (Ubuntu, RHEL and variants). Experience with a version control system (preferably Git) and using it to manage system configuration or automation. Experience with Continuous Integration or testing pipelines using GitLab, GitHub or similar. A solid hands-on understanding of the technologies underpinning cloud services (APIs, virtualisation of CPUs, IO, systems), virtual networks, block storage, resource management and monitoring. Experience with IAC automation tools (Terraform/OpenTofu, Ansible, Packer). Good communication and presentation skills, and experience dealing with end-users of IT services. An ability to work independently on critical infrastructure with minimal oversight, and with a focus on end-user availability. Desirable but not required: Experience with Openstack cloud platforms. Experience with solutions for monitoring and observability. e.g. Grafana, Prometheus, OpenSearch/ElasticSearch, Loki. Experience with High Performance Computing (HPC) environments using SLURM or similar batch workload solutions. Programming experience with Python3 utilising classes and inheritance. Benefits In addition to a competitive salary, Graphcore offers flexible working, a generous annual leave policy, private medical insurance and health cash plan, a dental plan, pension (matched up to 5%), life assurance and income protection. We have a generous parental leave policy and an employee assistance programme (which includes health, mental wellbeing, and bereavement support). We offer a range of healthy food and snacks at our central Bristol office and have our own barista bar! We welcome people of different backgrounds and experiences; we’re committed to building an inclusive work environment that makes Graphcore a great home for everyone. We offer an equal opportunity process and understand that there are visible and invisible differences in all of us. We can provide a flexible approach to interview and encourage you to chat to us if you require any reasonable adjustments. Sponsorship Applicants for this position must hold the right to work in the UK. Unfortunately at this time, we are unable to provide visa sponsorship or support for visa applications.
Graphcore
VerifiedJob Summary We are seeking a highly experienced Senior / Lead Linux Engineering Support Engineer to lead and develop a small team supporting engineering systems within a fast-paced AI-focused environment. This role combines deep Linux expertise with strong leadership, automation, and DevOps practices to ensure systems are reliable, scalable, and supportable at scale. A key aspect of the role is establishing and operating within a configuration-as-code environment, where system configuration and operational processes are managed through automation, pipelines, and source control rather than manual administration. You will be responsible for leading incident response, driving operational improvements, and setting standards for how Linux systems are managed and supported across the organisation. While the role includes leadership responsibilities, it will initially require a hands-on approach, including direct involvement in troubleshooting, system support, and automation efforts, while building team capability and scaling processes. Working closely with engineering teams, platform engineers, and infrastructure specialists, you will ensure systems remain stable, performant, and aligned with evolving business and product delivery needs. The Team You’ll be joining a multi-disciplinary team with strong technical skills and a very supportive culture. We work closely together, regularly share knowledge, and your skills will make a direct impact on our business. It’s an exciting and pivotal moment for us right now, with plenty of new projects ahead. If you're looking to solve interesting problems and see your work deliver real-world results, this is the team for you. Responsibilities and Duties Lead, mentor, and develop a team of Linux Engineering Support Engineers, establishing clear roles, responsibilities, and ways of working Own and oversee support for Linux-based systems and engineering environments, ensuring stability, performance, and availability Act as an escalation point for complex technical issues and outages, providing hands-on support where required Diagnose and resolve high-impact system and interoperability issues across mixed and distributed environments Perform hands-on investigation and troubleshooting to understand issues and drive effective solutions Lead incident response activities, including triage, coordination, and resolution Own and drive Root Cause Analysis (RCA) processes, ensuring preventative improvements are identified and implemented Establish and improve incident management processes, driving operational maturity and reliability Drive adoption of automation and configuration-as-code practices across Linux systems Ensure system changes are delivered through controlled, auditable processes wherever possible Oversee development and implementation of automation solutions for system management and operational tasks Promote and enforce use of Git-driven workflows and CI/CD pipelines for configuration and operational processes Identify and prioritise opportunities to reduce manual effort through automation and improved tooling Work closely with engineering teams to support development environments and system requirements Act as a senior technical liaison between engineering teams and infrastructure/platform functions Support onboarding of new systems, services, and environments using standardised and automated approaches Ensure system configurations remain consistent and aligned with defined standards and governance Oversee integration points (e.g. identity, CI/CD, tooling) and ensure issues are resolved effectively Identify and drive improvements in system performance, scalability, and maintainability Contribute to and enforce documentation, standards, and operational best practices Ensure systems meet audit, compliance, and governance requirements, with full traceability of changes Essential Extensive experience administering and supporting Linux-based systems in complex technical or engineering environments Strong troubleshooting skills across operating systems, networking, storage, and application layers Proven experience diagnosing and resolving complex technical issues, including across mixed or distributed environments Proven experience handling major incidents and outages, including leading resolution and contributing to Root Cause Analysis (RCA) Strong experience with automation and scripting (e.g. Bash, Python, or similar) Strong experience with configuration management or infrastructure-as-code tools (e.g. Ansible, Terraform, Puppet, or similar) Experience working with configuration-as-code practices and Git-driven workflows Experience designing, implementing, or supporting CI/CD pipelines for configuration and operational processes Strong understanding of system interoperability across distributed environments Experience working within defined standards, governance frameworks, and controlled processes Strong communication skills and ability to work closely with engineering, platform, and infrastructure teams Experience mentoring or supporting the development of other engineers Ability to operate effectively across time zones in a distributed organisation Proven ability to operate independently, set direction, and deliver outcomes Desirable Experience leading or coordinating incident response activities Experience working alongside DevOps, platform, or infrastructure engineering teams Experience with monitoring, observability, and logging systems Experience supporting AI/ML or high-performance computing environments Understanding of identity and access management concepts Experience building or scaling operational processes or support functions Experience administering and supporting Linux-based systems in a technical or engineering environment
Showing 1–8 of 19 opportunities
Get the latest internships, jobs, scholarships and research opportunities delivered to your inbox.
Get early access to verified notifications & deadlines.
Only relevant and verified career opportunities.
We value your inbox. Zero marketing clutter.
Students & scholars actively advancing.