Base Career helps you apply smarter for this job.
Key skills for this role
Based at Trakm8's head office in Coleshill, the IT & Service Delivery Engineer is responsible for the availability, performance, security and resilience of Trakm8's telematics and optimisation platforms and of the corporate IT services.
The role spans a mixed environment of corporate office, on-premises data centres and AWS.
Customer platforms process data in real time, alongside the standard business IT services used by employees.
Reliable delivery of both is critical.
Trakm8 runs a small, multi-skilled team covering both IT and Service Delivery.
Between them the team supports over 800 servers, eight production platforms, plus UAT, QA and Dev environments and around 60 users, whilst supporting services 24x7.
Based at Trakm8's head office in Coleshill, the IT & Service Delivery Engineer is responsible for the availability, performance, security and resilience of Trakm8's telematics and optimisation platforms and of the corporate IT services. The role spans a mixed environment of corporate office, on-premises data centres and AWS. Customer platforms process data in real time, alongside the standard business IT services used by employees. Reliable delivery of both is critical.
Trakm8 runs a small, multi-skilled team covering both IT and Service Delivery. Between them the team supports over 800 servers, eight production platforms, plus UAT, QA and Dev environments and around 60 users, whilst supporting services 24x7.
Monitor the health, capacity and performance of the telematics and optimisation platforms, corporate IT services and supporting infrastructure, taking action to maintain agreed levels of availability and throughput.
Investigate and resolve incidents across Linux and Windows servers, databases, networks, storage, container platforms, end user devices and application deployments, owning each from first report or alert through to root cause and permanent fix.
Skip the repetitive application forms
Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.
Trusted by over 500,000 job seekers on Base Career
More from this employer
, GBR
, GBR
Northampton, GBR
, AUS
Abingdon, GBR
Milton Keynes, GBR
Milton Keynes, GBR
, GBR
Carry out structured root cause analysis for recurring or significant issues, implementing corrective and preventative actions.
Communicate incident progress, risks and resolution clearly to technical and non-technical stakeholders.
Design, build, configure and maintain secure, resilient server infrastructure across data centre, virtualised and AWS environments.
Manage and improve containerised services, networking, storage, databases and load balancing.
Own backup, replication and restore testing, and contribute to capacity planning and disaster recovery, keeping recovery points and recovery times fit for purpose and evidenced.
Work with development, platform and service teams to support reliable application deployment and end-to-end service performance.
Administer Active Directory, Group Policy, DNS and internal certificate services across the domain.
Own server and workstation patching and endpoint protection coverage, including approvals, safeguard holds, deployment rings.
Administer accounts, access rights and licensing, and maintain inventory, imaging and software deployment tooling.
Respond directly to user requests and access issues, taking ownership through to resolution within the team’s operating model.
Maintain internal IT services including LAN, wireless, VOIP and mobile telephony, and support IT procurement, supplier management and hardware refresh.
Develop and maintain automation and scripts for repeatable operational tasks.
Maintain effective monitoring, alerting and dashboards across metrics and log platforms, reviewing thresholds and coverage to identify issues early and reduce avoidable incidents.
Build self-healing and auto-remediation where it is safe to do so, for example automated rebalancing of workloads in response to queue lag or resource pressure.
Identify technical debt and resilience gaps and recommend proportionate improvements, making effective use of modern tooling including AI-assisted development and automation.
Operate infrastructure with a security-first approach, applying access controls, hardening and vulnerability remediation in line with company policies.
Support logging, monitoring and security tooling, and assist with investigation and remediation of security alerts.
Plan and deliver changes and project work through the RFC and release process, booking into agreed release windows and providing test, rollback and post-implementation evidence.
Produce recurring operational and service reporting covering availability, capacity, patching, backup and security, and maintain accurate asset, licence and configuration records.
Create and maintain clear technical documentation, runbooks and support procedures for new and existing solutions.
The environment currently includes the following technologies. The list provides context for the role and is not intended to mean that experience in every technology is essential. The expectation is competence in several of these groups and the ability to pick up the rest.
Oracle Linux and Windows operating systems across VMware vSphere, Proxmox and AWS
Kubernetes and Docker, using MetalLB and Kong API gateway
MySQL, Cassandra (Scylla), CockroachDB, Redis, Kafka, Elasticsearch, PostgreSQL and RabbitMQ
Ansible and AWX, with scripting using Bash, Python or Perl
TCP/IP networking, DNS, routing, VPNs, BGP, SMTP, HTTP and HTTPS
Palo Alto firewalls; HAProxy, NGINX and dynamic DNS services
Grafana, Prometheus, Graphite (ClickHouse), InfluxDB and Icinga, with alerting into OpsGenie
ELK for application logging and Wazuh for platform host security monitoring
Git, Bitbucket, Jira and Bamboo CI/CD pipelines
iSCSI SAN, storage pools, volumes and snapshots
AWS services including EC2, Route53, RDS, ELB, VPC, Multi-AZ subnets, VPN, Transit Gateway and Customer Gateway
Windows Server and Windows desktop estates, Active Directory, Group Policy, DNS, DHCP, DFS and file and print services with Linux (Ubuntu)
Microsoft 365 and Entra ID, Intune, including Exchange Online, SharePoint, Teams, multi-factor authentication and Conditional Access
VMware vSphere and vCenter, with Veeam Backup & Replication and object storage for offsite copies
Endpoint management, patching and software deployment tooling, with BitLocker disk encryption
CrowdStrike Falcon endpoint protection and UTMStack SIEM
Palo Alto firewalls managed through Panorama, with GlobalProtect remote access
LAN, wireless, VOIP and mobile telephony across multiple UK sites
Puppet for configuration management and PowerShell for scripting, with Icinga and OpsGenie alerting shared across both estates
Internal certificate services, TLS and PKI, split-horizon DNS, LDAP and Kerberos
MS SQL behind business systems, plus the Atlassian suite, service desk, intranet and digital signage platforms
Depth in Kubernetes administration, cluster operations and container networking.
AWS, VMware or other virtualisation and cloud platforms.
Databases in production, such as MySQL, PostgreSQL, MS SQL, Cassandra, Kafka or Elasticsearch.
API gateway technologies such as Kong and measuring or reporting API performance.
Telemetry, IoT, high-throughput data platforms or similarly critical real-time services.
Storage, SAN and data centre experience covering physical security, networking and compute.
Endpoint management, software deployment and patching tooling such as PDQ, WSUS or Intune.
Backup and recovery tooling such as Veeam, including restore testing.
Providing IT support to end users directly, with the confidence to deal with colleagues at any level.
Configuration management with Ansible, Puppet or AWX.
Monitoring, alerting and dashboarding with Icinga, Grafana, Prometheus or similar.
Infrastructure-as-code, CI/CD tooling or automated testing for infrastructure changes.
AI-assisted coding and automation tooling used to accelerate operational and scripting work.
Firewall policy administration, ideally Palo Alto, and load balancing or reverse proxy work.
Endpoint protection, vulnerability management or SIEM platforms.
System build following NIST standards, hardening, SELinux and server firewalls.
Information security frameworks such as ISO 27001 or Cyber Essentials.
Working within a formal RFC and change window process, including rollback planning and evidence.
Moving live services off unsupported platforms and end-of-life operating systems without an outage.
Disaster recovery, capacity planning and service continuity testing.
Calm, methodical and solutions-focused, particularly when responding to live service issues.
Takes ownership and follows actions through to resolution.
Proactive in identifying risks, improvement opportunities and emerging capacity or resilience concerns.
Security-conscious and careful when working with critical systems and customer data.
Collaborative and approachable, with a willingness to share knowledge and support colleagues.
Organised and adaptable, able to balance planned work with changing operational priorities.
Committed to clear documentation, continuous learning and improving how the team works.
We’re a close-knit team who care about each other, our customers and making a positive impact, our values drive who we are and the decisions we make:
Caring & Committed – We show up for each other and deliver with dedication
Make a Difference – We create impact that matters
Relentless – We push forward until we succeed
Do the Right Thing – We choose integrity every time
Joy of the Journey – We enjoy the ride and celebrate along the way
As a Trakm8 Group employee, you have the following H&S responsibilities:
To comply with the Group H&S policy and procedures
To take care of your own health and safety and that of people who may be affected by what you do (or what you do not do)
To co-operate with others on health and safety
To not interfere with, or misuse, anything provided for your health, safety or welfare
To wear/use the correct Protective Equipment at all times
To follow the training you have received at all times
As a Trakm8 Group employee, you have the following Information Security responsibilities:
To follow the Group Information Security policies at all times
To follow the Trakm8 Secure Development Policy at all times;
To ensure that passwords under your responsibility are kept confidential
To keep Group or Company information secure and confidential at all times
To inform your manager immediately if you detect, suspect or witness an incident that may be a breach of security
References
Pre-Employment Credit Check
Basic Disclosure
The above job description is meant to describe the general nature and level of work being performed; it is not intended to be construed as an exhaustive list of all responsibilities, duties and skills required for the position. All job descriptions are subject to possible modifications in line with the needs of the business.
Verified company details for this employer are not available yet.
Full-time
Mid
Onsite
Apply faster on company sites with our extension.