Author Archives: Edward Davies

About Edward Davies

My name is Ed and I am a DevOps Engineer based in the midlands of England. I have a strong focus on AWS, Infrastructure as code and serverless technologies. I enjoy keeping fit by running plus watching my football team Wolves on the weekends!

From 4 Months to 2 Years: Scaling Systems and Stepping into Seniority at Bazaarvoice

A follow-on from the smash hit that was “My First 4 Months at Bazaarvoice.”

While my first post focused on onboarding and adapting to a global organization, this post shifts focus to developing at scale, navigating major organizational shifts, and my journey to becoming a Senior Engineer.

What’s Been Going On? (The High-Level View)

To say the past two years have been busy would be an understatement. We’ve been operating at a massive scale, balancing long-term infrastructure modernization with the daily realities of keeping the lights on across long lived systems.

Here is a snapshot of what the Platform Engineering team and I have tackled:

  • Infrastructure Modernization: Migrated over 100 Kubernetes deployments from a long standing, self-managed KOPS cluster over to multiple AutoMode EKS clusters.
  • Cloud & Cost Optimization: Improved our overall AWS posture by enforcing consistent tagging standards, reducing waste, and redesigning infrastructure for our systems.
  • CI/CD Overhauls: Helped multiple development teams migrate from Jenkins setups to modern GitHub Actions pipelines.
  • Multi-Cloud Networking: Connected our AWS and GCP footprints via a private cross-cloud interconnect.
  • Developer Empowerment: Developed an Internal Developer Platform (IDP) to consolidate and accelerate our application development lifecycles.

Navigating the Team Merger: Cultivating a Shared Brain

Halfway through my tenure, our Platform Engineering organization underwent a major shift. Originally, we operated as two small, distinct teams: Cloud Engineering (my team, focused on our AWS footprint, cloud governance, and least-privilege access across 50+ accounts) and Developer Experience (DevEx).

Unifying into a single, cohesive team was a huge win for our operational health. Not only did it give us a far more sustainable on-call rotation, but it also allowed us to share deep system knowledge and enhance system resilience. 

When we merged into a single Platform team, we inherited 60+ new applications and services, including the long standing KOPS Kubernetes cluster and scattered CI/CD tooling.

Bridging this massive knowledge gap wasn’t going to happen overnight. To tackle it, we designed a deliberate knowledge-transfer strategy:

  • Targeted KT Sessions: Deep dives into large, complex systems led by the original maintainers.
  • Split On-Call Rotations: Pairing a DevEx member with a Cloud Engineer (Primary/Secondary) so engineers could safely shadow unfamiliar issues.
  • Leveraging Subject Matter Experts (SMEs): Recognizing that no single engineer can learn every bespoke system in a few months, we leaned on individual strengths. Everyone naturally upskills over time by collaborating with the resident experts.

Game Days: Turning Chaos into a Sandbox

While product development teams have hackathons to build new features, our platform team has benefited immensely from Game Days. These are safe, controlled sandbox environments where the team can work through real-world failure scenarios without risking production.

They have proven incredibly useful for two things: handing over newly developed systems, and reverse-engineering complex, long-standing architectures. 

Case Study: Cross-Cloud Interconnect

My teammate and I were responsible for building our Cross-Cloud Interconnect, a private connection linking our GCP and AWS VPCs. While we knew the architecture inside and out, the rest of the team were focused on other projects and lacked context.

To fix this, I led a series of Game Days. In a sandbox environment, I intentionally broke the infrastructure so other engineers could get hands-on experience troubleshooting errors like BGP session failures, overloaded capacity, and connection scaling. We also practiced modifying the Infrastructure as Code (IaC) that defines the system.

This served as the perfect final handover. It allowed us to stress-test our runbooks & documentation, fill in any missing gaps, and ensure the entire team felt confident supporting the system.

The outcome of this was that we had 100% up time across the Black Friday/Cyber Monday weekend where Cross Cloud Interconnect handled 2 Billion+ requests.

We used this same chaos engineering model to build confidence in our long standing systems rebuilding them from scratch in a sandbox to map out hidden dependencies and understand how they behaved under load.

Embedding with Development Teams

In Platform Engineering, it’s easy to become siloed in your own ecosystem and lose sight of what product teams are actually building. Over the past year, we made a conscious effort to break down those walls and embed directly with our internal “customers.”

Instead of just handing over infrastructure from afar, we partnered closely with development teams on two major initiatives:

  • Simplifying Architecture (Lambda to ECS Fargate): We collaborated with a development team to migrate a complex, multi-Lambda and Step Function setup into a streamlined AWS ECS Fargate deployment. Rather than treating this as a one-off fix, we used it as a blueprint for the future. We built reusable Terraform modules and shared GitHub Actions workflows, creating a repeatable “paved path” that any team at Bazaarvoice can now leverage to adopt Fargate instantly.
  • Standardizing the Core (The KOPS to EKS Migration): Migrating applications over to our new EKS clusters wasn’t just a technical shift; it was a chance to see the daily challenges of dev teams working with heterogeneous applications and infrastructure. Working side-by-side with various teams gave us insights into their unique tech stacks. We used these insights to drive organizational consistency, aligning disparate development workflows as we moved them into the new EKS ecosystem.

By shifting from an “isolated platform” mindset to a collaborative, high-touch model, we didn’t just ship improved  infrastructure, we built stronger relationships across the entire engineering organization.

Developing in the Era of AI

In the two years I’ve been here, AI tooling has shifted from a novelty to an absolute necessity. Tools like GitHub Copilot have evolved from simple unit-test generators into foundational workflow drivers. If you aren’t actively leveraging AI in your development workflow today, you are unintentionally slowing yourself down. 

This was majorly highlighted at Platform Con 2026 where a majority of talks and workshops were AI focused. Meaning it isn’t just a shift at our org, but across Platform Engineering as a whole. 

Recognizing this shift, I gave an internal talk sharing tips and tricks for maximizing Copilot’s utility covering prompt engineering, model selection, and the critical importance of providing rich repository context.

In the world of LLMs, things move fast. With the emergence of MCP (Model Context Protocol) servers, agentic workflows and repository-level configuration files, the landscape is shifting again, and keeping our engineering practices aligned with these advancements is a continuous focus.

Agentic Case Study: Automating IAM Audits

Historically, managing IAM permission changes meant tedious, manual reviews from the Platform team. By offloading the first pass to an AI agent, we now automatically scan every IAM PR for least-privilege violations and compliance policy drift. This cuts down manual toil for our team while giving developers instant feedback on security concerns.

The Path to Senior Engineer

Looking back at my first four months, my focus was inward: learning the communication loops (like moving “work out of DMs” into public Slack channels), understanding team work-ownership, and treating internal developers as stakeholders.

Moving into a Senior Engineer role over these past two years required shifting that focus outward. Seniority at Bazaarvoice hasn’t just been about handling larger technical footprints like EKS migrations or cloud networking; it’s been about delivering impact at an organizational level rather than a team level.

Whether it was leading Game Days to up skill my peers, building reusable project scaffolds that abstract away complexity for hundreds of developers, the goal has been the same: make the right way the easy way for every engineer at Bazaarvoice.

It’s been a fast paced, incredibly rewarding two years and I’m excited for what’s next. Want to get in on the action? Head over to our careers page

My First 4 months at Bazaarvoice as a DevOps Engineer

I joined Bazaarvoice as a DevOps engineer into the Cloud engineering team in September 2023. It has been a very busy first 4 months learning a lot in terms of technical and soft skills. In this post I have highlighted my key learnings from my start at BV.

Communication

One of the key takeaways I have taken is no work in the DMs. I, and I imagine many others are used to asking questions via direct message to whom we believe would be most knowledgeable on the subject. Which can often lead to a goose chase of getting different names from various people until you find who can help. One thing my team Cloud Engineering utilizes is asking all work-related questions in public channels in slack. Firstly, this removes any wild goose chases as anyone who is knowledgeable on the matter can chime in not just a single person. Furthermore, by having these questions public it creates a FAQ page within slack. I often find myself now debugging or finding answers by searching questions/key words straight into the slack search bar and finding threads of questions which are addressing the same issues I’m facing. This means there are no repetitive answers and I do not have to wait on response times.

Bazaarvoice is a global organisation where I have team members across time zones. This essentially means you have only 50% of the day where myself and my colleagues in the US are online. So, using that time asking questions which have already been answered is not a productive use of time.

Work Ownership

Another concept which I have changed my views on is work ownership and pushing tickets forward as a team.

If you compare a Jira ticket from the first piece of work I started to my current Jira tickets 4 months in, you’ll notice now there is a stream of update comments in my current tickets. This feeds into the concept of the team owning the work rather than just myself. By having constant update comments if I fall ill or for whatever reason can’t continue a ticket a member of the team can easily get context on the state of the ticket by reading through the comments. This allows them to push the ticket forward themselves. Posting things like error messages and current blockers in the Jira comments also allows team members to offer their insight and input instead of keeping everything private.

As well as this upon finishing work, I would usually find another ticket to pick up, however what I now understand is that completing the teams work for the sprint is what’s important, using spare time now to help push the team’s tickets over the line with code reviews, jumping in huddles to troubleshoot issues as well as picking up tickets that colleagues haven’t got round to in the sprint. I now understand the importance of completing work as a team rather than an individual.

Treating internal stakeholders as customers

Bazaarvoice has many external customers and there are many teams who cater to these customers. However, in Cloud Engineering we are not an external customer facing team. Although, we do still have customers. This concept was strange to me at first where our colleagues in other teams were regarded as “customers”. The relationship is largely the same with how one would communicate and have expectations of external customers. We have SLA’s which are agreed upon as well as a “slack channel” for cloud related queries or escalations which the on-call engineer will handle. This customer relationship allows us to deliver efficiently with transparency. Another aspect of this is how we utilise a service request and playbook model. A service request is a common task/operation for our team to complete such as create an AWS account or create a VPC pairing. The service request template will extract all the necessary information from the customer needed to fulfil the request. This removes back and forth conversation between operator and customer, gathering the required information. Each service request is paired with a playbook for the operator to use. These playbooks include step by step instructions of how the operator can fulfil this request. Allowing someone like me from as early as week 2 to be able to fulfil customer requests.

Context shifting

In previous roles I would have to drop current workloads to deal with escalations or colleague questions. This requires a context shift where you must leave the current work and switch focus to something completely unrelated. Once the escalation/query is resolved then you must switch back. This can be tiresome and getting back into the original context of which I was working on takes time. This again feeds into the customer relationship the team has with internal stakeholders. A member of the team will be “on-call” for that week where they will handle all customer requests and answer customer queries. This allows the rest of the team to stick within their current context and deep focus on the task at hand without needing to switch focus. I have found this very beneficial when working on tasks which require a lot of deep focus and as such feel a lot more focussed in my delivery of meaningful work.

Utilising the size of the organisation

The organisation’s codebase consists of 1.5k repositories containing code serving all kinds of functionality. This means when embarking on a new piece of work there is often a nice template from another team which can be used which has a similar context due to being in the same organisation (same security constraints, AWS accounts ect.). For example, I recently created a GitHub Actions release workflow for one of our systems but had problems with authenticating into AWS via the workflow. A simple GitHub search for what I was looking for allowed me to find a team who has tackled the same problem I am facing. Meaning I can see how my code differs from theirs and make changes accordingly. I have had a problem solved by another team without them even realising!

Learn by doing!

I find reading documentation about systems can only get you so far in terms of understanding. Real understanding, I believe, comes from being able to actively make changes to a system and deploy these changes as a new release. This is something me and my onboarding mentor utilised. They set me a piece of work which I could go away and try in the mornings then we’d come together in the afternoon in a slack huddle to review my progress (making use of the time zone differences). This is a model that certainly worked for me and something I would try with any new onboardees that I may mentor.

Conclusion

Overall I have thoroughly enjoyed my first 4 months as a DevOps Engineer at Bazaarvoice particularly working on new technologies and collaborating with my team mates. But what has shocked me most is how much I had to improve on my communication and ways of working, something that is not taught much in the world of engineering with most the emphasis being on technical skills.

I hope you have enjoyed reading this and that you can take something away from your own ways of working!