What Is GitOps? Principles, Benefits, and How It Works

With GitOps, you can roll back your cluster with git revert and explain each approved change through a commit log. Routing normal changes through Git and running an in-cluster agent that detects drift from the repository gives infrastructure changes the same review and rollback discipline as application code, and Git preserves their history.

This guide covers the four GitOps principles, the pull-based reconciliation workflow, and the tools and practices you need to get started.

What Is GitOps?

The term GitOps blends “Git” and “operations,” and it names an operational model where a Git repository holds the declarative desired state of your infrastructure and applications while an agent running inside the target environment continuously pulls that state and applies it.

Git acts as the single source of truth, which means any change missing from the repository stays outside the approved cluster state.

Teams often describe their setup with the phrase “we keep our config in Git,” which usually points to a continuous integration and continuous delivery (CI/CD) pipeline that reads from a repository without the agent that handles the reconciliation work.

Weaveworks coined the term in a 2017 post that described GitOps as using developer tooling to drive operations. Storing manifests in Git leaves a gap between the desired state in source control and the current state in the live system, and closing that gap is what the agent exists to do. After Weaveworks closed in 2024, stewardship of the definition passed to the Cloud Native Computing Foundation (CNCF) OpenGitOps working group.

The Four GitOps Principles

The OpenGitOps working group operates under the CNCF App Delivery Special Interest Group and published the principles as OpenGitOps v1.0.0. The group produced them with 34 co-authors and input from more than 60 companies, with founding members including Amazon, Azure, GitHub, Red Hat, and Weaveworks.

The four principles fall into two halves. The first half covers how you describe and store the desired state, which the Declarative and Versioned and Immutable principles address through your repository and its commit history. The second half covers how that state reaches the cluster and stays there, which the Pulled Automatically and Continuously Reconciled principles address through the in-cluster agent.

Storing configuration in a repository satisfies the first two principles, and running an agent that watches and applies that state satisfies the other two.

Declarative

To manage a system with GitOps, you must express its desired state declaratively. You describe what should exist, a Deployment with three replicas, and the tooling works out how to get there.

An imperative kubectl scale deployment checkout --replicas=5 changes the cluster but leaves no artifact a controller can compare against later; a manifest with replicas: 5 does. Kubernetes manifests, Helm charts, Kustomize overlays, Terraform files, and other declarative formats all qualify.

Versioned and Immutable

Principle two requires a state store that enforces immutability and versioning and retains a complete version history. In Git, each desired-state commit records its author and timestamp. The diff captures the change, while your pull request policy can add a reviewer and approval record.

git revert handles rollback: you revert the commit, and the agent applies the previous known-good state at its next reconciliation. The v1.0.0 glossary calls the repository a “state store” and notes that Git is the canonical example, but teams can use any other system that meets these criteria.

Pulled Automatically

Principle three requires software agents to pull the desired state declarations from the source automatically. Argo CD and Flux are the two CNCF graduated reconciliation engines that implement this pattern for Kubernetes, and both run inside the cluster and fetch from the repository, which lets your CI system avoid credentials that can apply cluster changes.

A push-based approach creates a god-mode scenario in which the CI/CD pipeline holds credentials for deployments. With a read-only outbound connection to Git, the agent lets you keep the cluster’s control plane off the public network.

Continuously Reconciled

The fourth principle requires software agents to observe actual system state continuously and attempt to apply the desired state. Reconciliation means ensuring the actual state of a system matches its desired state, and unlike trigger-driven CI/CD, any divergence triggers reconciliation in GitOps.

The agent treats two sources of divergence identically: a new commit changed the desired state, or someone changed the live state. The second case is configuration drift, and the agent reverts a hotfix you apply with kubectl edit at the next cycle unless you also commit the change to Git.

How GitOps Works: The Pull-Based Reconciliation Loop

A GitOps workflow runs on a loop where Git holds the desired state and an in-cluster agent works to make the live cluster match it. Every change follows the same path from a pull request to a merge commit to a reconciliation cycle, which is what turns a repository into the deployment interface for the cluster.

Say you open a pull request that bumps the checkout service’s image tag from v2.0 to v2.1 in a Kustomize overlay. A reviewer approves it, the pull request merges, and the repository now describes a state the cluster doesn’t yet have. The in-cluster operator notices the gap on its next poll, applies the diff, and reports the application as synced.

The same loop covers two situations that look different on the surface but resolve the same way. A new merge commit changes the desired state, and the agent brings the cluster forward to match. Someone runs kubectl edit against the live cluster, and the agent treats that as drift from the repository and reverts it at the next interval. Flux, one of the two reconciliation engines introduced earlier, applies this behavior to any kubectl edit/patch/delete change unless you suspend reconciliation or push the change to Git.

Pull-Based vs. Push-Based Deployment

Push-based deployment is the pre-GitOps default. A CI job authenticates to the cluster and runs kubectl apply or helm upgrade when the pipeline fires.

By itself, that model does not continuously detect deviations between the cluster and the repository between runs. It also requires the CI system to hold cluster credentials, which is the exposure the pull model removes from CI.

Multi-environment promotion works in either model, though you can model environments as directories rather than branches so that promotion becomes a file copy from envs/staging to envs/prod.

Both engines support this pattern, and Flux serves as the example here because its dependency and health-check primitives make gated promotion the most direct to configure. With Flux, you can merge an infrastructure change to staging first and promote it to production only after the cluster reconciles and passes conformance tests. Argo CD reaches the same outcome through sync waves, hooks, or its ApplicationSet controller, which coordinate the promotion across environments rather than gating it inline.

GitOps vs. DevOps: Key Differences

DevOps and GitOps sit at different levels of the delivery stack, which is why teams often adopt them together rather than choosing between them. DevOps describes the cultural and organizational shift that gets development and operations teams working from shared goals, while GitOps is a specific continuous delivery technique that runs inside that culture.

Adopting GitOps changes how the deployment step actually executes by making Git the control plane and handing continuous deployment to an in-cluster agent. It fits as one implementation choice within a broader DevOps practice.

GitOps and infrastructure as code (IaC) overlap on the declarative front but differ in scope. IaC defines desired state in files, and a workflow that stops at terraform apply still requires you or your automation to run another plan whenever you need to check for drift. Argo CD works as a two-way reconciliation engine that stays aware of changes at the destination and closes that gap by watching the live state on an interval, whereas one-way IaC automation only fires when the source changes.

Key Benefits of GitOps

The benefits follow from the four principles rather than from any particular tool. Your pull request creates a review path for infrastructure changes, while continuous reconciliation keeps checking whether the live state matches the approved state:

  • Shorter deployment cycles and recovery times: Your pull request becomes the deployment unit, so shipping a change follows the same review path as a code change. Continuous reconciliation can also shorten mean time to repair (MTTR) by making rollback a version-controlled action.
  • Rollbacks without a runbook: git revert restores the last known-good state and the agent applies it.
  • Stronger security: With a pull-based setup, no cluster credentials need to sit in CI, and the operator’s kubectl apply permissions limit what a compromised Git repository can change.
  • Audit and compliance: The Git log records the approved desired-state history for the environment, though it does not capture every transient live-state mutation. Your auditor can trace an approved ingress timeout change to a commit hash rather than a Slack thread.
  • Reduced configuration drift: Continuous reconciliation reverts one-off kubectl fixes, which forces every lasting fix into Git or out of the cluster. Teams that skip continuous reconciliation and automatic rollback also skip this benefit.
  • Self-documenting operations: The repository is the current answer to “what is running in production.”

These benefits turn your repository into a versioned record of every approved change and give the cluster a continuously enforced desired state. Git holds what your team agreed to run, and the agent keeps checking the live environment against that record so approvals and reality stay in step.

GitOps Tools

The tooling splits into three layers, and only the first is GitOps-specific. Argo CD and Flux provide the reconciliation layer. Both engines are CNCF graduated projects:

  • Argo CD: Argo CD watches one or more repositories, compares rendered manifests with live cluster state, marks applications OutOfSync when they differ, and syncs automatically or on approval.
  • Flux: Flux uses a set of controllers, the flux command-line interface (CLI), and Git; it reconciles GitRepository, Kustomization, and HelmRelease resources and manages its own installation from Git.
  • Configuration tools: Helm templates charts and Kustomize patches base manifests with per-environment overlays, and their rendered output is what the engine applies. A kustomize edit set image command in a CI step is the typical way a new build enters the repository.
  • Git hosting: GitHub, GitLab, or any standard Git host works, since the engines need repository access, any required credentials, and optionally a webhook endpoint for faster notifications.

These layers separate reconciliation from manifest rendering and repository hosting. The two engines differ on defaults more than on model.

The table below compares their default reconciliation intervals, drift correction behavior, interfaces, and self-management approaches. Which interface and reconciliation settings you prefer will usually drive the choice.

Argo CDFlux
Default reconcile intervaltimeout.reconciliation controls polling, with jitter affecting the timingKustomization runs every five minutes, and .spec.interval sets the interval
Drift correction defaultYou opt in with selfHeal: true, prune: trueFlux reverts drift without extra configuration
InterfaceWeb interfaceflux CLI and Git
Self-managementYou install it from a pinned manifestFlux manages its own installation via flux bootstrap

How to Get Started with GitOps

Getting a GitOps pipeline into a working state on Kubernetes takes a handful of concrete steps, and each one maps to a decision the engines expect you to have made before they can reconcile anything.

The sequence below moves from the underlying platform to a single running application, and then to the growth patterns that keep the pipeline manageable as more teams and clusters come online:

  1. Provision a Kubernetes cluster and a Git repository for manifests. You may keep the manifest repository separate from application source so that config changes flow through their own review path.
  2. Install one of the two reconciliation engines in the cluster. Argo CD installs into its own namespace from the published manifest, and you should consider pinning a version for production. The Flux bootstrap installation commits Flux’s own manifests to your repository so every later change, including Flux upgrades, goes through a Git push.
  3. Point the engine at a directory holding one application and turn on automated sync. This gives you a full end-to-end loop against a small surface area before you widen the scope.
  4. Move to a directory-per-environment layout once the single-application setup is stable. Promotion between environments then becomes a file copy from envs/staging to envs/prod under the same review process.
  5. Adopt Argo CD’s ApplicationSet controller or Flux’s multi-tenancy configuration when several teams or clusters share the pipeline. Both patterns let you template application definitions across environments and tenants without duplicating manifests.

Readiness for any of this comes back to Principle one. If you still apply anything by hand or by script, you have to convert it to declarative state before the engine has something to reconcile against.

GitOps pipelines manage what you deploy, and Coralogix covers how it behaves and watches post-deploy telemetry data for drift-related regressions across logs, metrics, and traces. That lets you manage desired state and runtime behavior through complementary workflows.

GitOps Tracks What Deployed; Telemetry Shows How It Ran

A healthy status in Argo CD means the resources exist and report readiness, which is a narrower claim than most teams read into it. Successful responses from the checkout path require separate runtime verification, and a bad change can auto-sync to production while every health check passes. Closing that gap means correlating deploy commits with telemetry data, and the OpenTelemetry (OTel) CI/CD conventions include a commit revision attribute for tagging telemetry with the commit that shipped it.

Once telemetry carries the commit that shipped it, an observability layer can turn a live regression back into a specific change in Git.

Olly, Coralogix’s autonomous observability agent, cross-references live system behavior with the code changes in Git and runs root cause analysis down to the line that caused the regression. Pairing a GitOps pipeline with Kubernetes monitoring lets you tell a bad deploy from a bad reconciliation, then line up latency spikes with nearby pipeline events and Git activity such as commits or pushes. Deployment markers add context when your pipeline sends those events to Coralogix.

Start a free 14-day Coralogix trial and point it at a GitOps-managed cluster to see the same commit hash that Argo CD or Flux synced show up alongside the logs, metrics, and traces from the workloads it deployed.

Frequently Asked Questions About GitOps

What is the difference between GitOps and Jenkins?

Jenkins is a continuous integration tool that builds, tests, and publishes artifacts; in a push-based pipeline it also runs the deploy step against the cluster. GitOps moves that step inside. Jenkins commits the new image tag to the config repository, and Argo CD or Flux syncs it.

What are the four principles of GitOps?

The four principles are Declarative, Versioned and Immutable, Pulled Automatically, and Continuously Reconciled. A pipeline that stores manifests in Git and never watches the cluster meets half the definition.

Does GitOps only work with Kubernetes?

GitOps also works outside Kubernetes. Its principles apply to any infrastructure that you can observe and describe declaratively. Non-Kubernetes targets can rely on the Tofu Controller for Terraform or Crossplane for cloud resources, both of which still run inside a Kubernetes cluster.

What is Argo CD and how does it relate to GitOps?

Argo CD is a CNCF graduated continuous delivery tool that implements pull-based GitOps for Kubernetes: it watches a repository, diffs the desired manifests against live cluster state, and applies the difference. Flux provides the main alternative with the same reconciliation model.

How does GitOps handle secrets?

You should never commit plaintext secrets to Git. Common patterns include encrypting secrets in the repository with Sealed Secrets or another encryption tool, or storing them in an external manager such as HashiCorp Vault and syncing them in with the External Secrets Operator. You should prefer destination-cluster secrets rather than injecting them during manifest generation, since generated manifests sit in plaintext in Argo CD’s Redis cache.

The Importance of Communication in Software Development Teams

Programming is often thought of purely as a problem-solving activity. This may be true for the lone coder on their homegrown software in their garage, but in the multi-person environment of an Agile team, such problem solving must be collaborative. In this article, we’ll look at the role of communication in software development, particularly in an Agile framework.

Covid-19 has forced an unprecedented shift to remote working so we’ll finish up with a discussion of how Agile can be implemented in a remote setting.

Communication and Agile Development

Agile is increasingly becoming the gold standard for software development teams. With its emphasis on close-knit, self-organizing teams that work informally and respond quickly to change, Agile reserves a critical role for communication in software development.

This is because, as stated in the Agile Manifesto, people and their interactions are the driving force behind successful software development. Teams are self-organizing and need to quickly respond to change. This means that good communication is a cornerstone of successful software development. 

Daily Standups and Scrum Meetings

According to Principles Behind the Agile Manifesto, face-to-face interaction is the most effective form of communication in software development. The main way of facilitating this is via daily standups or Scrum meetings. These involve the entire team meeting at the beginning of each day in a sprint. 

Engineers take turns to explain what they have been doing and any problems they have encountered.

Whiteboarding

Whiteboards play a critical role in facilitating collaborative communication in software development. That’s because whiteboards are a low-tech, flexible, and intuitive way to facilitate communication in software development teams.

They allow developers to write down ideas and sketch diagrams as soon as inspiration strikes. The whiteboard’s physical prominence makes it a public platform for team communication; anything written on there can be seen by every team member and inspected at leisure.

Since Covid-19, many teams have switched to remote working. Software such as AWW board and Stormboard provide virtual whiteboarding solutions that allow the benefits of whiteboarding to be used anywhere in the world.

Pair Programming

Pair programming is the ultimate form of teamwork and communication in software development. It typically involves two programmers sitting side by side at one computer. One programmer, called the driver, writes the code. The other programmer, the navigator, watches what the driver is doing, spotting errors and suggesting improvements.

At first blush, pair programming seems uneconomical. Why would you get two (expensive) programmers to do the work of one? While this is the knee-jerk reaction of many managers, Alistair Cockburn and Laurie Williams tell a different story in their research, “The Costs and Benefits of Pair Programming”.  They cite several examples, including a study on a class of software engineering students to show that pair programming is economical, results in fewer coding errors, and higher team satisfaction.

Pair Programming Remotely

Surprisingly, pair programming can be done remotely.  XPairtise is an Eclipse plugin created by two computer scientists specifically for implementing pair programming remotely.  This tool allows distributed programmers to initiate pair programming sessions.

Its shared editor allows the navigator to see what the driver is typing. The navigator can’t edit the code, but they can assist the driver by highlighting fragments with the remote cursor and making suggestions using the embedded chat feature.

XPairtise is also integrated with Skype, enabling “over the shoulder” interaction by voice. Moreover, even when a programmer uses XPairtise in single-user mode, their session is available for their team to comment on, meaning the experience of each engineer always benefits the group.

Observability

Communication in software development isn’t just between human beings. It’s vitally important for engineers to know what’s going on inside the systems and platforms they use. Observability is what turns a system from a black box into a problem-solving tool.

The three pillars of observability are logging, tracing, and metrics. If you want to learn more about the role tracing plays, we’ve covered it in a previous post. We’ve also written about converting metrics into insights. Coralogix enhances traditional log and metric management with the power of machine learning.

Moving Out of the Basement

One of the principles of the Agile manifesto is “business people and developers must work together daily throughout the project.”

The importance of team communication in software development goes beyond the team itself. It extends to external teams and clients as well. For this to work, the development team needs to make sure their goals are aligned with the end-user by being in constant contact with interested stakeholders.

A Tale of Two Companies

Consider Oath, a global software company with teams in many different countries, including the US.  Their software projects must have good communication baked in from the start. As software project manager Cristina Otero comments, “it’s really important to communicate goals and expectations to everyone to make sure every piece of the machine is in sync.”

Google is another company that has tackled the problem of inter-stakeholder communication in software development head-on.  They were faced with the problem of making sure their site reliability engineering teams collaborated successfully with product development and service and infrastructure teams.

In their essay on the topic, Google compares their SRE team to a human API. Just as an API mediates between different components of an application, the SRE team needs mediation to facilitate information transfer between different project teams.

Google’s main method of doing this is via production meetings. These are weekly meetings where the SRE team “articulates to itself—and to its invitees—the state of the service(s) in their charge, so as to increase general awareness among everyone who cares, and to improve the operation of the service(s).”

Google’s production meetings are never longer than an hour, enough time to get ideas across without impairing engineer productivity with pointless meetings. While normally face-to-face, two SRE teams can meet by video, something particularly relevant during the Covid crisis.

Doing Agile Remotely

When Covid-19 arrived on the scene last year, companies were forced to shift to remote working almost overnight. Agile software teams were no exception. They’ve been faced with the challenge of taking the organic framework of Agile with its emphasis on problem solving through close-knit interpersonal interaction and making it work in a remote setting. 

In this section, we’ll look at remote communication strategies that facilitate the frictionless communication that software development teams rely on.

Communication through public channels

Agile software development depends on public communication channels to facilitate the diffusion of ideas throughout the entire team. 

In the office, everyone is aware of spoken conversations, as well as being free to examine the whiteboard. At home, everyone is at their own computers in their own rooms. Outside of using a tool such as Slack or Zoom, developers are not automatically aware of what the rest of their team is doing.

Text or Video?

In this situation, Agile’s classic emphasis on face-to-face communication has to be revised.  Kamil Lenonek argues against using video calling in favor of publicly accessible chat forums like Slack.

While video calling may seem a natural substitute for face-to-face interaction, it suffers from being private. Because everyone’s in their own house, there is no chance for the team to overhear conversations and any insights generated can’t diffuse to the rest of the team.

Instead, team members should use communication spaces where the whole team is present, such as Slack channels, even for things that they might normally do on a one-to-one basis. 

Even things that seem routine, such as setting up a piece of software, can be helpful to the entire team. Leonik gives an example of a Facebook message where somebody asked for help setting up Ghost. Leonink himself was looking for a ghost tutorial and this message seemed the perfect cue for one. But the messenger asked for help in private.

The downside of private communication is that any knowledge uncovered can’t be shared.  While many developers are scared to make all their communications public, perhaps because they don’t want to look stupid in front of their colleagues, Leonik urges them to do so.

By making insights publicly accessible, the entire team can benefit. Agile’s frictionless communication, which came perilously close to being snuffed out, can once again resume. 

Keeping it Together Remotely

The keystone of Agile is the ability for groups to solve problems that no one could tackle on their own. Any remote implementation of Agile must recognize that empathy and group morale are just as important as computer skills and Wi-Fi connections.

It’s no easy task.  According to a McKinsey report, remote work results in reduced team cohesion. Their survey found that 80% had impaired work relationships due to reduced communication and 84% reported workplace challenges dragging on for days or more.

There are lots of ways these issues can be addressed. For example, a US bank implemented virtual happy hours where team members can have a drink over Zoom and talk about whatever’s on their minds.

With hard work and dedication, a fully remote software team can communicate effectively. After all, just because a team is distributed, doesn’t mean it’s not close-knit.

Wrapping Up

With its emphasis on face-to-face interaction and collaborative problem solving, communication is fundamental to Agile software development.  We’ve seen how communication is facilitated through Scrum meetings, pair programming, and whiteboarding.

Covid-19 has put remote working in the spotlight. We’ve seen how important public communication is in this context. We’ve also seen the critical need for empathy which remote working has made more important than ever. By embracing the communication principles suggested by Agile, you and your team will continue to excel. 

Basic Principles of Kanban Software Development

Kanban software development is a popular, highly visual framework that falls under the Agile methodology. It requires real-time communication of availability and capacity, allowing full transparency of all work in a team. Tasks are represented visually on a Kanban board by cards. This allows all team members to see the status of every project at any time, including not started, blocked and completed tasks.  

The Kanban methodology helps development teams improve workflow in real-time and complete more work. 

This article will explore the principles of Kanban software development.

What is Kanban?

The name Kanban comes from two Japanese words, ‘Kan’ meaning sign and ‘Ban’ meaning board. It was developed in the late 1940s by Taiichi Ohno, a Japanese engineer, on the shop floor of Toyota. In order to reduce operating costs, they began experimenting with a system that processed small amounts of raw inventory quickly. By matching inventory with demand, the Kanban system helped Toyota achieve higher quality and throughput. 

Today, Kanban software development is used by teams all over the world to complete more work in less time, with a focus on customer value.

Principles of Kanban Software Development 

The methodology is guided by four key principles.

1. Visualize Workflow

Visualizing workflow in Kanban means not only visualizing the process, but visualizing each piece of work (represented by a Kanban card) as it moves through that process. Development Teams find that making the work visible, along with blockers, bottlenecks and queues leads to increased communication, collaboration and problem solving.

Visualizing all work also helps establish accountability and transparency across the team. A blocked work item can be identified with the person or team allocated to fix it. This is invaluable, especially for development teams, which require coordinated efforts to complete their work efficiently.

To begin improving workflows, it is important to visually map the current process. Only then will opportunities for improvement become obvious. For example, the delays of a development team receiving business approval for technical designs are easier to identify and alleviate. Visualization continues once Kanban is implemented and serves as a way to communicate the state of projects and work items.

2. Limit Work in Progress

kanban work in progress

Kanban software development teams can easily get overwhelmed with their long list of work items; such as coding project work, fixing production issues, reducing code debt and maintenance work. So teams often struggle to prioritize work in an organized way.

Therefore, an important concept in the Kanban process is limiting work in process by setting a WIP Limit. By limiting how much unfinished work is in progress, can reduce the time it takes an item to travel through the system. This also avoids problems caused by task switching and reduces the need to constantly re-prioritize items.

Setting a WIP Limit for both individuals and teams can help development teams move faster, reduce errors and collaborate more effectively.

3. Focus on Flow

When the first two principles of Kanban software development are in place, work flows freely. However, you need to focus your attention on interruptions in the flow.

These represent opportunities for additional visualization and process improvements. To assist flow, pull work through the pipeline. Don’t start any new work until a work item is completed. 

Blockers

Sometimes a work item cannot continue to flow through the pipeline due to an unexpected reason, so it becomes blocked. For example, the new project code which is ready for UAT, cannot be deployed to the test environment because of a system failure. Therefore, UAT cannot start. This work item would be marked as ‘Blocked’ on the Board until the issue is resolved. The number of days blocked would be recorded for reporting later. It is a good practice for blockers to be analysed in regular refinement sessions to identify any process improvements.

4. Continuous Improvement

Kanban requires constant monitoring and analysis to look for the next best way to improve development flow and remove waste. Conditions, resources and customer demands change over time, so it is always important to continually assess flow and WIP Limits and look for blockers and bottlenecks that can be removed.

Making small and meaningful changes will help the team perform more efficiently and effectively. 

Kanban Board and Cards

Kanban software development centers on the Kanban board, either a physical one in the office located in a prominent position for the whole team to view and access or as in recent times, a virtual one using software. The main purpose of the board is to create a shared understanding of value flow. However, the board does more than visualize workflow steps, it also highlights where bottlenecks are formed in the process.

To visualize your work, identify distinct steps or stages your work goes through as it moves from ‘To Do’ to ‘Done’. These will become columns on your board. Some boards may just have ‘To Do’, ‘In Progress’ and ‘Done’ columns. For a Kanban software development team, some examples of the columns created on the board include: ‘Design’, ‘Code’, ‘Unit Test’, ‘Code Review’, ‘UAT’ & ‘Prod’.

Each project or initiative should be allocated a horizontal swimlane. 

Pro-tip: If using a physical board, it is beneficial to use a magnetic one so magnets can be used as well as post it notes, which tend to fall off after a while. 

Kanban Cards

Every work item is visualized as a separate card placed on the Kanban board. Cards are moved from left to right to show progress and to help coordinate teams performing the work. The main purpose of the card is to give team members something to track as it moves throughout the workflow. This makes invisible work visible. 

Each card can contain critical information about the work item it represents. If a physical board is being used, then there is a limit of how much information can be displayed on each card. One way around this is for a team to also use issue tracking software like Jira and to raise a digital ticket too for each work item. This method allows for the meaty details to be contained in the Jira ticket and for the bare essentials details to be displayed on the card like Jira id, work item name and an estimate of effort. The advantage of tracking software is the history of a work item is stored for future reference and all team members can access and update the tickets.  

Another advantage is it allows Agile methodology to be performed so epics can be created to group work items, for example, in Jira, an epic is created for each project. Each epic can then be assigned a swimlane on the physical board. 

Avatars and Markers

There are different flavors of Kanban Boards and Cards. One flexible way is to assign Avatars for each team member often chosen by each individual and maybe as part of a theme. The avatars can be attached to magnets and are separate from the Kanban Cards. Each member only has one avatar as they can only work on one piece of work at a time. Their avatar is positioned on top of the card they are currently working on.

Magnetic Markers can also be used to highlight work on a card since the last stand up meeting i.e. a small green marker for any work completed and red markers for any blockers. These quickly allow the presenter at the stand up meetings to identify the progress or lack of progress and any blockers of work since the last stand up meeting for discussion. Once discussed, the green markers are removed from the cards, ready for today’s work to be tracked. 

Meetings

Once a board has been created, it becomes the centerpiece for at least two regular meetings:

  • Daily stand ups: Involves all team members, to discuss progress since the last stand up. These are usually scheduled in the mornings.
  • Refinement sessions: Involves all tech leads to discuss WIP Limits, board policies, blockers, bottlenecks and ‘stale’ work that has had no recent movement. These should be weekly or bi-weekly.

These meetings aim to focus on optimizing the flow that can help development teams stay productive and efficient by moving existing cards off the board. 

Summary

Kanban software development focuses on visualizing the entire project on boards in order to increase project transparency and collaboration between team members. The aim is to control and manage the flow of work, represented by Kanban cards so that the number of work entering the process matches those being completed.

Follow the four principles of the Kanban process for your team to consistently deliver high quality software.

More Changes Mean More Challenges for Troubleshooting

The widespread adoption of Agile methodologies in recent years has allowed organizations to significantly increase their ability to push out more high quality software. The new fast-paced CI/CD solutions pipeline and lightweight microservices architecture enable us to introduce new updates at a rate and scale that would have seemed unimaginable just a few years ago.

Previous development practices revolved heavily around centralized applications and infrequent updates that were shipped maybe once a quarter or even once a year. 

But with the benefits of new technologies comes a new set of challenges that need to be addressed as well. 

While the tools for creating and deploying software have improved significantly, we are still in the process of establishing a more efficient approach to troubleshooting. 

This post takes a look at the tools that are currently available for helping teams keep their software functioning as intended, the challenges that they face with advances in the field, and a new way forward that will speed up the troubleshooting process for everyone involved.

Challenges Facing Troubleshooters 

On the face of it, the ability to make more software is a positive mark in the ledger. But with new abilities come new issues to solve. 

Let’s take a look at three of the most common challenges facing organizations right now.

The Move from the Monolith to Microservices 

Our first challenge is that there are now more moving parts where issues may arise that can complicate efforts to fix them quickly and efficiently. 

Since making the move to microservices and Kubernetes, we are now dealing with a much more distributed set of environments. We have replaced the “monolith” of the single core app with many smaller, more agile but dispersed apps. 

Their high level of distribution makes it more difficult to quickly track down where a change occurred and understand what’s impacting what. As the elements of our apps become more siloed, we lose overall visibility over the system structure.

Teams are Becoming More Siloed

The next challenge is that with more widely distributed teams taking part in the software development process, there are much fewer people who are significantly familiar with the environment when it comes to addressing problems impacting the products. 

Each team is proficient in their own domain and segment of the product, but are essentially siloed off from the other groups. This leads to a lack of knowledge sharing across teams that can be highly detrimental when it comes time for someone to get called in to fix an issue stemming from another team’s part of the product. 

More Changes, More Often

Then finally for our current review, is the fact that changes are happening far more often than before. 

Patches no longer have to wait until it’s Tuesday to come out. New features on the app or adjustments to infrastructure like load balancing can impact the product and require a fix. With all this going on, it is easy for others to not be aware of these changes.

The lack of communication caused by siloes combined with the increase in the number of changes can create confusion when issues arise. Sort of like finding a specific needle in a pile of needles. It’s enough to make you miss the haystack metaphor. 

In Search of Context

In examining these challenges, we can quickly understand that they all center around the reality that those tasked with troubleshooting issues lack the needed context for fixing them. 

Developers, who are increasingly being called upon to troubleshoot, lack the knowledge of what has changed, by who, what it impacts, or even where to start looking in a potentially unfamiliar environment.

Having made the move to the cloud, developers have at their disposal a wide range of monitoring and observability tools covering needs related to tracing, logs, databases, alerting, security, and topology, providing them with greater visibility into their software throughout the Software Development Lifecycle.

When they get the call that something needs their attention, they often have to begin investigating essentially from scratch. This means jumping into their many tools and dashboards, pouring over logs to try and figure out the source of the problem. While these monitoring and observability tools can provide valuable insights, it can still be difficult to identify which changes impacted which other components.

Some organizations attempt to use communication tools like Slack to track changes. While we love the scroll as much as anyone else, it is far from a comprehensive solution and still lacks the connection between source and impact. Chances are that they will need to call in for additional help in tracking down the root of the problem from someone else on their team or a different team altogether.

In both cases, the person on-call still needs to spend significant and valuable time on connecting the dots between changes and issues. Time that might be better spent on actually fixing the problem and getting the product back online.  

Filling in the Missing Part of the DevOps Toolchain

The tooling available to help identify issues are getting much better at providing visibility, including to provide on-call responders with the context that will help them get to the source of the problem faster. 

Moving forward, we want to see developers gain a better, holistic understanding of their Kubernetes environments. So even if their microservices are highly distributed, their visibility over them should be more unified. 

Fixing problems faster means getting the relevant information into the hands of whoever is called up to address the issue. It should not matter if the incident is in a product that they themselves built or if they are opening it up for the first time. 

Reducing the MTTR relies on providing them with the necessary context from the tools that they are already using to monitor their environment. Hopefully by better utilizing our existing resources, we can cut down on the time and number of people required to fix issues when they pop up — allowing us to get back to actually building better products.