Generate summary with AI

Most configuration drift is quiet when it starts. A service getting re-enabled during an emergency fix doesn’t get caught unless someone specifically looks for it. Weeks later, two servers built from the same image behave differently under load, and nobody can say why. Uptime Institute’s 2026 Annual Outage Analysis found that configuration or change-management failure was the most commonly cited cause of major network-related outages, reported by 41% of organizations that had experienced a significant, serious, or severe network outage in the previous three years.
But preventing configuration drift isn’t as simple as auditing harder after the fact. Here’s everything you need to know to actually prevent it.
Why configuration drift causes problems
The hardest part of identifying drift is that it rarely looks like a fault. A disabled service and a policy applied to one server but not its peers can both leave a system running normally. The difference only matters when something depends on it, like a failover node that doesn’t behave like the primary, an application that works on one host and fails on another, or a patch that installs cleanly everywhere except the machines nobody realized were different.
Since the affected system still appears correctly configured, troubleshooting starts in the wrong place. Technicians check the application, the network path, or the latest deployment before anyone compares the system against what it’s supposed to be. That’s also why a working configuration isn’t proof of a correct one. A change can be technically functional and still be drift, which is why changes can go unnoticed until they cause an incident.
Where drift actually comes from
Configuration drift is any unapproved divergence between a system’s actual state and its formally approved baseline. That definition shifts the question from “is this different?” to “was this difference authorized?” A prevention model that treats every difference as a problem will drown in noise. One that compares against a controlled reference can separate legitimate change from drift.
Most production drift traces back to a handful of sources, including:
- Direct administrator changes made in a live session rather than through a managed path
- Emergency fixes applied under pressure and never documented or reverted
- Inconsistent deployments, where the same build lands differently on different systems
- Patches or upgrades that alter defaults as a side effect
- Vendor software that modifies services, settings, or permissions during installation or updates
- Misconfigured automation that applies the wrong state (or the right state to the wrong targets)
These sources don’t all need the same control. Procedural controls such as approval, testing, and documentation work for infrequent changes, where the process can reliably govern execution. Changes that are frequent, privileged, security-sensitive, or able to affect multiple systems at once need automated guardrails, because by the time a procedure catches them, the change has already spread. Sorting each source onto the right side of that line determines where automation has to do the enforcing.
The boundary is change control. A change becomes part of the intended state once it has been authorized, tested, documented, implemented, and incorporated into the updated baseline. Anything that alters a system outside that process is drift.
What to establish before you enforce anything
Drift prevention tools are only as good as the definitions behind them. Automation can compare a system against a baseline and reapply the approved state, but it can’t decide what “approved” means, who’s allowed to change it, or which systems it can actually reach.
Do these three main things first:
1: Define the baseline in enforceable terms
An enforceable baseline spells out exactly what the approved state is for each asset class, whether that’s domain controllers, standard workstations, web servers, or cloud virtual machines. Vague standards like “hardened” or “up to date” can’t be checked by a tool.
The technical state should cover:
- Operating system and application versions, including required patch levels
- Enabled services, open ports, and permitted protocols
- Local and privileged accounts
- Security settings and network configuration
- Installed software and required management agents
The baseline also needs governance details that tell automation what not to treat as drift, such as:
- Approved exceptions and machine-specific parameters
- A named owner for each baseline
- Version history showing what changed and when
- The change-control process used to modify it
Without that second list, configuration drift management breaks down fast. A tool that doesn’t know a particular server legitimately runs an extra service will either flag it on every check or just remove it without you knowing.
» Make sure you know the real cost of IT downtime
2: Set permissions and change boundaries that match real work
The permission model decides if drift can happen in the first place. Administrators should hold only the access their operational responsibilities require, with privileged roles and responsibilities kept separate from the accounts they use for everyday work. Someone who needs to restart services on application servers doesn’t need rights to modify domain security policy.
Change boundaries then define which configuration items can be modified during normal administration and which need elevated access or formal approval. High-impact changes involving security controls, identity, networking, production policies, or the baselines themselves should require independent authorization before implementation.
“The goal isn’t to make every change slow. Routine, low-risk tasks can run under preapproved change models with constrained permissions, so administrators can work efficiently without holding unrestricted configuration access.”
Dominique Locksley, Linux System Administrator at Adapt IT Holdings Limited
That balance matters, because if the approval process is heavier than the work, people find ways around it. It’s those workarounds that become drift or shadow IT.
3: Confirm every system can actually be managed
Enforcement across a mixed fleet only works if every system in scope is known, reachable, and supported by a management method. Before switching anything on, confirm four things:
- A complete asset inventory: Every system should be identifiable by owner, operating system, version, location, and management capability.
- Dependable connectivity for remote endpoints: Laptops that spend weeks off the corporate network need a management channel that reaches them wherever they are, or they fall out of scope without anyone noticing.
- Appropriate access for cloud resources: Cloud infrastructure needs API permissions for the IT management tooling and an authoritative configuration source, so the tool knows which definition is the real one.
- Documented handling for legacy systems: Older servers and legacy IT systems that can’t run modern AI agents or accept policy enforcement need recorded exceptions and compensating controls.
Only start automating once these prerequisites are verified. Unreachable or unsupported assets don’t show up as failures in a compliance report. They simply don’t show up at all, and a dashboard reporting full compliance across the systems it can see creates a false sense that the whole environment is aligned.
» Learn more about IT asset discovery and IT asset management
How to prevent configuration drift
Preventing drift takes three things working together:
- A single authoritative definition with only one approved way to change it
- Automation that detects divergence from that definition and corrects it
- Enforcement that holds up through the events that change the most systems at once
Each of these depends on the one before it, so here’s how you guarantee that the requirements are all met.
Give every configuration one source and one path to change
Drift often starts with copies. When teams keep their own versions of server images or configuration files, those versions evolve separately until nobody can say which one is correct. The fix has two parts: a single baseline, and a single way to change it.
Part 1: Standardize on one baseline per system class
Maintain one version-controlled base image or configuration for each supported system class. Everything in it falls into one of two groups:
- Shared requirements belong in the baseline: Security settings, management agents, logging, required services, and patch management configuration.
- Legitimate differences come in as documented variables or approved overlays: Hostnames, network addresses, environment settings, capacity, and application roles.
Since differences never become edits to the baseline itself, every exception stays explicit and reproducible. A change to shared configuration is made once and reaches every system that inherits it.
Part 2: Make the approved path the only path
A single source only holds if it’s also the only way in. The most effective safeguard is making the approved path easier than a local change with these methods:
- Route production changes through managed tooling: Changes should go through software deployment pipelines or configuration-management tools, with role-based permissions limiting who can open an unrestricted admin session at all.
- Let the repository enforce review: GitHub protected branches can require pull request reviews, passing status checks, signed commits, and restricted push access before a change reaches the protected branch.
- Scope service accounts tightly: Deployment scripts and vendor software should run under accounts limited to what their function requires, so an installer can’t rewrite settings it has no reason to touch.
- Protect configuration files on the system: File permissions, policy enforcement, and infrastructure integrity monitoring make critical files harder to modify outside the approved path.
- Detect what you can’t prevent: Where a local change is unavoidable, it should at least be caught, and that’s where automated enforcement takes over.
Detect and correct drift automatically
Automated enforcement turns the baseline into a closed loop that runs on a recurring schedule, similar to this:
- Collect the current configuration state automatically from all managed servers, endpoints, network devices, and cloud resources
- Compare the collected state against the approved baseline
- Classify each difference as an approved change, a documented exception, or unauthorized drift
- Correct confirmed drift through the tool that owns that configuration domain
- Recheck the corrected system to verify the baseline has actually been restored
- Fold approved changes into the baseline only after they’ve been tested and authorized
- Review drift events for recurring causes, such as manual administration, failed automation, or unmanaged assets
The last step keeps the loop from becoming a treadmill. Correcting the same setting on the same servers every week means something upstream keeps reintroducing it.
For endpoints and servers, Atera’s RMM covers the collect, compare, and correct steps in these ways:
- On-demand checks: Remote script execution runs baseline checks or corrections across selected devices or device groups whenever you need them.
- Recurring checks: Scripts that need to run regularly can be scheduled through IT automation profiles.
- Auto-healing scripts: These run automatically when a monitored condition crosses its threshold, such as a critical service stopping, and restore the expected state without waiting for a technician.
- Alert-driven tickets: Threshold-based alerts can create tickets automatically, so anything scripts can’t resolve reaches the queue with context attached.
- Script generation: When writing checks is the bottleneck, AI Copilot generates scripts from a plain-language description of what needs to be verified or corrected.
» Learn more about automating your ticket escalation process and automating ticket routing
Match controls to each infrastructure layer
The baseline principles stay the same across the environment, but network appliances, endpoints, and cloud infrastructure expose configuration differently and change on different lifecycles. Enforcement has to use controls native to each layer:
- Network appliances: Controlled configuration templates, centralized management, configuration backups, and command-level change auditing.
- Operating system endpoints: Configuration-management agents, device policies, security baselines, and automated remediation.
- Cloud infrastructure: Infrastructure-as-code, policy engines, API-based compliance checks, and restricted console access.
Network devices need particular attention because they’re often configured one at a time, which turns each device into its own accidental source of truth.
Keep enforcement consistent as the fleet changes
Steady-state enforcement handles gradual drift relatively easily. The bigger risk comes from events that change many systems at once, and from a fleet that outgrows device-by-device management.
Onboarding, OS upgrades, and fleet-wide patching introduce more change in a short window than almost anything else. Building the baseline into each workflow instead of correcting systems afterward can help you mitigate some of the drift.
Make sure you focus on:
- Onboarding: Apply role-based configurations automatically at provisioning, so new devices start compliant.
- OS upgrades: Test each upgrade against the baseline before deployment, then run compliance validation to catch changed defaults or settings the new version no longer supports.
- Fleet-wide patching: Deploy in staged rings with defined maintenance windows and post-installation checks.
Systems that fail validation at any point should be isolated for remediation. Letting them run “temporarily” in an exception state is how one-off exceptions harden into permanent drift.
Enforce by group (not device)
As the ratio of endpoints to technicians widens, configuration has to be applied by group rather than device by device. In Atera, the asset hierarchy provides that structure:
- Inheritance and precedence: Configuration policies assigned at the customer or site level are inherited by the folders and agents beneath it. A folder policy takes precedence over the parent-level policy, and a policy assigned directly to an agent takes precedence over both. Folders organized by function, location, or technical requirement let different baselines apply to different groups, and the hierarchy doubles as a record of which systems carry an approved exception.
- Permission boundaries: Assigning policies at the customer or folder level requires full admin access, while technicians can assign policies only to individual agents.
Configuration policies cover update and restart behavior. Keep three details in mind:
- Timing: Policies assigned at the customer or folder level are reapplied roughly every 12 hours. Policies assigned directly to an agent apply immediately.
- Reboot setting: Policies depend on the “Reboot if needed” option being enabled in the relevant IT automation profile.
- GPO precedence: Policies don’t override settings enforced by a domain Group Policy Object. If GPO already manages Windows Update, settle which tool owns those settings first.
IT automation profiles handle recurring patching and maintenance through the same hierarchy. Assigning profiles with staggered schedules to different folders gives you a practical way to stage patch rings.
» Here’s how to simplify group policy management with Atera
Drift prevention that scales with your fleet
Configuration drift doesn’t stop once the environment is clean. Every emergency fix, patch cycle, new device class, and vendor update is another chance for systems to slip away from their approved state. A baseline nobody reviews can quietly become the problem itself.
Atera gives IT teams and MSPs the pieces to run that loop from one place. Configuration policies set operate at the site or customer level, with exceptions assigned at the folder or agent level, while IT automation profiles keep patching and scheduled maintenance on track.
Related Articles
How to reduce alert fatigue in healthcare IT
In a hospital, a dismissed alert isn't a delayed ticket. It can be a patient deteriorating while the monitor keeps sounding. Cutting alert volume won't fix that on its own, and switching alerts off can make it worse.
Read nowHow to reduce alert fatigue across your IT team
Your technicians aren't ignoring alerts because they're careless. They're ignoring them because most alerts have taught them to. Once the stream stops being trustworthy, real failures slip past with the noise. Fixing it means deciding what deserves an interruption, tuning out transient spikes, collapsing alert storms, and automating the fixes you've run a hundred times.
Read nowHow to close the IT skills gap on your team
Course completions don't close skills gaps. Technicians finish the training, then hand the first unfamiliar failure straight back to a senior engineer. Real capability gets built on live tickets, incidents, and maintenance windows, with guidance that fades as competence grows.
Read nowHow to calculate cost per ticket (and why most teams get it wrong)
Most cost-per-ticket figures are wrong before anyone reads them. Missing overhead, tickets that were opened but never closed, spam and duplicates padding the count, and one month's costs divided by another month's tickets all make support look cheaper than it is. Fix both sides of the division and the number finally shows where technician time and money actually go.
Read nowEndless IT possibilities
Boost your productivity with Atera’s intuitive, centralized all-in-one platform










