SSM Agent Is Not Online: Fix "Unable to Connect to a Systems Manager Endpoint" (2026)

Logeshwaran
—

"SSM Agent is not online. The SSM Agent was unable to connect to a Systems Manager endpoint to register itself with the service" means AWS Systems Manager has never heard from the agent on your EC2 instance, or has stopped hearing from it. For an instance to show up in Systems Manager, four things must all be true: the SSM Agent is installed and running, the instance has IAM permissions to talk to Systems Manager, it can reach the ssm and ssmmessages endpoints over HTTPS (port 443), and it can read its own instance metadata. Most guides jump straight to security groups and VPC endpoints. Here is what they skip: on several common operating systems, the agent is not there at all. AWS's own Red Hat Enterprise Linux 8 and 9 AMIs ship without the SSM Agent, so a RHEL instance can sit in a perfectly configured network and never come online.

Ethan looks after AWS for a few local businesses, and one of them, an accountant two doors down from Jake's phone repair shop, needed a Red Hat server for a tax application. Ethan launched it in a private subnet next to an Amazon Linux server that had been managed by Systems Manager for a year, same VPC, same security group, same IAM role. The Amazon Linux server showed up in Fleet Manager in a minute. The Red Hat one never did, and Session Manager greeted him with the message above. He spent an hour checking VPC endpoints that were already fine. The problem was simpler: there was no agent to come online. This page is the order of checks he uses now, from the two-minute tests that find the cause, through each of the possible fixes, to confirming the instance is online and connecting to it.

⚡ Quick Answer

• Fastest diagnosis → Run the AWSSupport-TroubleshootManagedInstance runbook, or ssm-cli get-diagnostics on the instance. Two tools.

• Agent missing? → Not preinstalled on RHEL 8 and 9 AMIs, among others. Install it: sudo dnf install -y https://s3.amazonaws.com/ec2-downloads-windows/SSMAgent/latest/linux_amd64/amazon-ssm-agent.rpm. Install and start.

• Permissions → An instance profile with AmazonSSMManagedInstanceCore, or Default Host Management Configuration (IMDSv2 only). IAM fix.

• Network → Outbound 443 to ssm and ssmmessages, through a NAT gateway or VPC endpoints with private DNS on, and the endpoint security group allowing 443 in. Network fix.

• Is ec2messages still needed? → Optional in Regions launched before 2024, and not supported in Regions launched in 2024 or later. Endpoint list.

• Fixed it, still offline? → A failing agent backs off to retrying about once an hour. Restart it: sudo systemctl restart amazon-ssm-agent. Confirm Online.

Systems Manager needs no inbound ports on the instance. The agent makes every connection outward.

If this message has you staring at a console you thought you understood, take heart: there are only four requirements, and the tools below tell you which one is missing. You do not need to understand every piece of Systems Manager to fix it, and nothing about the instance itself is broken.

🧭 NEW HERE? READ THESE FIRST

If Systems Manager or the AWS networking words are new to you, these five pages make the rest easier:

📌 Bookmark this; the network section leans on the third and fourth.

What "SSM Agent is not online" actually means

Systems Manager does not reach into your instances. A small program on each instance, the SSM Agent, connects out to the Systems Manager service, registers the instance, and keeps checking in. Once registered, the instance is a managed node, and tools such as Session Manager, Run Command, Patch Manager and Fleet Manager can work with it. The service confirms each managed node is available by checking in with it every five minutes.

When registration never happens, or the check-ins stop, you see one of three symptoms:

  • Session Manager shows "SSM Agent is not online. The SSM Agent was unable to connect to a Systems Manager endpoint to register itself with the service", followed by hints about IAM, security groups and VPC endpoints.
  • Fleet Manager lists the instance with a ping status of Connection Lost.
  • Fleet Manager does not list the instance at all.

All three come from the same four requirements. AWS lists them for every managed node:

  1. The SSM Agent is installed and running on a supported operating system.
  2. An IAM instance profile, or Default Host Management Configuration, gives the instance permission to talk to Systems Manager.
  3. The agent can connect to a Systems Manager endpoint to register itself.
  4. The instance can read its own instance metadata (IMDS), which is how the agent learns its identity and gets credentials.

The hint text in the console mentions only IAM and the network, which is why people spend an hour on security groups when the agent was never installed.

Two small questions come up alongside this one. What does SSM stand for? Simple Systems Manager: the service was first called Amazon Simple Systems Manager, later Amazon EC2 Systems Manager, and the old abbreviation survived in the agent's name, the CLI command aws ssm and many resource names. And when people search for "SSM Agent status offline", they usually mean one of the symptoms above: a ping status of Connection Lost, or an instance missing from the list.

Find the cause in two minutes: the runbook and ssm-cli

AWS gives you two tools that check all four requirements for you. Use whichever fits your situation.

The runbook, when you cannot log in to the instance. AWSSupport-TroubleshootManagedInstance is an AWS-owned Automation runbook that checks why an EC2 instance is not reporting as managed. It reviews the instance's VPC configuration, security group rules, VPC endpoints, network ACLs, route tables and IAM instance profile, without needing access to the instance. Open it in Systems Manager, Automation, and run it with the instance ID, or from the CLI:

aws ssm start-automation-execution \
  --document-name AWSSupport-TroubleshootManagedInstance \
  --parameters "InstanceId=i-0123456789abcdef0" \
  --region us-east-1

It cannot see inside the operating system, so it will not tell you that the agent is missing; it checks everything around it. Note also that it does not evaluate IPv6 rules.

ssm-cli, when you can log in. SSM Agent 3.1.501.0 and later include ssm-cli, a diagnostic tool that checks the instance from the inside: instance metadata, connectivity to each endpoint, credentials, proxy settings, the agent service and its version. Connect by SSH, EC2 Instance Connect or the serial console, then run:

# Linux and macOS
sudo ssm-cli get-diagnostics --output table

# Windows (from C:\Program Files\Amazon\SSM)
.\ssm-cli.exe get-diagnostics --output table

The table lists each check with Success, Failed or Skipped and a note: EC2 IMDS, connectivity to the ssm, ec2messages and ssmmessages endpoints, connectivity to s3, kms, logs and monitoring, AWS credentials, the agent service, the proxy configuration, the SSM Agent version and, on Windows, the sysprep state. The s3, kms, logs and monitoring checks are for optional features; a failure there does not stop the instance registering. If ssm-cli is not found at all, the agent is either missing or older than 3.1.501.0, which already answers the question.

Ethan could reach the Red Hat server over SSH through a bastion. ssm-cli returned "command not found", and systemctl status amazon-ssm-agent said the service did not exist. Two minutes, one answer.

Cause 1: the SSM Agent is not installed, or not running

The SSM Agent is preinstalled on some AMIs from AWS and trusted third parties, not on all of them. According to AWS's list, you will usually find it on:

  • Amazon Linux 2 and Amazon Linux 2023, including the ECS-optimized and EKS-optimized AMIs
  • Ubuntu Server 18.04, 20.04, 22.04 LTS, 24.04 LTS and 25.04
  • SUSE Linux Enterprise Server 15.3 and later
  • AlmaLinux
  • Windows Server 2016, 2019, 2022 (not Nano) and 2025, and Windows Server 2012 R2 AMIs from November 2016 onward
  • macOS 13 Ventura, 14 Sonoma and 15 Sequoia

Red Hat Enterprise Linux is not on the list. AWS states plainly that its RHEL 8 and 9 AMIs do not come with the SSM Agent preinstalled. Debian, Oracle Linux, Rocky Linux and AMIs from AWS Marketplace or the community catalog are also ones to check rather than assume. Even on a listed AMI, the agent may be an older version, or installed but stopped.

Check whether it is installed and running:

Operating systemCheckStart it
Amazon Linux, RHEL, SLES, Debian, Oracle Linuxsudo systemctl status amazon-ssm-agentsudo systemctl enable amazon-ssm-agent then sudo systemctl start amazon-ssm-agent
Ubuntu (snap install)sudo systemctl status snap.amazon-ssm-agent.amazon-ssm-agent.servicesudo snap start amazon-ssm-agent
Windows ServerGet-Service AmazonSSMAgentStart-Service AmazonSSMAgent
macOSRead /var/log/amazon/ssm/amazon-ssm-agent.logSee AWS's macOS install steps

Install it on RHEL 8, 9 or 10. Python 2 or 3 must be present. Then run the command for your architecture; the ec2-downloads-windows folder in the URL looks wrong, but it is the correct global location for these files:

# x86_64
sudo dnf install -y https://s3.amazonaws.com/ec2-downloads-windows/SSMAgent/latest/linux_amd64/amazon-ssm-agent.rpm

# ARM64 (Graviton)
sudo dnf install -y https://s3.amazonaws.com/ec2-downloads-windows/SSMAgent/latest/linux_arm64/amazon-ssm-agent.rpm

sudo systemctl status amazon-ssm-agent

In a private subnet without internet access, that download needs a route to Amazon S3: a NAT gateway, or an S3 gateway endpoint with a policy that allows the bucket. For many instances, AWS recommends installation files from a bucket in your own Region, and the agent can be baked into a custom AMI so new instances start with it.

Install it at launch. To never meet this error on RHEL again, put the install command in the instance's user data, so it runs on first boot:

#!/bin/bash
dnf install -y https://s3.amazonaws.com/ec2-downloads-windows/SSMAgent/latest/linux_amd64/amazon-ssm-agent.rpm
systemctl enable --now amazon-ssm-agent

Install SSM Agent on Windows Server

Recent Windows Server AMIs from AWS include the agent, but a custom or older image may not. In PowerShell as an administrator, AWS's documented install downloads the latest 64-bit installer, runs it silently and restarts the service:

[System.Net.ServicePointManager]::SecurityProtocol = 'TLS12'
$progressPreference = 'silentlyContinue'
Invoke-WebRequest `
    https://s3.amazonaws.com/ec2-downloads-windows/SSMAgent/latest/windows_amd64/AmazonSSMAgentSetup.exe `
    -OutFile $env:USERPROFILE\Desktop\SSMAgent_latest.exe
Start-Process -FilePath $env:USERPROFILE\Desktop\SSMAgent_latest.exe -ArgumentList "/S" -Wait
rm -Force $env:USERPROFILE\Desktop\SSMAgent_latest.exe
Restart-Service AmazonSSMAgent

For a direct download in a browser, the same installer lives at https://s3.amazonaws.com/ec2-downloads-windows/SSMAgent/latest/windows_amd64/AmazonSSMAgentSetup.exe. These EC2 steps are for EC2 instances only; for on-premises Windows servers AWS provides a separate setup tool as part of hybrid activation.

Install SSM Agent on Ubuntu

On Ubuntu Server 18.04 through 25.04 the agent is a snap package, so the commands differ from the rest of Linux:

sudo snap install amazon-ssm-agent --classic
sudo snap list amazon-ssm-agent          # installed version
sudo snap start amazon-ssm-agent
sudo snap services amazon-ssm-agent      # is it active?

Check the SSM Agent version and whether it is running

Fleet Manager shows each managed node's agent version in its SSM Agent version column, and describe-instance-information returns it as AgentVersion. For an instance that is not online yet, check from inside the operating system:

Operating systemGet the installed SSM Agent version
Amazon Linux 2, CentOS, older RHELyum info amazon-ssm-agent
Amazon Linux 2023, RHEL 8 and laterdnf info amazon-ssm-agent
Ubuntu (snap)sudo snap list amazon-ssm-agent
Debianapt list amazon-ssm-agent
SUSEzypper info amazon-ssm-agent
macOSpkgutil --pkg-info com.amazon.aws.ssm
Windows Server& "C:\Program Files\Amazon\SSM\amazon-ssm-agent.exe" -version

Compare the number with the latest release in the agent's release notes on GitHub, and with the minimums this page mentions: 2.3.68.0 for Session Manager, 3.1.501.0 for ssm-cli, 3.2.582.0 for Default Host Management Configuration, and 3.3.40.0 for the switch from ec2messages to ssmmessages.

Keep it current. AWS releases agent updates as Systems Manager gains features, and an old agent can stop newer tools from working. Several features in this guide need a minimum version: ssm-cli needs 3.1.501.0, and Default Host Management Configuration needs 3.2.582.0.

Cause 2: the instance has no permission to talk to Systems Manager

An installed, running agent still cannot register without credentials. There are two ways to give them, and you need one.

Option A: an instance profile with the managed policy. Create an IAM role that EC2 can assume, attach the AWS managed policy AmazonSSMManagedInstanceCore, and attach the role to the instance as its instance profile:

  1. In IAM, create a role for the AWS service EC2.
  2. Attach the AmazonSSMManagedInstanceCore policy.
  3. Name the role, for example EC2-SSM-Core, and create it.
  4. In EC2, select the instance, choose Actions, Security, Modify IAM role, pick the role and save.
  5. Restart the agent, or wait a few minutes, for it to pick up the new credentials.

If the instance already has a role for your application, add the policy to that role rather than replacing it; an instance can have only one instance profile. Our instance profile guide explains how the role reaches the instance.

Option B: Default Host Management Configuration. This account-level setting lets Systems Manager manage every EC2 instance in a Region without an instance profile on each one. Turn it on in Fleet Manager, under Account management, Configure Default Host Management Configuration, and choose the default role. It has firm requirements:

  • Instances must use IMDSv2. IMDSv1 is not supported.
  • The SSM Agent must be 3.2.582.0 or later.
  • It must be turned on in each Region separately.
  • It can take up to 30 minutes for instances to start using it.
  • If an instance profile already allows ssm:UpdateInstanceInformation, the agent uses the instance profile instead.

The registration it creates is stored locally under /lib/amazon/ssm on Linux or C:\ProgramData\Amazon on Windows. Deleting those files breaks the instance's ability to get Default Host Management credentials, and you would need an instance profile or a new instance.

One trap: credentials left on the instance. The agent does not go straight to the instance profile. It looks for credentials in the same order as the AWS SDKs: environment variables first, then a shared credentials file, then an ECS task role, then the instance profile, and the Default Host Management role last. On Linux and macOS the agent runs as root, so the file that matters is /root/.aws/credentials. If someone once ran sudo aws configure on the server with an old access key, or a deleted user's key, the agent tries those keys and never reaches the perfectly good role you attached. Remove or rename that file, or the stray environment variables, and restart the agent.

Ethan has met this exactly once, on a server a previous contractor had used as a workstation. The role was right, the network was right, and the agent was still signing in as a user who had left the company a year earlier.

Cause 3: the agent cannot reach the Systems Manager endpoints

The agent connects out over HTTPS, port 443. Systems Manager never connects in, so you do not open any inbound port on the instance for it. What you need is a path out to the endpoints, and there are two ways to provide it.

Path 1: through the internet. In a public subnet, the instance needs a public IP address and a route to an internet gateway. In a private subnet, it needs a route to a NAT gateway. Either way, the security group and network ACLs must allow outbound 443. Our private subnet guide covers the NAT route step by step.

Path 2: through VPC endpoints. For instances with no internet access at all, create interface VPC endpoints for Systems Manager. Traffic then stays on the AWS network, and no NAT gateway is needed. Four details decide whether they work:

  • Create the right endpoints (listed in the next section), each associated with a subnet in every Availability Zone your instances use.
  • Private DNS must be on for each interface endpoint. It is on by default, but it can be switched off, and without it the agent still resolves the public address and fails.
  • The VPC must have enableDnsSupport and enableDnsHostnames turned on. Without them, private DNS for endpoints does not work.
  • The endpoints' security group must allow inbound 443 from the instances' subnet. This is the rule people forget, because it is on the endpoint, not the instance.

If you run your own DNS servers, add a conditional forwarder that sends amazonaws.com queries to the VPC's Amazon DNS server, or the endpoint names will not resolve to the private addresses.

Check the DNS attributes from the CLI:

aws ec2 describe-vpc-attribute --vpc-id vpc-0123456789abcdef0 --attribute enableDnsSupport
aws ec2 describe-vpc-attribute --vpc-id vpc-0123456789abcdef0 --attribute enableDnsHostnames

# turn them on if needed (one attribute per call)
aws ec2 modify-vpc-attribute --vpc-id vpc-0123456789abcdef0 --enable-dns-support "{\"Value\":true}"
aws ec2 modify-vpc-attribute --vpc-id vpc-0123456789abcdef0 --enable-dns-hostnames "{\"Value\":true}"

From the instance, a quick test of whether the endpoint resolves privately and answers on 443:

nslookup ssm.us-east-1.amazonaws.com        # should return private 10.x / 172.x / 192.168.x addresses
curl -sS -o /dev/null -w "%{http_code}\n" https://ssm.us-east-1.amazonaws.com
curl -sS -o /dev/null -w "%{http_code}\n" https://ssmmessages.us-east-1.amazonaws.com

Any HTTP status code back, even a 4xx, means the connection worked; a timeout means the path is blocked.

Ports at a glance. One of the most-viewed Systems Manager questions on developer forums simply asks which protocol and port the agent uses. The whole answer fits in three lines:

  • Outbound HTTPS, TCP 443, from the instance to the Systems Manager endpoints (and to S3, KMS and CloudWatch Logs if you use those features).
  • Instance metadata at 169.254.169.254, which never leaves the instance's host, but which a local firewall rule can still block.
  • No inbound ports at all. Not 22, not 3389, not 443. Session Manager shells, Run Command and even port forwarding ride on the connection the agent opened.

Which endpoints Systems Manager needs in 2026

Interface endpointNeeded forRequired?
com.amazonaws.region.ssmThe Systems Manager service itselfYes
com.amazonaws.region.ssmmessagesAgent messaging, Run Command, Session ManagerYes
com.amazonaws.region.ec2messagesThe older message delivery serviceOptional in Regions launched before 2024; not supported in Regions launched in 2024 or later
com.amazonaws.region.s3Agent updates, patching, scripts and output in S3Recommended; a gateway endpoint is enough
com.amazonaws.region.kmsKMS encryption for Session Manager or Parameter StoreOnly if you use it
com.amazonaws.region.logsSending session or command logs to CloudWatch LogsOnly if you use it
com.amazonaws.region.ec2VSS-enabled snapshots on WindowsOnly if you use it

The ec2messages row is where older guides mislead. Since agent version 3.3.40.0, Systems Manager uses ssmmessages instead of ec2messages whenever it is available. In Regions launched before 2024, ec2messages is now optional; in Regions launched in 2024 and later, its endpoints are not supported at all. Creating it does no harm where it exists, but leaving out ssmmessages breaks registration everywhere.

In an IPv6-only environment, the endpoint names change to the dual-stack forms ssm.region.api.aws, ssmmessages.region.api.aws and ec2messages.region.api.aws, which your outbound rules must allow.

The S3 endpoint policy, if you restrict it, must still allow the AWS-managed buckets the agent uses. A tight S3 endpoint policy that silently blocks those is a classic reason the agent registers but then fails to update. Our S3 VPC endpoint guide covers endpoint policies.

Cause 4: the instance cannot read its own metadata

The agent finds out which instance it is on, and gets the credentials from its instance profile, through the Instance Metadata Service at 169.254.169.254. If IMDS is turned off for the instance, or blocked by a local firewall rule, the agent cannot identify itself, and registration fails before the network is even tried. ssm-cli reports this as a failed "EC2 IMDS" check.

Check the instance's metadata options:

aws ec2 describe-instances --instance-ids i-0123456789abcdef0 \
  --query "Reservations[].Instances[].MetadataOptions"

HttpEndpoint must be enabled. If you use Default Host Management Configuration, HttpTokens must be required, which means IMDSv2, because that feature does not work with IMDSv1. Enabling IMDSv2-only is good security practice anyway, as long as everything else on the instance supports it.

Less common causes: proxies, logs and "too many open files"

A proxy in the way. If outbound traffic must go through an HTTP proxy, the agent has to be told about it, and ssm-cli includes a "Proxy configuration" check. Endpoint traffic to VPC endpoints usually should bypass the proxy, so a no-proxy setting for the metadata address and the endpoint names is part of the setup.

Read the agent's own logs. They usually name the failure in plain words:

  • Linux and macOS: /var/log/amazon/ssm/amazon-ssm-agent.log and /var/log/amazon/ssm/errors.log, plus /var/log/messages on some systems.
  • Windows: %PROGRAMDATA%\Amazon\SSM\Logs\amazon-ssm-agent.log and errors.log; turn on showing hidden files to see them in File Explorer.

"too many open files" on Linux. An agent log line like filewatcher listener encountered error when start watcher: too many open files means the per-user limit on inotify instances is exhausted. AWS's recommended setting is 8,192:

sudo sysctl fs.inotify.max_user_instances
sudo sysctl -w fs.inotify.max_user_instances=8192
printf '%s\n' 'fs.inotify.max_user_instances=8192' | sudo tee /etc/sysctl.d/99-amazon-ssm-agent.conf
sudo sysctl -p /etc/sysctl.d/99-amazon-ssm-agent.conf

Raising the general open-file limits does not fix this particular message; it is the inotify instance limit that matters.

Agent log errors, decoded

The agent's log is blunt once you know its vocabulary. These are the lines people paste into search engines most, and what each one is telling you:

Log line (shortened)What it meansFix
EC2RoleProvider Failed to connect to Systems Manager with SSM role credentials ... RequestManagedInstanceRoleToken: AccessDeniedException: Systems Manager's instance management role is not configured for accountThe agent found no instance profile, tried Default Host Management Configuration, and that is not turned on in this Region.Attach a role with AmazonSSMManagedInstanceCore, or turn on Default Host Management Configuration.
[CredentialRefresher] Retrieve credentials produced error: no valid credentials could be retrieved for ec2 identityThe same problem from another angle: no source of credentials worked.Same as above, and check IMDS is enabled.
Errors naming ssm:UpdateInstanceInformationThe agent has credentials, but the role lacks the permission that registers the instance.Add AmazonSSMManagedInstanceCore to the role.
Connection timeouts to ssm.region.amazonaws.com or ssmmessagesThe network path is blocked or DNS resolves the wrong address.The network section above.
Agent failed to assume any identityThe agent checked each way it can identify itself (an EC2 instance, an ECS task, an on-premises registration) and none worked. On EC2 this almost always means it cannot read instance metadata.Check IMDS is enabled and not blocked by a local firewall.
SSM Agent unable to acquire credentials: <error>...</error>The agent knows which instance it is on but cannot get credentials. The error inside the tags names the reason.The IAM section above; read the inner error.
Agent is in hibernate mode. Reducing logging. (in hibernate.log)Repeated failures pushed the agent into hibernation; it now retries rarely and logs once per back-off period.Fix the cause, then restart the agent.
filewatcher listener encountered error when start watcher: too many open filesThe per-user inotify instance limit is exhausted.Raise fs.inotify.max_user_instances to 8192.

Read the reason without logging in. The agent's source code shows that the first time it fails to get credentials, it writes the "SSM Agent unable to acquire credentials" line, with the underlying error, to the instance's serial port as well as its log file. On EC2, serial port output is what the console shows under Actions, Monitor and troubleshoot, Get system log. So for an instance you cannot reach at all, the cause is often one click away:

aws ec2 get-console-output --instance-id i-0123456789abcdef0 --latest --output text | grep -i "ssm agent"

The --latest option returns the most recent output on instances built on the Nitro System; without it you get a buffered copy that can lag behind.

What if you do not use Systems Manager at all? Amazon Linux and Ubuntu AMIs start the agent anyway, so an instance with no role writes these errors to its log every few minutes. One forum poster had a syslog alert firing on every ERROR line and was understandably tired of it. There are two clean answers: give the instance the AmazonSSMManagedInstanceCore permission so it registers properly, or stop the agent if you have no use for it:

sudo systemctl stop amazon-ssm-agent
sudo systemctl disable amazon-ssm-agent

Think twice before the second option. A registered agent is what lets you reach the server the day SSH breaks, and Session Manager and Run Command cost nothing extra on EC2.

The forum case, solved: RHEL 9 next to a working Amazon Linux instance

Search this error online and the most-read forum question with this exact title, viewed more than 125,000 times, describes the pattern perfectly: a RHEL 9 instance and an Amazon Linux instance in the same VPC, subnet and security group, with SSM endpoints in place and security groups open. Amazon Linux registers; RHEL stays offline. Most replies point to VPC endpoints and security groups, which the poster had already checked.

The difference between those two instances is not the network. Amazon Linux AMIs include the SSM Agent; AWS's RHEL 8 and 9 AMIs do not. When two instances share every network setting and only one registers, check the agent on the other one first:

  1. Connect to the RHEL instance by SSH, EC2 Instance Connect or the serial console.
  2. Run sudo systemctl status amazon-ssm-agent. "Unit not found" means it is not installed.
  3. Install it with the dnf command above for your architecture.
  4. Confirm the role on the instance includes AmazonSSMManagedInstanceCore.
  5. Wait a few minutes, then check Fleet Manager or the CLI for an Online ping status.

That was Ethan's fix for the accountant's server: one dnf command, and it appeared in Fleet Manager a few minutes later.

Locked out? Three ways into an instance while Session Manager is down

There is an irony in this error: the tool you would normally use to fix a server is the one that is not working, and many teams adopted Systems Manager precisely so they could close port 22. When you need a shell to install or restart the agent, these are the doors AWS leaves open.

1. EC2 Instance Connect Endpoint (best for private subnets). This creates a private tunnel from your computer into the VPC, so you can SSH to an instance with no public IP address, no bastion host and no NAT gateway. AWS charges nothing extra for the endpoint itself; you pay only for data transfer if the instance is in a different Availability Zone. You can create one endpoint per VPC and per subnet. Two security group rules make it work:

  • The endpoint's security group allows outbound TCP 22 to the instance's security group (or the VPC's address range).
  • The instance's security group allows inbound TCP 22 from the endpoint's security group.

Then connect with the key pair you launched the instance with. On AWS's RHEL and Amazon Linux AMIs the user name is ec2-user; Ubuntu uses ubuntu:

ssh -i my-key-pair.pem ec2-user@i-0123456789abcdef0 \
  -o ProxyCommand='aws ec2-instance-connect open-tunnel --instance-id i-0123456789abcdef0'

Connecting from the EC2 console's browser terminal instead, with keys handled for you, needs the EC2 Instance Connect package installed on the instance, which is not on every AMI. Your own key pair through open-tunnel works without it.

2. Your existing SSH path. A bastion host, a VPN or a direct SSH rule still works if you kept one. This is how Ethan reached the accountant's server: an SSH hop he had not yet retired.

3. The EC2 Serial Console (last resort). It reaches the instance even with networking completely broken. It is turned off by default and must be allowed at the account level, separately in each Region, and it works on instances built on the Nitro System. The catch for this particular problem: to log in to a Linux instance through it, the operating system needs a user with a password, and a freshly launched AMI has none. It is a rescue tool for servers you prepared in advance, not a quick door into a brand-new one.

If none of these is available and the instance holds nothing you need yet, the honest fastest fix is often to terminate it and launch again with the agent in its user data, as shown earlier.

Confirm the instance is online, and connect to it

Once the cause is fixed, check the ping status, either in Fleet Manager or from the CLI:

aws ssm describe-instance-information \
  --filters "Key=InstanceIds,Values=i-0123456789abcdef0" \
  --query "InstanceInformationList[].[InstanceId,PingStatus,AgentVersion,PlatformName]" \
  --output table

PingStatus is Online when everything works, ConnectionLost when check-ins stopped, and the instance is absent when it never registered.

Fixed it, but still offline? Restart the agent instead of waiting. An agent that has been failing for a while can drop into what AWS calls hibernation: it keeps retrying, but less and less often. AWS documents the retry interval starting at five minutes and stretching to once an hour by default, configurable up to once every 24 hours. So an instance whose role you fixed this morning may not try again until well after lunch. A restart makes it try at once:

# Linux
sudo systemctl restart amazon-ssm-agent

# Ubuntu (snap)
sudo snap restart amazon-ssm-agent

# Windows (PowerShell, as administrator)
Restart-Service AmazonSSMAgent

If you cannot reach the instance to restart the agent, rebooting the instance does the same job. AWS's own Session Manager guidance makes a related point: if the agent was already running when you attached the instance profile, you may need to restart it before the instance appears.

Then open a shell with Session Manager. In the console, choose the instance and Connect, Session Manager. From your own computer, install the Session Manager plugin for the AWS CLI, then:

aws ssm start-session --target i-0123456789abcdef0

Three more things the same connection can do from the AWS CLI, once the instance is online:

# Port forwarding: reach the instance's port 80 at http://localhost:9999
aws ssm start-session --target i-0123456789abcdef0 \
  --document-name AWS-StartPortForwardingSession \
  --parameters '{"portNumber":["80"],"localPortNumber":["9999"]}'

# Run a command without opening a shell
aws ssm send-command --document-name AWS-RunShellScript \
  --targets Key=InstanceIds,Values=i-0123456789abcdef0 \
  --parameters 'commands=["uptime"]'

# Use your normal ssh and scp through Session Manager (in ~/.ssh/config)
Host i-* mi-*
    ProxyCommand sh -c "aws ssm start-session --target %h --document-name AWS-StartSSHSession --parameters 'portNumber=%p'"

The SSH route still needs a key the instance accepts, but port 22 never has to be open to the internet.

A plain Session Manager shell needs no inbound security group rules and no SSH keys at all, which is the reason many teams adopt Systems Manager in the first place. Our Systems Manager guide covers what else becomes available once your instances are managed.

Errors that look like "not online", and how they differ

Several Session Manager errors sound similar but point somewhere else. If your message is one of these, the fix may be quicker than everything above.

"An error occurred (TargetNotConnected) when calling the StartSession operation"

The CLI's version of "not online" usually reads i-0123456789abcdef0 is not connected. It has two causes. The first is everything on this page: the instance is not fully set up for Session Manager. The second catches experienced people: the instance is in a different AWS account or Region from the one your CLI is using. If your CLI profile defaults to us-east-1 and the instance lives in eu-west-1, you get this error for a perfectly healthy instance. Add --region and, if needed, --profile:

aws ssm start-session --target i-0123456789abcdef0 --region eu-west-1 --profile client-account

"The instance you selected isn't configured to use Session Manager"

Here the instance is listed in Systems Manager but Session Manager refuses it. AWS gives two reasons: the instance profile does not include the permissions Session Manager needs, or the agent is older than version 2.3.68.0, the minimum Session Manager needs. AmazonSSMManagedInstanceCore covers the permissions; updating the agent covers the rest.

"SessionManagerPlugin is not found" on your own computer

This one is not about the instance at all. To start sessions from the AWS CLI, your computer needs the Session Manager plugin installed alongside the CLI. Install it with the command for your computer, then check with session-manager-plugin --version:

# Amazon Linux, RHEL, Fedora (x86_64; use linux_arm64 on ARM)
sudo dnf install -y https://s3.amazonaws.com/session-manager-downloads/plugin/latest/linux_64bit/session-manager-plugin.rpm

# Ubuntu, Debian (use ubuntu_arm64 on ARM)
curl "https://s3.amazonaws.com/session-manager-downloads/plugin/latest/ubuntu_64bit/session-manager-plugin.deb" -o "session-manager-plugin.deb"
sudo dpkg -i session-manager-plugin.deb

# macOS (Apple silicon; use mac/ instead of mac_arm64/ on Intel)
curl "https://s3.amazonaws.com/session-manager-downloads/plugin/latest/mac_arm64/session-manager-plugin.pkg" -o "session-manager-plugin.pkg"
sudo installer -pkg session-manager-plugin.pkg -target /
sudo ln -s /usr/local/sessionmanagerplugin/bin/session-manager-plugin /usr/local/bin/session-manager-plugin

# Windows: run the installer from
# https://s3.amazonaws.com/session-manager-downloads/plugin/latest/windows/SessionManagerPluginSetup.exe

The plugin also needs AWS CLI 1.16.12 or later. On Windows, if the command still is not found, add C:\Program Files\Amazon\SessionManagerPlugin\bin\ to your PATH and open a new terminal. Some antivirus software can freeze the plugin during port forwarding; excluding the plugin's folder fixes that.

No Connect button on the Session Manager tab in the EC2 console

On a new instance with no role, the EC2 console's Session Manager tab offers to create one through Quick Setup, which builds a host management configuration: an instance profile with the right permissions, attached for you, at no charge. EC2 can take several minutes to notice. AWS's advice is to wait about two minutes, reboot the instance if the button still has not appeared, and make sure only one host management configuration exists in Quick Setup; if there are two, delete the older one.

A blank screen after the session starts

The connection worked, but nothing appears. AWS lists four causes:

  • The instance's root volume is full, so the agent cannot work.
  • The console link mixes two Regions, such as a us-west-2 console address with region=us-west-1 in the URL.
  • Session logging is turned on, the instance uses VPC endpoints, and there is no s3 or logs endpoint for the session output to reach.
  • The S3 bucket or log group named in your session preferences has been deleted.

"document worker timed out" when a session starts

The full message ends with check [ssm-document-worker]/[ssm-session-worker] log for crash reason, and the agent's debug log usually shows failed to create channel: too many open files. Too many session worker processes have exhausted the inotify limit. Raise fs.inotify.max_user_instances to 8192 as shown above, or close sessions you no longer need: aws ssm describe-sessions --state Active lists them, and aws ssm terminate-session --session-id ends one.

For IT admins: keep every new instance managed

Fixing one instance is easy. Keeping a fleet managed takes a few defaults:

  • Turn on Default Host Management Configuration in every Region you use, so new instances with IMDSv2 and a recent agent are managed without anyone remembering an instance profile.
  • Make IMDSv2 the default for new instances, which Default Host Management Configuration requires and security teams prefer.
  • Bake the agent into your golden AMIs for operating systems that do not include it, such as RHEL, or install it in launch templates' user data.
  • Create the Systems Manager VPC endpoints once per VPC, with private DNS on and a shared endpoint security group allowing 443 from your VPC's address ranges.
  • Keep agents updated automatically, so newer features, and the fixes that come with them, reach every instance.
  • Watch for ConnectionLost in Fleet Manager, or wire the troubleshooting runbook to run when new instances launch, so a missing agent is found the day an instance is created, not the day someone needs a shell.

For automatic agent updates, AWS's own walkthrough schedules the AWS-UpdateSSMAgent document as a State Manager association. This one updates every instance tagged ssm-update=weekly at 2:00 a.m. UTC each Sunday:

aws ssm create-association \
  --name AWS-UpdateSSMAgent \
  --targets Key=tag:ssm-update,Values=weekly \
  --schedule-expression "cron(0 2 ? * SUN *)"

The update downloads from S3, which is one more reason to give private subnets an S3 gateway endpoint.

Systems Manager also manages on-premises servers and machines in other clouds through hybrid activations, with a service role instead of an instance profile. The same four requirements apply: agent, permissions, network and identity. One difference worth knowing: a hybrid node that is deregistered, or whose hardware changes enough to alter its fingerprint, can send its agent into hibernation, so a deregistered on-premises server needs registering again, not just a restart.

The "SSM Agent is not online" checklist

  1. Run AWSSupport-TroubleshootManagedInstance, or ssm-cli get-diagnostics if you can log in.
  2. Agent installed and running? Remember RHEL 8 and 9 AMIs do not include it.
  3. Instance profile with AmazonSSMManagedInstanceCore, or Default Host Management Configuration with IMDSv2 and agent 3.2.582.0 or later.
  4. Outbound 443 to ssm and ssmmessages through an internet gateway, a NAT gateway or VPC endpoints.
  5. VPC endpoints: private DNS on, endpoint security group allowing 443 inbound, a subnet in each Availability Zone.
  6. VPC enableDnsSupport and enableDnsHostnames on; custom DNS forwards amazonaws.com.
  7. IMDS enabled for the instance.
  8. Agent logs read for the exact error.
  9. Ping status Online in describe-instance-information.

Frequently asked questions

What does "SSM Agent is not online" mean?

It means the SSM Agent on the instance has not registered with AWS Systems Manager, or has stopped checking in. The cause is one of four: the agent is missing or stopped, the instance lacks IAM permissions, it cannot reach the Systems Manager endpoints, or it cannot read its instance metadata.

Why is my EC2 instance not showing in Fleet Manager?

The instance has never registered as a managed node. Check that the SSM Agent is installed and running, the instance has AmazonSSMManagedInstanceCore or Default Host Management Configuration, and it can reach the ssm and ssmmessages endpoints on port 443.

Is SSM Agent installed by default on EC2?

Only on some AMIs, including Amazon Linux 2 and 2023, recent Ubuntu Server, SUSE 15.3 and later, AlmaLinux, recent Windows Server and macOS. AWS's RHEL 8 and 9 AMIs do not include it.

How do I install SSM Agent on RHEL 9?

Run sudo dnf install -y https://s3.amazonaws.com/ec2-downloads-windows/SSMAgent/latest/linux_amd64/amazon-ssm-agent.rpm on x86_64, or the linux_arm64 file on Graviton, then check it with sudo systemctl status amazon-ssm-agent. Python 2 or 3 must be installed.

Which VPC endpoints does Systems Manager need?

com.amazonaws.region.ssm and com.amazonaws.region.ssmmessages are required. ec2messages is optional in Regions launched before 2024 and not supported in newer ones. An S3 endpoint is recommended for agent updates and patching.

Is the ec2messages endpoint still required?

No. Since SSM Agent 3.3.40.0, Systems Manager uses ssmmessages whenever it is available. ec2messages is optional in Regions launched before 2024, and its endpoints are not supported in Regions launched in 2024 or later.

What port does the SSM Agent use?

HTTPS on port 443, outbound only. The agent opens every connection to Systems Manager, so no inbound port needs to be open on the instance.

What IAM policy does SSM need?

Attach the AWS managed policy AmazonSSMManagedInstanceCore to the role in the instance's instance profile. Alternatively, Default Host Management Configuration provides permissions for every IMDSv2 instance in the Region without instance profiles.

What is Default Host Management Configuration?

An account-level Systems Manager setting that makes EC2 instances managed automatically, without instance profiles. Instances need IMDSv2 and SSM Agent 3.2.582.0 or later, it is turned on per Region, and it can take up to 30 minutes to apply.

How do I run ssm-cli get-diagnostics?

On Linux or macOS, run sudo ssm-cli get-diagnostics --output table. On Windows, change to C:\Program Files\Amazon\SSM and run .\ssm-cli.exe get-diagnostics --output table. It needs SSM Agent 3.1.501.0 or later.

What does the AWSSupport-TroubleshootManagedInstance runbook check?

It checks why an EC2 instance is not reporting as managed, reviewing its VPC configuration, security groups, VPC endpoints, network ACLs, route tables and IAM instance profile. It does not check inside the operating system or evaluate IPv6 rules.

Why does Fleet Manager show Connection Lost?

The instance registered before but its agent stopped checking in. Common reasons are a stopped agent, a removed IAM role, a changed security group or route, or deleted VPC endpoints.

Do I need a NAT gateway for Systems Manager?

No, if you create interface VPC endpoints for ssm and ssmmessages, with private DNS on. A NAT gateway is the other way to give private instances outbound access to the endpoints.

Why does my Amazon Linux instance work but my RHEL instance doesn't?

Amazon Linux AMIs include the SSM Agent, and AWS's RHEL 8 and 9 AMIs do not. With identical network settings, the RHEL instance simply has no agent to register. Install it and it comes online.

Where are the SSM Agent logs?

On Linux and macOS, /var/log/amazon/ssm/amazon-ssm-agent.log and errors.log. On Windows, %PROGRAMDATA%\Amazon\SSM\Logs\amazon-ssm-agent.log and errors.log.

How do I check if an instance is online in Systems Manager?

Run aws ssm describe-instance-information with a filter for the instance ID and read PingStatus. Online means it is managed; ConnectionLost means check-ins stopped.

Does Systems Manager need IMDS?

Yes. The agent reads the instance metadata service to learn the instance identity and get credentials. If IMDS is disabled the agent cannot register, and Default Host Management Configuration requires IMDSv2.

How do I check the SSM Agent version?

In Fleet Manager, read the SSM Agent version column. On the instance, run dnf info amazon-ssm-agent or yum info amazon-ssm-agent on Amazon Linux and RHEL, sudo snap list amazon-ssm-agent on Ubuntu, or amazon-ssm-agent.exe -version in C:\Program Files\Amazon\SSM on Windows.

How do I install SSM Agent on Windows?

Download AmazonSSMAgentSetup.exe from https://s3.amazonaws.com/ec2-downloads-windows/SSMAgent/latest/windows_amd64/, run it with /S for a silent install, then run Restart-Service AmazonSSMAgent in PowerShell.

What does SSM stand for in AWS?

Simple Systems Manager. The service was originally Amazon Simple Systems Manager, then Amazon EC2 Systems Manager, and is now AWS Systems Manager, but the SSM abbreviation stayed in the agent and CLI names.

What does "Agent failed to assume any identity" mean?

The SSM Agent could not identify itself as an EC2 instance, an ECS task or a registered on-premises machine. On EC2 the usual cause is that instance metadata is disabled or blocked.

How can I see why SSM Agent failed without logging in?

Open the instance's system log in the EC2 console, or run aws ec2 get-console-output. The agent writes "SSM Agent unable to acquire credentials" with the underlying error to the serial port the first time credentials fail.

Is the SSM Agent free?

Yes. The agent costs nothing, and Session Manager and Run Command have no extra charge for EC2 instances. Some other Systems Manager features, and advanced on-premises nodes, are billed.

What does TargetNotConnected mean in Session Manager?

The instance is not ready for Session Manager, or it is in a different AWS account or Region from the one your CLI is using. Check the agent, role and network, and add --region and --profile to aws ssm start-session.

What does "instance management role is not configured for account" mean?

The agent found no instance profile and fell back to Default Host Management Configuration, which is not turned on in that Region. Attach a role with AmazonSSMManagedInstanceCore, or turn on Default Host Management Configuration in Fleet Manager.

Why is my instance still offline after I attached the IAM role?

The agent may be in hibernation, retrying only about once an hour after repeated failures, or it started before the role was attached. Restart the agent or reboot the instance, then check the ping status again.

How do I connect to a private EC2 instance when Session Manager doesn't work?

Use an EC2 Instance Connect Endpoint with your key pair, which needs no public IP or bastion and has no extra charge, or an existing SSH path. The EC2 Serial Console also works if it is enabled and the instance has a password-based user.

Can I stop the SSM Agent if I don't use Systems Manager?

Yes. sudo systemctl stop amazon-ssm-agent and sudo systemctl disable amazon-ssm-agent stop the log errors on an instance with no role. Keeping it registered is free for Session Manager and Run Command on EC2, and gives you a way in if SSH breaks.

This error reads like a networking riddle, and sometimes it is one. More often it is one of four plain things, and the runbook or ssm-cli tells you which in a couple of minutes. Ethan's accountant client never knew there had been a problem; the server was in Fleet Manager before lunch, and Ethan added the agent to his Red Hat launch template that afternoon so it would not happen twice. If you are staring at "not online" right now, start with the question nobody asks first: is the agent actually there?

📌 If you keep one line from this page

"Not online" needs four things: the agent, the permission, a path to ssm and ssmmessages, and instance metadata. Check the agent first; RHEL does not ship with it.

Then let the AWSSupport-TroubleshootManagedInstance runbook check the network for you.

Revision note. Written October 7, 2026, with the endpoint rules for Regions launched since 2024 included. If you spent an evening on security groups for this, you are in very large company; the fix is usually smaller than the search.

Related