Recovery and Repair

This page covers what to do when an endpoint or a self-hosted server gets into a bad state: a system restore, a reimage, a damaged local database, a lost agent identity, a stalled update, or a long stretch with no connection to the cloud. Each tab describes what ET Ducky does on its own and what is left for you.

For diagnosing a problem on a managed machine rather than a problem with ET Ducky itself, see Troubleshooting Workflows.

An endpoint was restored to an earlier point

What happens next depends on one thing: whether the restore preserved the agent's configuration file at C:\ProgramData\ETDucky\Agent\AgentConfig.json, or on Linux /var/lib/etducky.

If the file survived, the machine keeps its identity and its credential. It reconnects on its next heartbeat and appears online again with no action from you. A restore that rolls the agent binary back to an older version is also self-correcting, because the agent checks for a newer release every six hours and installs it silently.

If the file did not survive, the machine has lost its identity. It cannot re-register on its own, because registration requires the one-time token the installer used and the running service does not have it. The agent detects this state and asks to be relinked to its existing record, which an administrator approves. See the Identity and Credentials tab for what that looks like and what approval changes.

Local telemetry buffered before the restore point is gone. Anything the agent had already sent is on the cloud side and is unaffected.

Cloned machines and golden images

An image captured after the agent was installed carries that agent's identity and credential inside it. Every machine deployed from that image presents the same identity, so they overwrite each other's status and only one appears in your fleet.

Build images with the agent absent, and install it as part of first boot or your deployment automation. If an image has already been deployed with an agent baked in, uninstall and reinstall the agent on each affected machine so every one enrols on its own.

ET Ducky notices a related case on its own. When a machine asks to be relinked to a record that is still sending heartbeats, the approval screen says so, because a live record means the original machine is running somewhere and approving would move its identity onto the copy.

An agent shows offline while the machine is clearly running

An agent is marked offline when no heartbeat has arrived for ten minutes. The record is never deleted automatically, so a permanently offline entry stays until you remove it.

Check the agent log at C:\ProgramData\ETDucky\Agent\logs first. A machine that is running but reporting nothing is usually one of three things: the service is stopped, the endpoint cannot reach the cloud, or the agent lost its identity and is waiting for re-enrollment approval. The last of those writes a warning to the log saying exactly that, and the endpoint appears in the queue described under Identity and Credentials.

If you removed an agent record while the machine was still running, approving its re-enrollment request restores the record rather than creating a second one.

Secure desktop settings after a restore

Two remote-control capabilities are off on every agent until an administrator turns them on: Allow Ctrl+Alt+Del, and Allow viewing and control of the secure desktop. Both live in the agent's configuration file, so a restore that rolled that file back rolls them back with it, and a restore that lost the file leaves them off. Either way the agent ends up in the safe state rather than an unexpectedly permissive one.

The Ctrl+Alt+Del capability is the one to check after a restore, because it is the only agent setting that also writes somewhere the configuration file does not reach. Windows will not let software raise the secure attention sequence unless a machine policy permits it, so while the setting is on the agent holds SoftwareSASGeneration set under HKLM\SOFTWARE\Microsoft\Windows\CurrentVersion\Policies\System. That value lives in the registry, and the setting that justifies it lives in a file, so a restore can easily move one without the other.

The agent reconciles the two every time the service starts. If the setting is on and the policy is missing or holds a value that does not work for a service, it is set. If the setting is off and the policy is present, it is reverted. There is nothing to do by hand in the ordinary case; restarting the service is enough to bring them back into step.

One case deserves a look. When the agent first sets the policy it records what was there beforehand in sas-policy-previous.txt, beside AgentConfig.json, and turning the setting off puts exactly that value back. If a restore took that file but left the registry value behind, the agent removes the policy rather than inventing a value to restore. That is the right default on a host that never had the policy set, which is the overwhelming majority. If this host deliberately had a non-default SoftwareSASGeneration before ET Ducky was installed, note it before you restore, because the record of it is what was lost.

The agent log names both events explicitly, with the old and new values, so the sequence is auditable after the fact rather than something to infer.

Golden images carry all of this. An image captured while these capabilities were on produces machines that have them on, with the registry policy already set and a recorded previous value that came from the machine the image was built on rather than from theirs. Building images with the agent absent, which is the guidance under Clones and images for identity reasons, avoids this as well.

The agent buffers telemetry it has not sent

Each agent keeps a local database at C:\ProgramData\ETDucky\Agent\events.db holding what it has collected but not yet delivered. That is what lets an endpoint keep monitoring through a network outage and send the backlog on reconnect.

Records the cloud has accepted are removed on a retention schedule, twenty four hours by default. Records that have never been delivered cannot age out that way, so they are bounded by a limit instead. When an agent goes over that limit it discards the oldest undelivered records first, raw event batches before correlated summaries, because a summary carries the distilled signal from the batches it was built from.

Discarded records are counted, so a gap is recorded rather than passing unnoticed, and the agent writes a warning naming how many were dropped.

A damaged local database

If the local database is unreadable, usually after an unexpected power loss or a disk fault, the agent moves it aside to events.db.corrupt- followed by a timestamp, together with its write-ahead files, creates an empty database, and carries on monitoring.

The damaged file is kept rather than deleted, because it is the only record of what was lost and support may want it. Telemetry it held had not been delivered and is not recoverable. Nothing else about the agent is affected: its identity, credential and configuration live in a separate file.

The agent only does this for genuine corruption. A file that is merely locked, or a disk with no free space, is a temporary condition and is retried rather than replaced.

Choosing a limit for your organization

The limit is set per organization under Configuration, on the Workspace tab, and applies to every agent in the organization. Changes reach agents on their next heartbeat. Only organization administrators can change it.

SettingValue
Default10,000 undelivered records
Minimum1,000
Maximum200,000

The trade is evidence against disk. A larger limit rides out a longer outage with no gap in what you can review afterwards, and uses more space on every endpoint. A smaller limit bounds what the agent can consume and shortens how long an outage can run before the oldest records start to go.

Raise it if your endpoints regularly lose connectivity for long stretches, or if you are investigating an incident and want to be certain nothing is dropped. The default suits a fleet with reliable connectivity.

When an endpoint loses its identity

An agent's identity and credential live in its configuration file. If that file is deleted or damaged, by a restore, a reimage, a disk fault, or a cleanup tool, the agent has nothing to authenticate with and cannot register again on its own, because registration needs the one-time token the installer used.

A damaged file is moved aside to AgentConfig.json.corrupt- followed by a timestamp rather than overwritten, and the agent recovers the previous identifier from it where the text is still readable, so the endpoint asks under the identity your organization already knows.

The agent then submits a re-enrollment request and waits. It keeps monitoring locally while it waits, buffering to the local database. It does not send anything until an administrator approves, and it sends no telemetry as part of the request itself.

Approving a re-enrollment request

Requests appear under Configuration, on the Agents tab, in a panel that is only shown when something is waiting. Each entry names the host, the agent record it is asking to be relinked to, the address it is asking from, and how long it has been asking.

ET Ducky matches the request to an existing record using a hardware fingerprint derived from the machine identifier, network adapter addresses and disk serial. That identifies a machine. It does not prove that a machine is the one it claims to be, since all three are values a host reports about itself, which is why approval is yours rather than automatic.

Before approving, confirm you recognise the host name and that something happened to that machine recently. If the entry warns that the record it is claiming is still online, stop. A record that is still receiving heartbeats means the original machine is running, and approving would take its identity away from a working endpoint. That is what a clone or a restored duplicate looks like.

Denying leaves the endpoint where it was. If you deny in error, reinstalling the agent on that machine enrols it fresh.

What approval changes

The endpoint collects a new credential on its next attempt, within a few minutes, and restarts its service to resume normal operation. The existing agent record is reused rather than replaced, so tags, history and anything else attached to it are kept. A record that had been removed is restored.

No credential is created at the moment you approve. One is issued only when the endpoint collects it, and only to a requester that can prove it is the same one that asked. An approval that is never collected expires after twenty four hours, and a request nobody acts on expires after seven days. Either can simply be made again.

Telemetry the endpoint buffered while it was waiting is sent once it is back.

The remote desktop helper updates separately

Remote desktop sessions are served by a helper program that is versioned and delivered on its own, not by the agent installer. It is cached under C:\Program Files\ETDucky\Agent\RdpHelper in a folder named for its version, which is required for it to capture elevation prompts correctly.

The agent checks for a newer helper every fifteen minutes until it has one cached, then every six hours. Upgrading the agent does not upgrade the helper, so immediately after an agent release the two can be at different versions for a few hours. Restarting the agent service triggers a check within thirty seconds if you need it sooner.

A downloaded helper is verified by checksum and by its signature before it is used. A file that fails either check is discarded and the previous helper keeps running, so a failed update leaves remote desktop working rather than broken.

An agent that is not updating

Agents check for a new release every six hours and install it silently. The installer stops the service, replaces the files and starts it again, preserving configuration, so the agent going away mid-update is expected.

If an agent stays on an old version, check that it is online at all, since an agent that cannot reach the cloud cannot check for updates. Next check the agent log for a download that failed verification. A release whose checksum does not match is refused rather than installed, which is deliberate, and the agent retries on its next cycle.

Self-hosted deployments serve their own installers. If your fleet is not updating there, see the Self-Hosted Server tab.

Endpoints that cannot reach the cloud

An agent that loses its connection keeps collecting and stores what it gathers locally. It retries on its normal interval and reconnects on its own when the network returns. Its live command channel backs off progressively when it cannot connect, up to thirty seconds between attempts, and falls back to periodic polling so commands still arrive.

Nothing needs restarting. If an endpoint stays offline once the network is healthy, confirm the ETDucky agent service is running and that outbound access to your ET Ducky host on port 443 is permitted, then check the agent log.

What survives an outage

Collected telemetry is buffered locally and delivered when the connection returns, within the limit described under Disk and Local Data. Once the backlog exceeds that limit the oldest undelivered records are discarded, so a very long outage leaves a gap at its beginning rather than at its end.

Heartbeats are not queued. An agent that was unreachable shows as offline for that period and comes back online on its first successful heartbeat, rather than backfilling the time it was away.

If you expect long outages, raise the buffer limit before they happen rather than after.

Backing up and restoring a Local-First instance

A compressed database dump is written to ./backups/ before every migration, so an upgrade always leaves a recovery point behind it. You can take one at any time with docker compose exec db pg_dump -U etducky etducky | gzip > manual-backup.sql.gz.

Restoring is destructive by design: it stops the stack, removes the database volume, brings the database back up alone, loads the dump, then starts the rest.

Back up the storage volume and the licence directory as well, not only the database. The storage volume holds the key that protects your stored secrets, including AI provider keys and integration tokens. A database restored without it comes up and runs, but those secrets cannot be read and have to be entered again.

Rolling back a version

Scheduled self-updates back up, pull, verify the image signature, migrate, then restart. If the new version fails its health check the previous image is started again automatically and the rollback is reported in the dashboard. A version that failed once is never retried unattended; review the logs, then save the update schedule again to re-arm it.

After a rollback the database schema may already have moved forward, since migrations run before the health check. The dump taken immediately before the update is your recovery point if the older version misbehaves against the newer schema.

A migration that fails stops there and the application deliberately does not start. Your data is untouched and the pre-migration dump is in ./backups/.

Pinning the server and the fleet

Pinning a specific image tag in your environment file holds the server where it is, and scheduled updates follow only the rolling stable tag, so a pinned version pauses them. The dashboard says when this is in effect.

Agent versions are pinned separately. Self-hosted agents ask your server for updates rather than ours, so leaving the downloads directory empty answers every check with no update available and holds your fleet at its installed version. Drop new installers in when you are ready and agents pick them up on their next six hour check.

Outbound internet access is required for sign-in regardless of pinning. ET Ducky is not an air-gapped product.