Capstone: Long Jobs Over SSH

GNU Screen & Linux Job Control

Chapter 10 · Capstone: Running and Rescuing Long Jobs Over SSH

This chapter puts the whole course together in one realistic day of work on a remote server. You'll start a long job safely, lose your connection, find your way back, rescue a job that was started the wrong way, watch everything from outside, and clean up. Each step names the chapter it comes from.

The scenario
You're migrating a database on a server called dbhost. You need to run a long import, keep an eye on disk space, and — as it turns out — rescue a compression job a colleague started straight in their SSH session before they had to leave.

Step 1: Prepare Your Environment (Chapter 9)

A one-off setup that makes every later step easier: a status line, so you always know you're inside screen and which window you're in, and plenty of scrollback.

# ~/.screenrc on dbhost startup_message off defscrollback 10000 hardstatus alwayslastline "%H | %-w[%n %t]%+w"

Step 2: Start the Long Job the Right Way (Chapters 2, 4, 9)

laptop$ ssh dbhost dbhost$ cd ~/migration dbhost$ screen -S migrate -L # named session, with logging # Inside the session, window 0: Ctrl-a A import # name the window $ mysql newdb < dump.sql # A second window to keep an eye on disk space: Ctrl-a c Ctrl-a A disk $ watch df -h Ctrl-a 0 # back to the import Ctrl-a _ # tell me when it goes quiet (finished or hung)

Step 3: Lose the Connection (Chapters 1, 7)

Your laptop goes to sleep and the SSH connection drops. Without screen, the terminal hang-up would send SIGHUP to the import and it would stop. Here, only the screen client is lost: the import runs inside a terminal owned by the screen server, which nothing has hung up.

Step 4: Find Your Way Back (Chapters 2, 3)

laptop$ ssh dbhost dbhost$ screen -ls There is a screen on: 41200.migrate (Attached) # the server hasn't noticed you left 1 Socket in /run/screen/S-philip. dbhost$ screen -r migrate # refused: the session still counts as attached dbhost$ screen -d -r migrate # detach the dead client, reattach here

You're back in window 0. Scroll up with Ctrl-a [ to see anything the import printed while you were away, or read screenlog.0 (Chapter 9).

Step 5: Investigate the Colleague's Job (Chapters 5, 6)

Your colleague messages: they started a large compression job directly in their SSH session, then had to leave. You're both working under the same shared deployment account, so you can see and act on the process. First, find it and check where it lives.

$ pgrep -af xz 41877 xz -9 archive.tar $ pstree -s -p 41877 systemd(1)───sshd(41820)───sshd(41855)───bash(41856)───xz(41877) $ ps -o pid,stat,tty,cmd -p 41877 PID STAT TT CMD 41877 R+ pts/3 xz -9 archive.tar

The ancestry shows no SCREEN process: xz is a direct child of an SSH login's shell, on that login's terminal pts/3, and in its foreground (+). If that SSH connection drops, it dies. Its job number, %1 in your colleague's shell, means nothing in yours — only the PID does.

Step 6: Rescue It Into Screen (Chapters 7, 8)

You can't press Ctrl-Z in someone else's terminal, but you can pull the process into a new window of your own session with reptyr.

# Check ptrace permission first $ cat /proc/sys/kernel/yama/ptrace_scope 1 # restricted: needs sudo # In a new window of the migrate session Ctrl-a c Ctrl-a A compress $ sudo reptyr 41877

xz now runs in the compress window. When your colleague's connection closes, nothing reaches it. A quick pstree -s -p 41877 would now show your session's SCREEN in its ancestry.

If reptyr weren't available
The fallback would be to have your colleague, before leaving, press Ctrl-Z, run bg and disown %1 (Chapter 7). The job would survive, but nobody could watch it again, only check that the process still exists and that its output file keeps growing.

Step 7: Watch Everything From Outside (Chapters 8, 9)

You detach and go to a meeting, but want to check in from your phone's SSH app without disturbing the layout.

# Add one more window without attaching: watch the import's table count $ screen -S migrate -X screen -t tables watch "mysql -e 'SHOW TABLES' newdb | wc -l" # Snapshot what the import window is showing right now $ screen -S migrate -p import -X hardcopy /tmp/import.txt $ tail /tmp/import.txt # See everything running in the session $ pstree -p -a 41200

Step 8: Finish and Clean Up (Chapters 2, 3, 4)

# The silence monitor from Step 2 reports that window 0 has gone quiet $ screen -r migrate # Check the import finished cleanly, then close windows you no longer need Ctrl-a k # in each finished window # When everything's done: end the session and tidy leftovers $ screen -S migrate -X quit $ screen -wipe $ screen -ls No Sockets found in /run/screen/S-philip.

Where Each Step Came From

ChapterUsed in this capstone
1 — Why Screen?Why the import survived the dropped connection
2 — Detach & ReattachNamed session; -d -r for the stale "Attached" session
3 — Listing Sessionsscreen -ls, the session ID, -X quit and -wipe
4 — Windows & RegionsNamed windows and silence monitoring
5 — Job Control BasicsWhy a job number couldn't be used across shells
6 — IDs Comparedpgrep, pstree -s and ps to locate the colleague's job
7 — Surviving LogoutSIGHUP, and the disown fallback
8 — Rescuing Running Jobsreptyr, ptrace_scope, and adding a window with -X screen
9 — Automating Screen.screenrc, -L logging, hardcopy, scrollback

What This Course Doesn't Cover

  • tmux in depth — the same ideas with different keys and a more modern design (Chapter 1).
  • Multi-user sessions between different accounts, which need administrator setup (Chapter 9).
  • Permanent services — anything that should start at boot and restart when it fails belongs in systemd (see the systemd in Depth course).
  • Job schedulers such as cron, at and systemd timers, for jobs that should run at a set time without you.

Hands-On Exercises

Exercise 1

Recreate Steps 2–4 on your own machine, using a harmless stand-in for the import (such as a loop that prints the date every 10 seconds). Simulate the dropped connection by closing the terminal window, and recover the session. Note the exact message screen -ls gives you.

📄 View solution
Exercise 2

Recreate Steps 5–6: in one terminal start yes > /dev/null outside screen; from a second terminal, find it, prove it's not in screen, and move it into a screen window. Then prove it is in screen.

📄 View solution
Exercise 3

Write a short checklist, in your own words, for starting any job on a remote server that might take more than ten minutes. It should cover starting, checking progress, surviving a disconnection, and cleaning up.

📄 View solution

Course Quick Reference

  • Start: screen -S name (add -L to log); screen -dmS name cmd to start detached
  • Leave and return: Ctrl-a d; screen -r name; screen -d -r name if it's still marked Attached
  • List and identify: screen -ls; IDs are PID.name; $STY and $WINDOW inside
  • Windows: Ctrl-a c, n/p, ", A, k; splits S, |, Tab; monitoring M and _
  • Job control: &, Ctrl-Z, jobs -l, fg, bg, kill %n, $!
  • Find processes: pgrep -af, pstree -s -p PID, pstree -p -a SCREEN_PID
  • Survive logout: nohup, bg then disown, setsid; check systemd's KillUserProcesses
  • Rescue into screen: reptyr PID inside a window (check ptrace_scope)
  • Automate: -X screen, -p, -X stuff $'cmd\n', -X hardcopy, ~/.screenrc
  • Clean up: screen -S name -X quit, screen -wipe

Course Complete

This completes GNU Screen & Linux Job Control, 10/10 chapters.