Skip to content

Drv/Ip: send to a disconnected TCP peer raises SIGPIPE and terminates the application on Linux #6056

Description

@bitWarrior
F´ Version v4.3.0 (code unchanged on devel at 172aeb2)
Affected Component Drv/Ip (TcpServerSocket, TcpClientSocket, IpSocket); seen through Drv.TcpServer

Problem Description

A deployment using Drv.TcpServer as its GDS link can be terminated by SIGPIPE when the GDS disconnects. The process dies from the signal: no F´ event, no assert, and no FatalHandler.

TcpServerSocket::sendProtocol (Drv/Ip/TcpServerSocket.cpp:136) and TcpClientSocket::sendProtocol (Drv/Ip/TcpClientSocket.cpp:99) call ::send(fd, data, size, SOCKET_IP_SEND_FLAGS), and the default config sets SOCKET_IP_SEND_FLAGS = 0 (default/config/IpCfg.hpp:26). On Linux, a send() on a stream socket whose peer has closed raises SIGPIPE unless the call passes MSG_NOSIGNAL, and SIGPIPE's default action terminates the process. Nothing in the framework ignores or handles SIGPIPE (git grep SIGPIPE finds nothing).

If a telemetry send lands after the peer closes but before the read task has seen EOF and closed the descriptor, the application dies. In our testing this happens most often when a client connects and disconnects quickly. After a longer-lived connection the read task usually closes the socket first, but the race is still there.

A related detail: if SIGPIPE is ignored, the failing send() returns EPIPE. IpSocket::send (Drv/Ip/IpSocket.cpp:166) only treats EBADF/ECONNRESET as SOCK_DISCONNECTED, so EPIPE becomes SOCK_SEND_ERROR, and the sender does not close the connection. The read task still closes it once it sees EOF.

Context / Environment

Operating System: Linux
CPU Architecture: x86_64
Platform: Linux-6.14.0-37-generic-x86_64-with-glibc2.39
Python version: 3.12.3
CMake version: 3.26.0
Pip version: 26.2.1
Pip packages:
    fprime-tools==4.3.0
    fprime-gds==4.3.0
    fprime-fpp==3.3.0
Project submodules:
    https://github.com/nasa/fprime.git @ v4.3.0

Seen in a Linux deployment (Raspberry Pi Zero 2 target, tested natively on x86_64) whose comDriver is Drv.TcpServer, with telemetry downlinked through ComStub at 1 Hz.

How to Reproduce

  1. Build a native Linux deployment that uses Drv.TcpServer as its com driver and downlinks telemetry (e.g. TlmChan in a 1 Hz rate group). Leave SIGPIPE at its default disposition.
  2. Start it listening, e.g. ./MyDeployment -a 127.0.0.1 -p 50000.
  3. Connect a TCP client and close it immediately, then wait about 1.5 s. Repeat a few times:
    import socket, time
    for i in range(6):
        socket.create_connection(("127.0.0.1", 50000)).close()
        time.sleep(1.5)
  4. The application often exits with signal 13 (SIGPIPE; Python's subprocess reports return code -13), usually within the first few connections. In our runs it died in 4 of 6 trials that used this pattern.

Drv.TcpClient (used by Ref) shares the same send path, so a GDS server that closes the connection should have the same effect. We did not reproduce that case separately.

Expected Behavior

A peer disconnect should make the send fail with a status (SOCK_DISCONNECTED), the connection should be closed, and the application should keep running and accept or reconnect the next GDS connection.

Possible Fixes

  • Pass MSG_NOSIGNAL on platforms that have it, for example by defaulting SOCKET_IP_SEND_FLAGS to MSG_NOSIGNAL on Linux. This only affects the socket calls, not process-wide signal handling. macOS has no MSG_NOSIGNAL; setting SO_NOSIGPIPE on the socket after socket()/accept() is the equivalent there.
  • Treat EPIPE like ECONNRESET in IpSocket::send (and UdpSocket::send), returning SOCK_DISCONNECTED so the sender closes the connection promptly.

Workaround used for now: signal(SIGPIPE, SIG_IGN); in the deployment's Main.cpp. With it, the repeated connect/disconnect test above passes, and without it the test fails.

AI Usage

This issue was investigated and drafted with Claude Code (Anthropic), an AI coding agent, at the request of and on behalf of a project maintainer. The agent found the crash while upgrading a deployment to v4.3.0, reproduced it with a test (which fails without the workaround and passes with it), and traced it to the lines cited above in v4.3.0 and devel. The reproduction results come from actual runs; the possible fixes are suggestions and have not been implemented or tested in F´ itself.

IAMAI

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Type

    No type

    Projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions