|
|
| F´ Version |
v4.3.0 (code unchanged on devel at 172aeb2) |
| Affected Component |
Drv/Ip (TcpServerSocket, TcpClientSocket, IpSocket); seen through Drv.TcpServer |
Problem Description
A deployment using Drv.TcpServer as its GDS link can be terminated by SIGPIPE when the GDS disconnects. The process dies from the signal: no F´ event, no assert, and no FatalHandler.
TcpServerSocket::sendProtocol (Drv/Ip/TcpServerSocket.cpp:136) and TcpClientSocket::sendProtocol (Drv/Ip/TcpClientSocket.cpp:99) call ::send(fd, data, size, SOCKET_IP_SEND_FLAGS), and the default config sets SOCKET_IP_SEND_FLAGS = 0 (default/config/IpCfg.hpp:26). On Linux, a send() on a stream socket whose peer has closed raises SIGPIPE unless the call passes MSG_NOSIGNAL, and SIGPIPE's default action terminates the process. Nothing in the framework ignores or handles SIGPIPE (git grep SIGPIPE finds nothing).
If a telemetry send lands after the peer closes but before the read task has seen EOF and closed the descriptor, the application dies. In our testing this happens most often when a client connects and disconnects quickly. After a longer-lived connection the read task usually closes the socket first, but the race is still there.
A related detail: if SIGPIPE is ignored, the failing send() returns EPIPE. IpSocket::send (Drv/Ip/IpSocket.cpp:166) only treats EBADF/ECONNRESET as SOCK_DISCONNECTED, so EPIPE becomes SOCK_SEND_ERROR, and the sender does not close the connection. The read task still closes it once it sees EOF.
Context / Environment
Operating System: Linux
CPU Architecture: x86_64
Platform: Linux-6.14.0-37-generic-x86_64-with-glibc2.39
Python version: 3.12.3
CMake version: 3.26.0
Pip version: 26.2.1
Pip packages:
fprime-tools==4.3.0
fprime-gds==4.3.0
fprime-fpp==3.3.0
Project submodules:
https://github.com/nasa/fprime.git @ v4.3.0
Seen in a Linux deployment (Raspberry Pi Zero 2 target, tested natively on x86_64) whose comDriver is Drv.TcpServer, with telemetry downlinked through ComStub at 1 Hz.
How to Reproduce
- Build a native Linux deployment that uses
Drv.TcpServer as its com driver and downlinks telemetry (e.g. TlmChan in a 1 Hz rate group). Leave SIGPIPE at its default disposition.
- Start it listening, e.g.
./MyDeployment -a 127.0.0.1 -p 50000.
- Connect a TCP client and close it immediately, then wait about 1.5 s. Repeat a few times:
import socket, time
for i in range(6):
socket.create_connection(("127.0.0.1", 50000)).close()
time.sleep(1.5)
- The application often exits with signal 13 (SIGPIPE; Python's
subprocess reports return code -13), usually within the first few connections. In our runs it died in 4 of 6 trials that used this pattern.
Drv.TcpClient (used by Ref) shares the same send path, so a GDS server that closes the connection should have the same effect. We did not reproduce that case separately.
Expected Behavior
A peer disconnect should make the send fail with a status (SOCK_DISCONNECTED), the connection should be closed, and the application should keep running and accept or reconnect the next GDS connection.
Possible Fixes
- Pass
MSG_NOSIGNAL on platforms that have it, for example by defaulting SOCKET_IP_SEND_FLAGS to MSG_NOSIGNAL on Linux. This only affects the socket calls, not process-wide signal handling. macOS has no MSG_NOSIGNAL; setting SO_NOSIGPIPE on the socket after socket()/accept() is the equivalent there.
- Treat
EPIPE like ECONNRESET in IpSocket::send (and UdpSocket::send), returning SOCK_DISCONNECTED so the sender closes the connection promptly.
Workaround used for now: signal(SIGPIPE, SIG_IGN); in the deployment's Main.cpp. With it, the repeated connect/disconnect test above passes, and without it the test fails.
AI Usage
This issue was investigated and drafted with Claude Code (Anthropic), an AI coding agent, at the request of and on behalf of a project maintainer. The agent found the crash while upgrading a deployment to v4.3.0, reproduced it with a test (which fails without the workaround and passes with it), and traced it to the lines cited above in v4.3.0 and devel. The reproduction results come from actual runs; the possible fixes are suggestions and have not been implemented or tested in F´ itself.
IAMAI
develat 172aeb2)Drv/Ip(TcpServerSocket,TcpClientSocket,IpSocket); seen throughDrv.TcpServerProblem Description
A deployment using
Drv.TcpServeras its GDS link can be terminated by SIGPIPE when the GDS disconnects. The process dies from the signal: no F´ event, no assert, and noFatalHandler.TcpServerSocket::sendProtocol(Drv/Ip/TcpServerSocket.cpp:136) andTcpClientSocket::sendProtocol(Drv/Ip/TcpClientSocket.cpp:99) call::send(fd, data, size, SOCKET_IP_SEND_FLAGS), and the default config setsSOCKET_IP_SEND_FLAGS = 0(default/config/IpCfg.hpp:26). On Linux, asend()on a stream socket whose peer has closed raises SIGPIPE unless the call passesMSG_NOSIGNAL, and SIGPIPE's default action terminates the process. Nothing in the framework ignores or handles SIGPIPE (git grep SIGPIPEfinds nothing).If a telemetry send lands after the peer closes but before the read task has seen EOF and closed the descriptor, the application dies. In our testing this happens most often when a client connects and disconnects quickly. After a longer-lived connection the read task usually closes the socket first, but the race is still there.
A related detail: if SIGPIPE is ignored, the failing
send()returnsEPIPE.IpSocket::send(Drv/Ip/IpSocket.cpp:166) only treatsEBADF/ECONNRESETasSOCK_DISCONNECTED, soEPIPEbecomesSOCK_SEND_ERROR, and the sender does not close the connection. The read task still closes it once it sees EOF.Context / Environment
Seen in a Linux deployment (Raspberry Pi Zero 2 target, tested natively on x86_64) whose
comDriverisDrv.TcpServer, with telemetry downlinked throughComStubat 1 Hz.How to Reproduce
Drv.TcpServeras its com driver and downlinks telemetry (e.g.TlmChanin a 1 Hz rate group). Leave SIGPIPE at its default disposition../MyDeployment -a 127.0.0.1 -p 50000.subprocessreports return code-13), usually within the first few connections. In our runs it died in 4 of 6 trials that used this pattern.Drv.TcpClient(used byRef) shares the same send path, so a GDS server that closes the connection should have the same effect. We did not reproduce that case separately.Expected Behavior
A peer disconnect should make the send fail with a status (
SOCK_DISCONNECTED), the connection should be closed, and the application should keep running and accept or reconnect the next GDS connection.Possible Fixes
MSG_NOSIGNALon platforms that have it, for example by defaultingSOCKET_IP_SEND_FLAGStoMSG_NOSIGNALon Linux. This only affects the socket calls, not process-wide signal handling. macOS has noMSG_NOSIGNAL; settingSO_NOSIGPIPEon the socket aftersocket()/accept()is the equivalent there.EPIPElikeECONNRESETinIpSocket::send(andUdpSocket::send), returningSOCK_DISCONNECTEDso the sender closes the connection promptly.Workaround used for now:
signal(SIGPIPE, SIG_IGN);in the deployment'sMain.cpp. With it, the repeated connect/disconnect test above passes, and without it the test fails.AI Usage
This issue was investigated and drafted with Claude Code (Anthropic), an AI coding agent, at the request of and on behalf of a project maintainer. The agent found the crash while upgrading a deployment to v4.3.0, reproduced it with a test (which fails without the workaround and passes with it), and traced it to the lines cited above in v4.3.0 and
devel. The reproduction results come from actual runs; the possible fixes are suggestions and have not been implemented or tested in F´ itself.IAMAI