Fedora Account System
Red Hat Associate
Red Hat Customer
Description of problem: The audisp-remote failure options (notably stop and halt) do not work now, but did in previous versions. The remote_ending_action option does seem to work when set to remote_ending_action = stop but not when = halt. The network_failure_action does not work when set to network_failure_action = halt or = halt. I do not really care about any options other than halt... Version-Release number of selected component (if applicable): audispd-plugins-1.7.13-1.fc10.x86_64 How reproducible: Always Steps to Reproduce: 1. Configure collector machine to receive on say port 1213. Set tcp_listen_port = 1213 in /etc/audit/auditd.conf. 2. Configure sender to send to this machine with the correct "remote-server=" and "port = 1213" in the /etc/audisp/audisp-remote.conf file as well as "active = yes" in the /etc/audisp/plugins.d/au-remote.conf file. 3. Configure sender to halt on failure - set "network_failure_action = halt" in /etc/audisp/audisp-remote.conf. 4. Restart auditd service on collector. 5. Restart auditd service on sender. 6. Send in some sample events from sender like 'auditctl -m "from sender"'. 7. Verify collector is getting sender audit data (e.g. sender is named "sender1" then "ausearch -i -ts recent -n sender1 -m USER"). 8. Unplug network cable on sender. 9. Send in more events from sender1 machine. 10. Verify sender1 does not halt. Actual results: The sender machine does not halt (or the audisp-remote process stop) as configured. Expected results: Sender should halt after time period. Was using the defaults; previously it was almost immediate when the sender would stop. Additional info: This worked at least a few months back. We didn't regression test every release so I'm not sure when the behavior changed. This is pretty bad because the shutdown on failure behavior is required and this was leveraged to meet overall system requirements.
This happens because the connect in the init_sock routine blocks, and when the cable is pulled it cannot recover quickly but instead waits on the tcp timeout. I patched it locally to replace the connect with a non-blocking select with timeout, but my fix may be too specific to be included. I sent a copy of the patch to Steve for reference.
I made a simpler patch that I think solves the problem. The extra logging of errors and extraneous returned data from the server I consider a separate issue. The patch is in upstream commit 316: https://fedorahosted.org/audit/changeset/316
Audit-2.0.1 was built into rawhide to fix this problem. If there are still problems related to this, feel free to re-open the bug. Thanks for reporting the issue.