The async fn Future state machine overhead (~1.9us) dominated
MpscTunnelSender::send, while RingSink operations were only ~40ns.
Breakthrough: make send() an async fn that completes synchronously
on the first poll for the direct (ring tunnel) path. Uses
futures::task::noop_waker() to construct a dummy Context, then calls
Sink trait methods (poll_ready, start_send, poll_flush) directly.
RingSink always returns Ready immediately, so the waker is never
invoked and the async fn completes without yielding.
Channel mode (TCP/UDP/WG tunnels) still uses async send_async()
with proper backpressure. Ring tunnels detected via tunnel_info()
type check in PeerConn.
Results (4 threads, 1400B, 15s):
pps: 249K → 474K (+90%)
send_msg_by_ip: 3.53us → 1.67us (-53%)
send_msg_internal: 2.40us → 502ns (-79%)
MpscTunnelSender::send: 1.97us → 144ns (-93%)
All 207 peers:: tests pass. Netns-requiring tests (three_node,
credential) unchanged (require root).