<?xml version="1.0" encoding="UTF-8"?>
<!-- This template is for creating an Internet Draft using xml2rfc,
     which is available here: http://xml.resource.org. -->
<!DOCTYPE rfc SYSTEM "rfc2629.dtd" [
<!-- One method to get references from the online citation libraries.
     There has to be one entity for each item to be referenced. 
     An alternate method (rfc include) is described in the references. -->

<!ENTITY RFC2119 SYSTEM "http://xml.resource.org/public/rfc/bibxml/reference.RFC.2119.xml">
<!ENTITY RFC5666 SYSTEM "http://xml.resource.org/public/rfc/bibxml/reference.RFC.5666.xml">
]>
<?xml-stylesheet type="text/xsl" href="rfc2629.xslt" ?>
<?rfc strict="yes" ?>
<?rfc toc="yes"?>
<?rfc tocdepth="3"?>
<?rfc symrefs="yes"?>
<?rfc sortrefs="yes" ?>
<?rfc compact="yes" ?>
<?rfc subcompact="no" ?>
<rfc ipr="trust200902" 
     category="info"
     docName="draft-dnoveck-nfsv4-rpcrdma-rtissues-00">
  <front>
    <title abbrev="RPC/RDMA Round-trip Issues">
      Issues Related to RPC-over-RDMA Internode Round-trips
    </title>
    <author initials="D." surname="Noveck" fullname="David Noveck">
      <organization abbrev="HPE">
        Hewlett Packard Enterprise
      </organization>
      <address>
        <postal>
          <street>165 Dascomb Road</street> 
          <city>Andover</city>
          <region>MA</region>
          <code>01810</code>
          <country>USA</country>
        </postal>
        <phone>+1 781-572-8038</phone>
        <email>davenoveck@gmail.com</email>
      </address>
    </author>
    <date year="2016"/>

    <area>Transport</area>
    <workgroup>Network File System Version 4</workgroup>
    <abstract>
      <t>
        As currently designed and implemented, the RPC-over-RDMA 
        protocol requires use of multiple internode round trips to 
        process many common operations.  For example,
        NFS READ or WRITE operations require use of three internode
        round trips.  This document looks at this issue and discusses
        what can and what should be done to address it, both within the
        context of an extensible version of RPC-over-RDMA and possibly 
        outside that framework.
      </t>
    </abstract>
  </front>
  <middle>
    <section title="Preliminaries" anchor="PRELIM">	
      <section title="Requirements Language" anchor="INTRO-req">
        <t>
          The key words "MUST", "MUST NOT", "REQUIRED", "SHALL", 
          "SHALL NOT", "SHOULD", "SHOULD NOT", "RECOMMENDED", 
          "MAY", and "OPTIONAL" in this document are to be interpreted 
          as described in <xref target="RFC2119"/>.  
        </t>
      </section>
      <section title="Introduction" anchor="PRELIM-intro">
        <t>
          When many common operations are performed using RPC-over-RDMA,
          additional  inter-node round-trip latencies are required
          to take advantage of the performance benefits provided by
          RDMA Functionality.  
        </t>	
        <t>
          While the latencies involved are generally small, they are a reason
          for concern for two reasons.
        <list style="symbols">
          <t>
            With the ongoing improvement of persistent memory 
            technologies, such internode latencies, being fixed,
            can be expected to consume an increasing portion of the
            total latency required for processing NFS requests using 
            RPC-over-RDMA.
          </t>
          <t>
            High-performance transfers using NFS may be needed outside
            of a machine-room environment.  As RPC-over-RDMA is used in
            networks of campus and metropolitan scale, the internode
            round-trip time of sixteen microseconds per mile becomes 
            an issue.
          </t>
        </list>
        </t>	
        <t>
          Given this background, round trips beyond the minimum necessary
          need to be justified by corresponding benefits.  If they are
          not, work needs to be done to eliminate those excess round trips.
        </t>	
        <t>
          We are going to look at the existing situation with regard
          to round trip latency and make some suggestions as to  
          how the issue might be best addressed. We will consider 
          things that could be done in the near 
          future and also explore further possibilities that would
          require a longer-term approach to be adopted.
        </t>	
      </section>
    </section>
    <section title="Review of the Current Situation" anchor="CUR">
      <section title="Troublesome Requests" anchor="CUR-trouble">
        <t>
          We will be looking at four sorts of situations:
        <list style="symbols">
          <t>
            An RPC operation involving Direct Data Placement of request 
            data (e.g., an NFSv3 WRITE or corresponding NFSv4 COMPOUND).  
          </t>
          <t>
            An RPC operation involving Direct Data Placement of response 
            data (e.g., an NFSv3 READ or corresponding NFSv4 COMPOUND).  
          </t>
          <t>
            An RPC operation where the request data is longer than the 
            inline buffer limit.
          </t>
          <t>
            An RPC operation where the response data is longer than the 
            inline buffer limit.
          </t>
        </list>
        </t>
        <t>
          We will survey the resulting latencies in an RPC-over-RDMA
          Version One environment in <xref target="CUR-details"/>
          below.  
        </t>
      </section>
      <section title="Request Processing Details" anchor="CUR-details">
        <t>
          We'll start with the case of a request involving direct placement 
          of request data.  Processing proceeds as described below.  Although
          we are focused on internode latency, the time to perform
          a request also includes such things as interrupt latency, overhead
          involved in interacting with the RNIC, and the time for the server
          to execute the requested operation.
        <list style="symbols">
          <t>
            First, the memory to be accessed remotely is registered.  
            This is a local operation. 
          </t>
          <t>
            Once the registration has been done, 
            the initial send of the request
            can proceed.  Since this is in the
            context of connected operation, there is an internode 
            round-trip involved.  However, the next step can proceed
            after the initial transmission is received.  As a result, only
            the responder-bound side of the transmission contributes to
            overall operation latency.
          </t>
          <t>
            The responder, after being notified of the receipt of the request,
            uses RDMA READ to fetch the bulk data.  This
            involves an internode round-trip latency.  The responder then
            needs to be notified of the completion of the explicit RDMA
            operation
          </t>
          <t>
            The responder (after doing the actual operation) sends the 
            response.  Again,  as this is in the
            context of connected operation, there is an internode 
            round-trip involved.  However, the next step can proceed
            after the initial transmission is received by the requester.
          </t>
          <t>
            The requester, after being notified of the receipt of the response,
            deregisters the memory originally registered before the request
            was issued.  This is also a local operation.
          </t>
        </list>
        </t>
        <t>
          To summarize, if we exclude the actual server execution of the 
          request,  the latency consists of two round-trip
          internode latencies plus two-responder-side interrupt latencies 
          plus one requester-side interrupt latency plus any necessary 
          registration/de-registration overhead.  This is in contrast
          to a request not using explicit RDMA operations in which 
          there is a single inter-node round-trip latency and one 
          interrupt latency on the requester and the responder. 
        </t>
        <t>
          The processing of the other sorts of requests mentioned in 
          <xref target="CUR-trouble" /> is very similar.
        <list style="symbols">
          <t>
            The case of direct data placement of response data follows the 
            same pattern.  The only difference is that the transfer of the
            bulk data is performed using RDMA WRITE, rather than RDMA READ.
          </t>
          <t>
            Handling of a long request is also similar to the above.  The
            memory associated with a position-zero read chunk is registered,
            transferred using RDMA READ, and deregistered.  As a result we have
            the same overhead and latency issues associated with the case of
            direct data placement, without the corresponding benefits.
          </t>
          <t>
            Handling of a long response is a mirror image in that 
            RDMA WRITE is used, rather than RDMA READ.
          </t>
        </list>
        </t>
      </section>
    </section>
 
    <section title="Near-term Work" anchor="NEAR">
      <t>
        We are going to consider how the latency issues discussed
        in <xref target="CUR"/> might be addressed in the context of
        an extensible version of RPC-over-RDMA, such as that proposed 
        in <xref target="rpcrdmav2"/>.
      </t>
      <t>
        In <xref target="NEAR-target"/>, we will establish a performance
        target for the troublesome requests, based on the performance of
        requests that do not involve long messages or direct data placement. 
      </t>
      <t>
        We will then consider how extensions 
        might be defined to bring latency and overhead for the requests 
        discussed in <xref target="CUR-trouble"/> into line with those
        for other requests.  There will be two specific classes of 
        requests to address:
      <list style="symbols">
        <t>
          Those that do not involve direct data placement will be addressed
          in <xref target="NEAR-cont"/>.  In this case, there are no 
          compensating benefits justifying the higher latency and overhead. 
        </t>
        <t>
          The more complicated case of requests that do involve direct data
          placement is discussed in <xref target="NEAR-sbddp"/>.  In this case,
          direct data placement could serve as a compensating benefit, and the 
          important question to be addressed is whether Direct Data Placement
          can be effected without the additional round-trip latencies.
        </t>
      </list>
      </t>
      <t>
        The optional features to deal with each of the classes of messages
        discussed above could be implemented separately.  However, in 
        the handling of RPCs with very large amounts of bulk data, the
        two features are synergistic.  This fact makes it desirable to 
        define the
        features as part of the same extension.  See <xref target="NEAR-syn"/>
        for details.
      </t>
      <section title="Target Performance" anchor="NEAR-target"> 
        <t>
          As our target, we will look at the latency and overhead 
          associated with other sorts of RPC requests, i.e. those that 
          do not use DDP, and that have request and response messages 
          which do fit within the buffer limit.
        </t>
        <t>
          Processing proceeds as follows: 
        <list style="symbols">
          <t>
            The initial send of the request is done.  Since this is in the
            context of connected operation, there is an internode 
            round-trip involved.  However, the next step can proceed
            after the initial transmission is received.  As a result, only
            the responder-bound side of the transmission contributes to
            overall operation latency.
          </t>
          <t>
            The responder, after being notified of the receipt of the request,
            performs the requested operation and sends the reply.
            As in the case of the request, there is an internode round-trip
            involved. However, the request can be considered complete upon
            receipt of the requester-bound transmission.  The 
            responder-bound acknowledgment does not contribute to request
            latency.
          </t>
        </list>
        </t>
        <t>
          In this case there is only a single internode round-trip latency 
          necessary to effect the RPC.  Total request latency includes 
          this round-trip
          latency plus interrupt latency on the requester and responder, plus
          the time for the responder to actually perform the requested 
          operation.
        </t>
        <t>
          Thus the delta between the operations discussed in 
          <xref target="CUR"/> and our baseline consists of:
        <list style="symbols">
          <t>
            One additional internode round-trip latency.
          </t>
          <t>
            One additional instance of responder-side interrupt latency
          </t>
          <t>
            The additional overhead necessary to do memory registration and
            deregistration.
          </t>
        </list>
        </t> 
          
      </section>
      <section title="Message Continuation" anchor="NEAR-cont">
        <t>
          Using multiple RPC-over-RDMA transmissions, in sequence, to
          send a single RPC message avoids the additional latency 
          associated with the
          use of explicit RDMA operations to transfer position-zero
          read chunks or reply chunks.
        </t>
        <t>
          Although transfer of a single request or reply in N transmissions
          will involve N+1 internode latencies, overall request 
          latency is not increased
          as it currently is, by requiring that operations involving multiple
          nodes be serialized.
        </t>
        <t>
          As an illustration, let's consider the case of a request involving 
          a response consisting of two RPC-over-RDMA transmissions.  Even
          though each of these transmissions is acknowledged, that 
          acknowledgement does not contribute to request latency.  The second
          transmission can be received by the requester and acted upon without
          waiting for either acknowledgment.
        </t>
        <t>
          This situation would require multiple receive-side interrupts but
          it is unlikely to result in extended interrupt latency.  With 1K
          sends (Version One), the second receive will complete about 200
          nanoseconds after the first assuming a 40Gb/s transmission rate.
          Given likely interrupt latencies, the first interrupt routine
          would be able 
          to note that the completion of the second receive had already 
          occurred.
        </t>
      </section>
      <section title="Send-based DDP" anchor="NEAR-sbddp">
        <t>
          In order to effect proper placement of request or reply
          data within the context of individual RPC-over-RDMA transmissions, 
          receive buffers
          must be structured to accommodate this function
        </t>
        <t>
          To illustrate the considerations that lead clients and servers 
          to choose particular buffer structures, we will use as examples,
          the cases of NFS READs and WRITEs of 8K data blocks (or the 
          corresponding NFSv4 COMPOUNDs).
        </t>
        <t>
          In such cases, the client and server need to have the DDP-eligible
          bulk data placed in 8K-aligned 8K buffer segments.  Rather than
          being transferred in separate transmissions using explicit RDMA
          operations, a message can be sent so that bulk data is received
          into an appropriate buffer segment.  In this case, it will be
          excised from the XDR payload stream, just as it is in the case of
          existing DDP facilities.
        </t>
        <t> 
          Consider a server expecting write requests which are mostly 
          X bytes long, exclusive of an 8K bulk data area.   In this case
          the payload stream will be less than X bytes and will fit in 
          buffer segment devoted to that purpose.  The bulk data needs to
          be placed in the subsequent buffer segment in order to align it
          properly, i.e. with 8K alignment in the DDP target buffer.  In
          order to place the data appropriately, the  sender (in the case, the 
          client needs) to 
          add padding of length X-Y bytes where Y is the length of payload
          stream for the current request.  The case of reads is exactly the
          same except that the sender adding the padding is the server.      
        </t>
        <t>                
          To provide send-based DDP as an RPC-over-RDMA extension, the 
          framework defined in <xref target="xcharext" /> could be used.
          A new "transport characteristic" could be defined which  
          allowed a participant to expose the structure of his receive 
          buffers and to identify the buffer segments capable of being
          used as DDP targets.  In addition, a new optional message header
          would have to be defined.  It would be defined to provide:
        <list style="symbols">
          <t>
            A way to designate DDP-eligible data item as 
            corresponding to target buffer segments, rather than memory 
            registered for RDMA.
          </t>
          <t>
            A way to indicate to the responder that it should place
            DDP-eligible data items in DDP-targetable buffer segments, rather 
            than in memory registered for RDMA.
          </t>
          <t>
            A way to designate a limited portion of an RPC-over-RDMA 
            transmission
            as constituting the payload stream.
          </t>
        </list>

        </t>
      </section>
      <section title="Feature Synergy" anchor="NEAR-syn">
        <t>
          While message continuation and send-based DDP each address an
          important class of commonly used messages, their combination
          allows simpler handling of some important classes of messages:
        <list style="symbols">
          <t>
            READs and WRITEs transferring larger IOs
          </t>
          <t>
            COMPOUNDs containing multiple IO operations.
          </t>
          <t>
            Operations whose associated payload stream is longer than
            the typical value. 
          </t>
        </list>
        </t>
        <t>
          To accommodate these situations, it seems that the definition of
          the headers for message continuation need to interact with data
          structures for send-based DDP as follows:
        <list style="symbols">
          <t>
            The header type for the message starting a chained group contains
            DDP-directing structures which support both send-based DDP
            as well as DDP using Explicit RDMA operations.
          </t>
          <t>
            Buffer references for Send-based DDP should be relative to
            the start of the transmission group and should allow transitions
            between buffer segments in different receive buffers. 
          </t>
          <t> 
            The header type for messages within a chained group should not 
            have DDP-related fields but should rely on the initial message
            of the group for DDP-related functions.
          </t>
          <t>
            The portion of each received transmission devoted to the 
            payload stream should be part of the header for each message 
            within a chained group.  The payload stream for the message 
            as a whole should be the concatenation of those for each
            transmission.
          </t>
        </list>
        </t>
      </section>

    </section>
    <section title="Possible Future Development" anchor="FUTURE">
      <t>
          Although the reduction of explicit RDMA operation reduces the number
          of inter-node round trips and eliminates sequences of operations 
          in which multiple round-trip latencies are serialized with server
          interrupt latencies, the use of connected operations means that
          round-trip latencies will always be present, since each
          message is acknowledged.
        </t>
        <t>
          One avenue that has been considered is use of
          unreliable-datagram (UD) transmission in environments where the
          "unreliable" transmission is sufficiently reliable that RPC
          replay can deal with a very low rate of message loss.  
          For example, UD in 
          Infiniband specifies a low enough rate of frame loss to make
          this a viable approach, particularly given NFSv4.1's EOS support.
        </t>
        <t>
          With this sort of arrangement, request latency is still the same.
          However, since the acknowledgements are not serving any 
          substantial function, it is tempting to consider removing them,
          as they do take up some transmission bandwidth, that might be 
          used otherwise, if the protocol were to reach the goal of 
          effectively using the underlying medium.
        </t>
        <t>
          The size of such wasted transmission bandwidth depends on the 
          average messages size and many implementation considerations
          regarding how acknowledgments are done.  In any case, given
          expected message sizes, the wasted transmission bandwidth will
          be very small.
        </t>
        <t>
          When RPC messages are quite small, acknowledgments may be of
          concern.  However, in that situation, a better response would
          be transfer multiple RPC messages within a single RPC-over-RDMA
          transmission. 
        </t>
        <t>
          When multiple RPC messages are combined into a single transmission,
          the overhead of interfacing with the RNIC, particularly the 
          interrupt handling overhead, is amortized over multiple RPC
          messages. 
        </t>
        <t>
          Although this technique is quite outside the spirit of existing
          RPC-over-RDMA implementations, it appears possible to define new
          header types capable of supporting this sort of transmission, 
          using the extension framework described in 
          <xref target="rpcrdmav2" />.
        </t>
      </section>


    <section title="Summary" anchor="CON">
      <t>
        We've examined the issue of round-trip latency and concluded:
      <list style="symbols">
        <t>
          That the number of round trips per se is not as important as the
          contribution of any extra round-trips to overall request latency.
        </t>
        <t>
          That the latency issue can be addressed using the extension
          mechanism provided for in <xref target="rpcrdmav2" />.
        </t>
      </list>
      </t>
      <t>
        As it seems that the features sketched out could put internode 
        latencies for a large class of requests back to the baseline value 
        for the RPC paradigm, more detailed definition of the required
        extension functionality is in order.
      </t>
      <t>
        We've also looked at round-trips at the physical level, in that
        acknowledgments are sent in circumstances where there is no obvious
        need for them.  With regard to these, we have concluded:
      <list style="symbols">
        <t>
          That these acknowledgements do not contribute to request latency.
        </t>
        <t>
          That while UD transmission can remove 
          acknowledgements of limited value, the
          performance benefits are not sufficient to justify the disruption
          that this would entail.
        </t>
        <t>
          That issues with transmission bandwidth overhead in a small-message
          environment are better addressed by combining 
          multiple RPC messages in
          a single RPC-over-RDMA transmission.  
          This is particularly so, because 
          such a step is likely to reduce overhead in such environments
          as well 
        </t>
      </list>
      </t>
      <t>
        As the features described involve the use of alternatives to 
        explicit RDMA
        operations, in performing direct data placement and in transferring
        messages that are larger than the receive buffer limit, it is 
        appropriate to understand the role that such operations
        are expected to have once the extensions discussed in this document are
        fully specified and implemented.
      </t>
      <t>
        It is important to note that these extensions are OPTIONAL and are 
        expected to remain so, while support for explicit RDMA operations
        will remain an integral part of RPC-over-RDMA.
      </t>
      <t>
        Given this framework, the degree to which explicit RDMA operations 
        will be used will reflect future implementation choices and needs.  
        While
        we have been focusing on cases in which other options might be more 
        efficient in some cases, it worth looking also at the cases in 
        which explicit RDMA operations are likely to remain preferable:
      <list style="symbols">
        <t>
          In some environments in which direct data placement to memory of
          a certain alignment does not meet application requirements 
          and in which data needs to
          be read into a particular address on the client.  Also, 
          large physically contiguous
          buffers may be required in some environments. In these situations,
          send-based DDP is not an option. 
        </t>
        <t>
          Where large transfers are to be done, there will be limits to the
          capacity of send-based DDP to provide the required functionality,
          since the basic pattern using send/receive is to allocate 
          a pool of memory to contain 
          receive buffers in advance of issuing requests.
          While this issue can be mitigated by use of message continuation,
          tying up large numbers of credits for a single request can cause
          difficult issues as well.  As a result, send-based DDP may 
          be restricted to 
          "small" IO's although the definition of "small" in this context
          is inevitably 
          somewhat elastic.
        </t>
      </list>
      </t>
    </section>
    <section title="Security Considerations" anchor="SEC">
      <t>
        This document does not raise any security issues.
      </t>
    </section>
    <section title="IANA Considerations" anchor="IANA">
      <t>
        This document does not require any actions by IANA.
      </t>
    </section>
  </middle>
  <back>
    <references title="Normative References">
      &RFC2119;
      <reference anchor="rfc5666bis" 
                 target="http://www.ietf.org/id/draft-ietf-nfsv4-rfc5666bis-05.txt">
        <front>
          <title>
            Remote Direct Memory Access Transport for Remote Procedure Call
          </title>

          <author initials="C." surname="Lever" role="editor">
            <organization>Oracle</organization>
          </author>
          <author initials="W." surname="Simpson">
            <organization>DayDreamer</organization>
          </author>
          <author initials="T." surname="Talpey">
            <organization>Oracle</organization>
          </author>
          <date month="April" year="2016" />
        </front>
        <annotation>
          Work in progress.
        </annotation>

      </reference>
    </references>
    <references title="Informative References">
      &RFC5666;
      <reference anchor="rpcrdmav2" 
                 target="http://www.ietf.org/id/draft-cel-nfsv4-rpcrdma-version-two-00.txt">
        <front>
          <title>
            RPC-over-RDMA Version Two
          </title>

          <author initials="C." surname="Lever" role="editor">
            <organization>Oracle</organization>
          </author>
          <author initials="D." surname="Noveck">
            <organization>Hewlett Packard Enterprise</organization>
          </author>
          <date month="April" year="2016" />
        </front>
        <annotation>
          Work in progress.
        </annotation>
      </reference>
      <reference anchor="xcharext" 
                 target="http://www.ietf.org/id/draft-dnoveck-nfsv4-rpcrdma-xcharext-00.txt">
        <front>
          <title>
            RPC-over-RDMA Extension to Manage Transport Characterisitcs
          </title>

          <author initials="D." surname="Noveck">
            <organization>Hewlett Packard Enterprise</organization>
          </author>
          <date month="April" year="2016" />
        </front>
        <annotation>
          Work in progress.
        </annotation>
      </reference>
    </references>
    <section title="Acknowledgements" anchor="ACK">
      <t>
        The author gratefully acknowledges the work of Brent Callaghan and
        Tom Talpey producing the original RPC-over-RDMA Version One 
        specification <xref target="RFC5666" /> and also Tom's work in
        helping to clarify that specification. 
      </t>
      <t>
        The author also wishes to thank Chuck Lever for his work resurrecting 
        NFS support for RDMA in <xref target="rfc5666bis"/>, and for 
        helpful discussion regarding RPC-over-RDMA latency issues.
    
      </t>
    </section>
  </back>
</rfc>
