评价此页
non_blockingpin_memory() 在 PyTorch 中的应用"

PyTorch 中 non_blockingpin_memory() 的良好使用指南#

创建日期:2024年7月31日 | 最后更新:2026年4月01日 | 最后验证:2024年11月05日

作者Vincent Moens

简介#

将数据从 CPU 传输到 GPU 是许多 PyTorch 应用程序的基础。用户必须了解在不同设备间移动数据的最有效工具和选项。本教程考察了 PyTorch 中设备间数据传输的两种关键方法:pin_memory() 和带有 non_blocking=True 选项的 to()

您将学到什么#

可以通过异步传输和内存固定(memory pinning)来优化从 CPU 到 GPU 的张量传输。然而,这里有一些重要的注意事项。

  • 使用 tensor.pin_memory().to(device, non_blocking=True) 的速度可能会比直接使用 tensor.to(device) 慢两倍。

  • 通常情况下,tensor.to(device, non_blocking=True) 是提升传输速度的有效选择。

  • 虽然 cpu_tensor.to("cuda", non_blocking=True).mean() 可以正确执行,但尝试使用 cuda_tensor.to("cpu", non_blocking=True).mean() 将导致错误的输出。

前言#

本教程中报告的性能情况取决于构建教程所使用的系统。尽管结论适用于不同系统,但具体观察结果可能会因可用硬件的不同而略有差异,特别是在较旧的硬件上。本教程的主要目标是提供一个理解 CPU 到 GPU 数据传输的理论框架。然而,任何设计决策都应根据具体情况,并参考基准测试的吞吐量测量结果以及手头任务的具体需求进行调整。

import torch

assert torch.cuda.is_available(), "A cuda device is required to run this tutorial"

本教程需要安装 tensordict。如果您的环境中尚未安装 tensordict,请在单独的单元格中运行以下命令进行安装:

# Install tensordict with the following command
!pip3 install tensordict

我们将首先概述这些概念背后的理论,然后转向这些功能的具体测试示例。

背景#

内存管理基础#

当在 PyTorch 中创建 CPU 张量时,该张量的内容需要放置在内存中。我们在此讨论的内存是一个相当复杂的概念,值得仔细研究。我们区分由内存管理单元处理的两种内存类型:RAM(为简单起见)和磁盘上的交换空间(可能是也可能不是硬盘)。磁盘和 RAM(物理内存)中的可用空间共同构成了虚拟内存,这是对所有可用资源的一种抽象。简而言之,虚拟内存使得可用空间大于仅从 RAM 中获取的空间,并创造了一种主内存比实际更大的错觉。

在正常情况下,常规 CPU 张量是可分页的(pageable),这意味着它被划分为称为“页面”的块,这些块可以存在于虚拟内存中的任何位置(无论是在 RAM 还是磁盘上)。如前所述,这样做的好处是内存看起来比实际的主内存更大。

通常,当程序访问不在 RAM 中的页面时,会发生“缺页中断”(page fault),操作系统(OS)随后会将该页面调入 RAM(即“换入”或“页入”)。反过来,OS 可能不得不换出(或“页出”)另一个页面以便为新页面腾出空间。

与可分页内存相反,固定内存(Pinned memory,或称页锁定/不可分页内存)是一种不能被换出到磁盘的内存。它允许更快且更可预测的访问时间,但缺点是它比可分页内存(即主内存)更受限制。

CUDA 与(非)可分页内存#

为了理解 CUDA 如何将张量从 CPU 复制到 CUDA,让我们考虑上述两种情况:

  • 如果内存是页锁定的,设备可以直接访问主内存中的该内存。内存地址是明确定义的,需要读取这些数据的函数可以得到显著加速。

  • 如果内存是可分页的,所有页面在被发送到 GPU 之前都必须先调入主内存。此操作可能需要时间,且不如在页锁定张量上执行时那样可预测。

更准确地说,当 CUDA 将可分页数据从 CPU 发送到 GPU 时,它必须在进行传输之前首先创建该数据的一个页锁定副本。

带有 non_blocking=True 的异步与同步操作(CUDA cudaMemcpyAsync#

当执行从主机(例如 CPU)到设备(例如 GPU)的复制时,CUDA 工具包提供了相对于主机同步或异步执行这些操作的方式。

在实践中,调用 to() 时,PyTorch 总是调用 cudaMemcpyAsync。如果 non_blocking=False(默认值),则在每次 cudaMemcpyAsync 后都会调用 cudaStreamSynchronize,从而使对 to() 的调用在主线程中阻塞。如果 non_blocking=True,则不会触发同步,主机上的主线程也不会被阻塞。因此,从主机角度来看,可以同时将多个张量发送到设备,因为线程不需要等待一次传输完成即可启动下一次传输。

注意

通常,传输在设备端是阻塞的(即使在主机端不是):在执行另一个操作时,设备上的复制无法发生。然而,在某些高级场景下,可以在 GPU 端同时进行复制和内核执行。如下例所示,必须满足三个要求才能实现这一点:

  1. 设备必须至少有一个空闲的 DMA(直接内存访问)引擎。现代 GPU 架构(如 Volterra、Tesla 或 H100 设备)具有不止一个 DMA 引擎。

  2. 传输必须在单独的、非默认的 CUDA 流上进行。在 PyTorch 中,可以使用 Stream 来处理 CUDA 流。

  3. 源数据必须位于固定内存中。

我们通过在以下脚本上运行分析(profiles)来演示这一点。

import contextlib

from torch.cuda import Stream


s = Stream()

torch.manual_seed(42)
t1_cpu_pinned = torch.randn(1024**2 * 5, pin_memory=True)
t2_cpu_paged = torch.randn(1024**2 * 5, pin_memory=False)
t3_cuda = torch.randn(1024**2 * 5, device="cuda:0")

assert torch.cuda.is_available()
device = torch.device("cuda", torch.cuda.current_device())


# The function we want to profile
def inner(pinned: bool, streamed: bool):
    with torch.cuda.stream(s) if streamed else contextlib.nullcontext():
        if pinned:
            t1_cuda = t1_cpu_pinned.to(device, non_blocking=True)
        else:
            t2_cuda = t2_cpu_paged.to(device, non_blocking=True)
        t_star_cuda_h2d_event = s.record_event()
    # This operation can be executed during the CPU to GPU copy if and only if the tensor is pinned and the copy is
    #  done in the other stream
    t3_cuda_mul = t3_cuda * t3_cuda * t3_cuda
    t3_cuda_h2d_event = torch.cuda.current_stream().record_event()
    t_star_cuda_h2d_event.synchronize()
    t3_cuda_h2d_event.synchronize()


# Our profiler: profiles the `inner` function and stores the results in a .json file
def benchmark_with_profiler(
    pinned,
    streamed,
) -> None:
    torch._C._profiler._set_cuda_sync_enabled_val(True)
    wait, warmup, active = 1, 1, 2
    num_steps = wait + warmup + active
    rank = 0
    with torch.profiler.profile(
        activities=[
            torch.profiler.ProfilerActivity.CPU,
            torch.profiler.ProfilerActivity.CUDA,
        ],
        schedule=torch.profiler.schedule(
            wait=wait, warmup=warmup, active=active, repeat=1, skip_first=1
        ),
    ) as prof:
        for step_idx in range(1, num_steps + 1):
            inner(streamed=streamed, pinned=pinned)
            if rank is None or rank == 0:
                prof.step()
    prof.export_chrome_trace(f"trace_streamed{int(streamed)}_pinned{int(pinned)}.json")

将这些分析跟踪加载到 Chrome (chrome://tracing) 中显示以下结果:首先,让我们看看如果可分页张量在主流中发送到 GPU 后再对 t3_cuda 执行算术运算会发生什么。

benchmark_with_profiler(streamed=False, pinned=False)

使用固定张量并不会对跟踪产生太大影响,这两个操作仍然是按顺序执行的。

benchmark_with_profiler(streamed=False, pinned=True)

将可分页张量发送到单独流上的 GPU 也是一个阻塞操作。

benchmark_with_profiler(streamed=True, pinned=False)

只有当在单独流上将固定张量复制到 GPU 时,它才能与在主流上执行的另一个 CUDA 内核重叠。

benchmark_with_profiler(streamed=True, pinned=True)

PyTorch 的视角#

pin_memory()#

PyTorch 提供了通过 pin_memory() 方法和构造函数参数创建并发送张量到页锁定内存的可能性。在初始化了 CUDA 的机器上,CPU 张量可以通过 pin_memory() 方法转换为固定内存。重要的是,pin_memory 在主机的主线程上是阻塞的:它会等待张量被复制到页锁定内存后,才会执行下一个操作。新的张量可以直接通过诸如 zeros()ones() 等构造函数在固定内存中创建。

让我们检查一下固定内存并发送张量到 CUDA 的速度。

import torch
import gc
from torch.utils.benchmark import Timer
import matplotlib.pyplot as plt


def timer(cmd):
    median = (
        Timer(cmd, globals=globals())
        .adaptive_autorange(min_run_time=1.0, max_run_time=20.0)
        .median
        * 1000
    )
    print(f"{cmd}: {median: 4.4f} ms")
    return median


# A tensor in pageable memory
pageable_tensor = torch.randn(1_000_000)

# A tensor in page-locked (pinned) memory
pinned_tensor = torch.randn(1_000_000, pin_memory=True)

# Runtimes:
pageable_to_device = timer("pageable_tensor.to('cuda:0')")
pinned_to_device = timer("pinned_tensor.to('cuda:0')")
pin_mem = timer("pageable_tensor.pin_memory()")
pin_mem_to_device = timer("pageable_tensor.pin_memory().to('cuda:0')")

# Ratios:
r1 = pinned_to_device / pageable_to_device
r2 = pin_mem_to_device / pageable_to_device

# Create a figure with the results
fig, ax = plt.subplots()

xlabels = [0, 1, 2]
bar_labels = [
    "pageable_tensor.to(device) (1x)",
    f"pinned_tensor.to(device) ({r1:4.2f}x)",
    f"pageable_tensor.pin_memory().to(device) ({r2:4.2f}x)"
    f"\npin_memory()={100*pin_mem/pin_mem_to_device:.2f}% of runtime.",
]
values = [pageable_to_device, pinned_to_device, pin_mem_to_device]
colors = ["tab:blue", "tab:red", "tab:orange"]
ax.bar(xlabels, values, label=bar_labels, color=colors)

ax.set_ylabel("Runtime (ms)")
ax.set_title("Device casting runtime (pin-memory)")
ax.set_xticks([])
ax.legend()

plt.show()

# Clear tensors
del pageable_tensor, pinned_tensor
_ = gc.collect()
Device casting runtime (pin-memory)
pageable_tensor.to('cuda:0'):  0.3708 ms
pinned_tensor.to('cuda:0'):  0.3170 ms
pageable_tensor.pin_memory():  0.1156 ms
pageable_tensor.pin_memory().to('cuda:0'):  0.4390 ms

我们可以观察到,将固定内存张量转换为 GPU 的速度确实比可分页张量快得多,因为在底层,可分页张量在发送到 GPU 之前必须先复制到固定内存中。

然而,与某种普遍看法相反,在将可分页张量转换为 GPU 之前对其调用 pin_memory() 不会带来显著的速度提升;相反,此调用通常比直接执行传输要慢。这是合理的,因为我们实际上是在要求 Python 执行一个 CUDA 在从主机到设备复制数据时无论如何都会执行的操作。

注意

PyTorch 对 pin_memory 的实现依赖于通过 cudaHostAlloc 在固定内存中创建一个全新的存储,在极少数情况下,这可能比 cudaMemcpy 分块转换数据更快。同样,这里的观察结果可能会根据可用硬件、发送张量的大小或可用 RAM 的数量而有所不同。

non_blocking=True#

如前所述,许多 PyTorch 操作可以选择通过 non_blocking 参数相对于主机异步执行。

在此,为了准确评估使用 non_blocking 的好处,我们将设计一个稍微复杂的实验,因为我们想要评估在调用和不调用 non_blocking 的情况下,将多个张量发送到 GPU 的速度有多快。

# A simple loop that copies all tensors to cuda
def copy_to_device(*tensors):
    result = []
    for tensor in tensors:
        result.append(tensor.to("cuda:0"))
    return result


# A loop that copies all tensors to cuda asynchronously
def copy_to_device_nonblocking(*tensors):
    result = []
    for tensor in tensors:
        result.append(tensor.to("cuda:0", non_blocking=True))
    # We need to synchronize
    torch.cuda.synchronize()
    return result


# Create a list of tensors
tensors = [torch.randn(1000) for _ in range(1000)]
to_device = timer("copy_to_device(*tensors)")
to_device_nonblocking = timer("copy_to_device_nonblocking(*tensors)")

# Ratio
r1 = to_device_nonblocking / to_device

# Plot the results
fig, ax = plt.subplots()

xlabels = [0, 1]
bar_labels = [f"to(device) (1x)", f"to(device, non_blocking=True) ({r1:4.2f}x)"]
colors = ["tab:blue", "tab:red"]
values = [to_device, to_device_nonblocking]

ax.bar(xlabels, values, label=bar_labels, color=colors)

ax.set_ylabel("Runtime (ms)")
ax.set_title("Device casting runtime (non-blocking)")
ax.set_xticks([])
ax.legend()

plt.show()
Device casting runtime (non-blocking)
copy_to_device(*tensors):  20.0170 ms
copy_to_device_nonblocking(*tensors):  14.9903 ms

为了更好地理解这里发生了什么,让我们分析这两个函数。

from torch.profiler import profile, ProfilerActivity


def profile_mem(cmd):
    with profile(activities=[ProfilerActivity.CPU]) as prof:
        exec(cmd)
    print(cmd)
    print(prof.key_averages().table(row_limit=10))

让我们先看看使用常规 to(device) 的调用堆栈:

print("Call to `to(device)`", profile_mem("copy_to_device(*tensors)"))
copy_to_device(*tensors)
-------------------------  ------------  ------------  ------------  ------------  ------------  ------------
                     Name    Self CPU %      Self CPU   CPU total %     CPU total  CPU time avg    # of Calls
-------------------------  ------------  ------------  ------------  ------------  ------------  ------------
                 aten::to         4.38%       1.090ms       100.00%      24.851ms      24.851us          1000
           aten::_to_copy        12.66%       3.145ms        95.62%      23.761ms      23.761us          1000
      aten::empty_strided        20.85%       5.181ms        20.85%       5.181ms       5.181us          1000
              aten::copy_        24.99%       6.211ms        62.11%      15.435ms      15.435us          1000
    cudaStreamIsCapturing         1.75%     434.873us         1.75%     434.873us       0.435us          1000
          cudaMemcpyAsync        15.24%       3.787ms        15.24%       3.787ms       3.787us          1000
    cudaStreamSynchronize        20.13%       5.002ms        20.13%       5.002ms       5.002us          1000
-------------------------  ------------  ------------  ------------  ------------  ------------  ------------
Self CPU time total: 24.851ms

Call to `to(device)` None

现在看看 non_blocking 版本:

print(
    "Call to `to(device, non_blocking=True)`",
    profile_mem("copy_to_device_nonblocking(*tensors)"),
)
copy_to_device_nonblocking(*tensors)
-------------------------  ------------  ------------  ------------  ------------  ------------  ------------
                     Name    Self CPU %      Self CPU   CPU total %     CPU total  CPU time avg    # of Calls
-------------------------  ------------  ------------  ------------  ------------  ------------  ------------
                 aten::to         5.24%       1.050ms        99.88%      20.027ms      20.027us          1000
           aten::_to_copy        14.85%       2.978ms        94.64%      18.977ms      18.977us          1000
      aten::empty_strided        24.63%       4.939ms        24.63%       4.939ms       4.939us          1000
              aten::copy_        34.13%       6.843ms        55.16%      11.060ms      11.060us          1000
    cudaStreamIsCapturing         2.44%     488.307us         2.44%     488.307us       0.488us          1000
          cudaMemcpyAsync        18.59%       3.728ms        18.59%       3.728ms       3.728us          1000
    cudaDeviceSynchronize         0.12%      24.111us         0.12%      24.111us      24.111us             1
-------------------------  ------------  ------------  ------------  ------------  ------------  ------------
Self CPU time total: 20.051ms

Call to `to(device, non_blocking=True)` None

使用 non_blocking=True 时结果无疑更好,因为所有传输都在主机端同时启动,并且只执行一次同步。

其收益将取决于张量的数量和大小,以及所使用的硬件。

注意

有趣的是,阻塞式的 to("cuda") 实际上执行的是与带有 non_blocking=True 时相同的异步设备转换操作(cudaMemcpyAsync),只是在每次复制后都有一个同步点。

协同效应#

既然我们已经指出将已经在固定内存中的张量传输到 GPU 比从可分页内存传输更快,并且我们知道异步执行这些传输也比同步执行更快,我们可以对这些方法的组合进行基准测试。首先,让我们编写几个新的函数,分别对每个张量调用 pin_memoryto(device)

def pin_copy_to_device(*tensors):
    result = []
    for tensor in tensors:
        result.append(tensor.pin_memory().to("cuda:0"))
    return result


def pin_copy_to_device_nonblocking(*tensors):
    result = []
    for tensor in tensors:
        result.append(tensor.pin_memory().to("cuda:0", non_blocking=True))
    # We need to synchronize
    torch.cuda.synchronize()
    return result

对于较大张量的较大批次,使用 pin_memory() 的好处更为显著。

tensors = [torch.randn(1_000_000) for _ in range(1000)]
page_copy = timer("copy_to_device(*tensors)")
page_copy_nb = timer("copy_to_device_nonblocking(*tensors)")

tensors_pinned = [torch.randn(1_000_000, pin_memory=True) for _ in range(1000)]
pinned_copy = timer("copy_to_device(*tensors_pinned)")
pinned_copy_nb = timer("copy_to_device_nonblocking(*tensors_pinned)")

pin_and_copy = timer("pin_copy_to_device(*tensors)")
pin_and_copy_nb = timer("pin_copy_to_device_nonblocking(*tensors)")

# Plot
strategies = ("pageable copy", "pinned copy", "pin and copy")
blocking = {
    "blocking": [page_copy, pinned_copy, pin_and_copy],
    "non-blocking": [page_copy_nb, pinned_copy_nb, pin_and_copy_nb],
}

x = torch.arange(3)
width = 0.25
multiplier = 0


fig, ax = plt.subplots(layout="constrained")

for attribute, runtimes in blocking.items():
    offset = width * multiplier
    rects = ax.bar(x + offset, runtimes, width, label=attribute)
    ax.bar_label(rects, padding=3, fmt="%.2f")
    multiplier += 1

# Add some text for labels, title and custom x-axis tick labels, etc.
ax.set_ylabel("Runtime (ms)")
ax.set_title("Runtime (pin-mem and non-blocking)")
ax.set_xticks([0, 1, 2])
ax.set_xticklabels(strategies)
plt.setp(ax.get_xticklabels(), rotation=45, ha="right", rotation_mode="anchor")
ax.legend(loc="upper left", ncols=3)

plt.show()

del tensors, tensors_pinned
_ = gc.collect()
Runtime (pin-mem and non-blocking)
copy_to_device(*tensors):  395.4713 ms
copy_to_device_nonblocking(*tensors):  309.1386 ms
copy_to_device(*tensors_pinned):  317.6252 ms
copy_to_device_nonblocking(*tensors_pinned):  299.9406 ms
pin_copy_to_device(*tensors):  565.4348 ms
pin_copy_to_device_nonblocking(*tensors):  324.7638 ms

其他复制方向(GPU -> CPU,CPU -> MPS)#

到目前为止,我们一直假设从 CPU 到 GPU 的异步复制是安全的。这通常是正确的,因为 CUDA 会自动处理同步,以确保在读取时所访问的数据是有效的(只要张量位于可分页内存中)。

然而,在其他情况下我们不能做同样的假设:当张量放置在固定内存中时,在调用主机到设备传输后修改原始副本可能会损坏 GPU 上接收到的数据。同样,当传输在相反方向(从 GPU 到 CPU,或从除 CPU 或 GPU 之外的任何设备到非 CUDA 管理的 GPU(如 MPS))进行时,如果不进行显式同步,则无法保证 GPU 上读取的数据是有效的。

在这些场景中,这些传输不能保证在数据访问时复制已完成。因此,主机上的数据可能是不完整或错误的,实际上变成了垃圾数据。

让我们首先用一个固定内存张量来演示这一点。

DELAY = 100000000
try:
    i = -1
    for i in range(100):
        # Create a tensor in pin-memory
        cpu_tensor = torch.ones(1024, 1024, pin_memory=True)
        torch.cuda.synchronize()
        # Send the tensor to CUDA
        cuda_tensor = cpu_tensor.to("cuda", non_blocking=True)
        torch.cuda._sleep(DELAY)
        # Corrupt the original tensor
        cpu_tensor.zero_()
        assert (cuda_tensor == 1).all()
    print("No test failed with non_blocking and pinned tensor")
except AssertionError:
    print(f"{i}th test failed with non_blocking and pinned tensor. Skipping remaining tests")
1th test failed with non_blocking and pinned tensor. Skipping remaining tests

使用可分页张量总是有效的。

i = -1
for i in range(100):
    # Create a tensor in pageable memory
    cpu_tensor = torch.ones(1024, 1024)
    torch.cuda.synchronize()
    # Send the tensor to CUDA
    cuda_tensor = cpu_tensor.to("cuda", non_blocking=True)
    torch.cuda._sleep(DELAY)
    # Corrupt the original tensor
    cpu_tensor.zero_()
    assert (cuda_tensor == 1).all()
print("No test failed with non_blocking and pageable tensor")
No test failed with non_blocking and pageable tensor

现在让我们演示 CUDA 到 CPU 在没有同步的情况下也无法产生可靠的输出。

tensor = (
    torch.arange(1, 1_000_000, dtype=torch.double, device="cuda")
    .expand(100, 999999)
    .clone()
)
torch.testing.assert_close(
    tensor.mean(), torch.tensor(500_000, dtype=torch.double, device="cuda")
), tensor.mean()
try:
    i = -1
    for i in range(100):
        cpu_tensor = tensor.to("cpu", non_blocking=True)
        torch.testing.assert_close(
            cpu_tensor.mean(), torch.tensor(500_000, dtype=torch.double)
        )
    print("No test failed with non_blocking")
except AssertionError:
    print(f"{i}th test failed with non_blocking. Skipping remaining tests")
try:
    i = -1
    for i in range(100):
        cpu_tensor = tensor.to("cpu", non_blocking=True)
        torch.cuda.synchronize()
        torch.testing.assert_close(
            cpu_tensor.mean(), torch.tensor(500_000, dtype=torch.double)
        )
    print("No test failed with synchronize")
except AssertionError:
    print(f"One test failed with synchronize: {i}th assertion!")
0th test failed with non_blocking. Skipping remaining tests
No test failed with synchronize

通常情况下,只有当目标是支持 CUDA 的设备且原始张量位于可分页内存中时,到设备的异步复制才可以在无需显式同步的情况下保持安全。

总而言之,使用 non_blocking=True 时,将数据从 CPU 复制到 GPU 是安全的;但对于任何其他方向,尽管仍然可以使用 non_blocking=True,但用户必须确保在访问数据之前执行设备同步。

实践建议#

根据我们的观察,现在可以总结出一些初步建议:

通常,non_blocking=True 将提供良好的吞吐量,无论原始张量是否位于固定内存中。如果张量已经在固定内存中,则可以加速传输;但如果从 Python 主线程手动将其发送到固定内存,则这在主机上是一个阻塞操作,因此会抵消使用 non_blocking=True 的大部分好处(因为 CUDA 反正也会进行 pin_memory 传输)。

现在人们可能会合理地问 pin_memory() 方法有什么用。在下一节中,我们将进一步探讨如何利用它来更进一步加速数据传输。

其他注意事项#

众所周知,PyTorch 提供了一个 DataLoader 类,其构造函数接受 pin_memory 参数。考虑到我们之前关于 pin_memory 的讨论,您可能想知道,如果内存固定本质上是阻塞的,DataLoader 是如何设法加速数据传输的。

关键在于 DataLoader 使用了一个单独的线程来处理从可分页内存到固定内存的数据传输,从而防止了主线程的任何阻塞。

为了说明这一点,我们将使用来自同名库的 TensorDict 原语。调用 to() 时,默认行为是将张量异步发送到设备,随后紧接着执行一次 torch.device.synchronize() 调用。

此外,TensorDict.to() 包含一个 non_blocking_pin 选项,该选项会启动多个线程在进行 to(device) 之前执行 pin_memory()。如下例所示,这种方法可以进一步加速数据传输。

from tensordict import TensorDict
import torch
from torch.utils.benchmark import Timer
import matplotlib.pyplot as plt

# Create the dataset
td = TensorDict({str(i): torch.randn(1_000_000) for i in range(1000)})

# Runtimes
copy_blocking = timer("td.to('cuda:0', non_blocking=False)")
copy_non_blocking = timer("td.to('cuda:0')")
copy_pin_nb = timer("td.to('cuda:0', non_blocking_pin=True, num_threads=0)")
copy_pin_multithread_nb = timer("td.to('cuda:0', non_blocking_pin=True, num_threads=4)")

# Rations
r1 = copy_non_blocking / copy_blocking
r2 = copy_pin_nb / copy_blocking
r3 = copy_pin_multithread_nb / copy_blocking

# Figure
fig, ax = plt.subplots()

xlabels = [0, 1, 2, 3]
bar_labels = [
    "Blocking copy (1x)",
    f"Non-blocking copy ({r1:4.2f}x)",
    f"Blocking pin, non-blocking copy ({r2:4.2f}x)",
    f"Non-blocking pin, non-blocking copy ({r3:4.2f}x)",
]
values = [copy_blocking, copy_non_blocking, copy_pin_nb, copy_pin_multithread_nb]
colors = ["tab:blue", "tab:red", "tab:orange", "tab:green"]

ax.bar(xlabels, values, label=bar_labels, color=colors)

ax.set_ylabel("Runtime (ms)")
ax.set_title("Device casting runtime")
ax.set_xticks([])
ax.legend()

plt.show()
Device casting runtime
td.to('cuda:0', non_blocking=False):  391.5463 ms
td.to('cuda:0'):  310.5651 ms
td.to('cuda:0', non_blocking_pin=True, num_threads=0):  311.5503 ms
td.to('cuda:0', non_blocking_pin=True, num_threads=4):  301.4078 ms

在此示例中,我们正在将许多大张量从 CPU 传输到 GPU。此场景非常适合利用多线程 pin_memory(),这可以显著增强性能。然而,如果张量很小,多线程带来的开销可能超过收益。同样,如果只有几个张量,在单独线程上固定张量的优势也会变得有限。

另外需要注意的是,虽然在固定内存中创建永久缓冲区来缓冲张量后再传输到 GPU 似乎很有优势,但这并不一定能加速计算。由于复制数据到固定内存所造成的固有瓶颈仍然是一个限制因素。

此外,将位于磁盘(无论是共享内存还是文件)上的数据传输到 GPU 通常需要一个中间步骤,即将数据复制到固定内存(位于 RAM 中)。在这种上下文中,为大数据传输使用 non_blocking 可能会显著增加 RAM 消耗,从而可能导致不利影响。

在实践中,没有万能的解决方案。结合 non_blocking 传输使用多线程 pin_memory 的有效性取决于多种因素,包括特定的系统、操作系统、硬件以及正在执行的任务的性质。在尝试加速 CPU 和 GPU 之间的数据传输或比较不同场景的吞吐量时,这里有一份需要检查的因素列表:

  • 可用核心数

    有多少 CPU 核心可用?系统是否与其他可能竞争资源的用户或进程共享?

  • 核心利用率

    CPU 核心是否被其他进程大量占用?应用程序是否在进行数据传输的同时执行其他 CPU 密集型任务?

  • 内存利用率

    当前使用了多少可分页和页锁定内存?是否有足够的剩余内存来分配额外的固定内存而不影响系统性能?请记住,没有什么是免费的,例如 pin_memory 会消耗 RAM,并可能影响其他任务。

  • CUDA 设备能力

    GPU 是否支持多个 DMA 引擎进行并行数据传输?所使用的 CUDA 设备有哪些特定的功能和限制?

  • 发送张量的数量

    在典型操作中传输了多少张量?

  • 要发送的张量大小

    传输的张量大小是多少?少量大张量或大量小张量可能无法从相同的传输程序中受益。

  • 系统架构

    系统的架构是如何影响数据传输速度的(例如,总线速度、网络延迟)?

此外,在固定内存中分配大量张量或大尺寸张量会占用相当大一部分 RAM。这减少了分页等其他关键操作的可用内存,从而可能对算法的整体性能产生负面影响。

结论#

在整个教程中,我们探讨了从主机发送张量到设备时影响传输速度和内存管理的几个关键因素。我们了解到,使用 non_blocking=True 通常会加速数据传输,并且如果实现得当,pin_memory() 也可以提高性能。然而,这些技术需要仔细的设计和校准才能发挥作用。

请记住,分析您的代码并密切关注内存消耗对于优化资源使用和获得最佳性能至关重要。

更多资源#

如果您在使用 CUDA 设备时遇到内存复制问题,或者想了解更多关于本教程所讨论内容的详细信息,请查阅以下参考资料:

脚本的总运行时间:(1 分 3.163 秒)