Skip to content

Torch

dtype serialization utilities for tensor clients.

This module provides a stable mapping between PyTorch torch.dtype objects and compact integer codes ("dcodes"). These codes can be used to serialize tensor metadata for network protocols, file formats, or any client-server interaction involving tensors.

Mappings

int8        <-> 12
int16       <-> 13
int32       <-> 14
int64       <-> 15
float32     <-> 24
float64     <-> 25
complex64   <-> 37
complex128  <-> 38

Example

from torch import float32, int64 from dtypes import dcodeof, dtypeof dcodeof(float32) 24 dtypeof(15) torch.int64

dcodeof(dtype)

Get the integer serialization code ("dcode") for a given torch.dtype.

Parameters

dtype : torch.dtype The PyTorch dtype to encode.

Returns

int The corresponding integer code. Returns 0 if the dtype is unsupported.

Examples

dcodeof(torch.float32) 24 dcodeof(torch.int16) 13 dcodeof(torch.bool) 0 # not supported

Source code in pytannic/torch/types.py
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
def dcodeof(dtype: dtype):
    """
    Get the integer serialization code ("dcode") for a given `torch.dtype`.

    Parameters
    ----------
    dtype : torch.dtype
        The PyTorch dtype to encode.

    Returns
    -------
    int
        The corresponding integer code. Returns `0` if the dtype is unsupported.

    Examples
    --------
    >>> dcodeof(torch.float32)
    24
    >>> dcodeof(torch.int16)
    13
    >>> dcodeof(torch.bool)
    0  # not supported
    """

    if dtype == int8:
        return 12
    elif dtype == int16:
        return 13
    elif dtype == int32:
        return 14
    elif dtype == int64:
        return 15
    elif dtype == float32:
        return 24
    elif dtype == float64:
        return 25
    elif dtype == complex64:
        return 37
    elif dtype == complex128:
        return 38
    else:
        return 0     

dtypeof(code)

Get the torch.dtype corresponding to a serialization code ("dcode").

Parameters

code : int The integer code to decode.

Returns

torch.dtype The corresponding dtype.

Raises

ValueError If the code is not recognized.

Examples

dtypeof(24) torch.float32 dtypeof(15) torch.int64 dtypeof(99) Traceback (most recent call last): ... ValueError: Unknown code

Source code in pytannic/torch/types.py
 97
 98
 99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
def dtypeof(code: int) -> dtype:

    """
    Get the `torch.dtype` corresponding to a serialization code ("dcode").

    Parameters
    ----------
    code : int
        The integer code to decode.

    Returns
    -------
    torch.dtype
        The corresponding dtype.

    Raises
    ------
    ValueError
        If the code is not recognized.

    Examples
    --------
    >>> dtypeof(24)
    torch.float32
    >>> dtypeof(15)
    torch.int64
    >>> dtypeof(99)
    Traceback (most recent call last):
        ...
    ValueError: Unknown code
    """

    if code == 12:
        return int8
    elif code == 13:
        return int16
    elif code == 14:
        return int32
    elif code == 15:
        return int64
    elif code == 24:
        return float32
    elif code == 25:
        return float64
    elif code == 37:
        return complex64
    elif code == 38:
        return complex128
    else:
        raise ValueError("Unknown code")

Tensor serialization utilities for network transport.

This module defines a compact binary protocol for sending and receiving PyTorch tensors across a network. A serialized tensor is laid out as:

[Header][Metadata][Raw Buffer]
  • The Header is defined in pytannic.header.Header and contains basic framing information (magic number, version, checksum, payload size).
  • The Metadata block encodes dtype, shape, and buffer size.
  • The Raw Buffer contains contiguous tensor bytes in row-major order.

Only CPU tensors are supported. GPU tensors are automatically moved to CPU before serialization.

Examples

import torch from pytannic.torch.tensor import serialize, deserialize x = torch.arange(6, dtype=torch.int32).reshape(2, 3) data = serialize(x) y = deserialize(data) torch.equal(x, y) True

Metadata dataclass

Tensor metadata for serialization.

Attributes

dcode : int Encoded dtype (see pytannic.torch.types.dcodeof). offset : int Byte offset into the raw buffer (currently always 0). nbytes : int Number of bytes in the raw tensor buffer. rank : int Tensor rank (number of dimensions). shape : tuple[int, ...] Tensor shape as a tuple of dimension sizes.

Source code in pytannic/torch/tensors.py
 47
 48
 49
 50
 51
 52
 53
 54
 55
 56
 57
 58
 59
 60
 61
 62
 63
 64
 65
 66
 67
 68
 69
 70
 71
 72
 73
 74
 75
 76
 77
 78
 79
 80
 81
 82
 83
 84
 85
 86
 87
 88
 89
 90
 91
 92
 93
 94
 95
 96
 97
 98
 99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
@dataclass
class Metadata: 
    """
    Tensor metadata for serialization.

    Attributes
    ----------
    dcode : int
        Encoded dtype (see `pytannic.torch.types.dcodeof`).
    offset : int
        Byte offset into the raw buffer (currently always 0).
    nbytes : int
        Number of bytes in the raw tensor buffer.
    rank : int
        Tensor rank (number of dimensions).
    shape : tuple[int, ...]
        Tensor shape as a tuple of dimension sizes.
    """

    dcode: int
    offset: int
    nbytes: int
    rank: int
    shape: tuple[int, ...]  

    def pack(self) -> bytes:
        """
        Serialize the metadata into a binary blob.

        Returns
        -------
        bytes
            Packed metadata using little-endian struct format.
        """
        return pack(self.format, self.dcode, self.offset, self.nbytes, self.rank, *self.shape)

    @classmethod
    def unpack(cls, data: bytes):  
        """
        Deserialize a `Metadata` instance from binary data.

        Parameters
        ----------
        data : bytes
            Binary blob containing packed metadata.

        Returns
        -------
        Metadata
            A new `Metadata` instance.
        """
        head = calcsize("<B Q Q B")
        dcode, offset, nbytes, rank = unpack("<B Q Q B", data[:head])
        if rank == 0:
            shape = ()
        else:
            shape = unpack(f"<{rank}Q", data[head:head + rank * 8])
        return cls(dcode, offset, nbytes, rank, shape) 

    @property
    def format(self) -> str: 
        """
        Struct format string for packing/unpacking tensor metadata.

        The format is `"<B Q Q B{rank}Q"`, which corresponds to the
        following C struct layout in little-endian order (no padding):

        .. code-block:: cpp

            struct Metadata<Tensor> {
                uint8_t dcode;        // 1 byte
                size_t offset;  // 8 bytes
                size_t nbytes;  // 8 bytes
                uint8_t rank;         // 1 byte
                size_t shape[rank]; // 8 bytes each
            };

        Layout (before shape):
        ----------------------
        - dcode  : 1 byte
        - offset : 8 bytes
        - nbytes : 8 bytes
        - rank   : 1 byte
        --------------------------------
        Fixed size = 18 bytes + (8 * rank) for shape

        Shape array:
        ------------
        - Each dimension size is stored as an `unsigned long long` (8 bytes).
        - The number of entries equals `rank`.

        Returns
        -------
        str
            A format string of the form `"<B Q Q B{rank}Q"`.
        """ 
        return f"<B Q Q B{self.rank}Q"

format property

Struct format string for packing/unpacking tensor metadata.

The format is "<B Q Q B{rank}Q", which corresponds to the following C struct layout in little-endian order (no padding):

.. code-block:: cpp

struct Metadata<Tensor> {
    uint8_t dcode;        // 1 byte
    size_t offset;  // 8 bytes
    size_t nbytes;  // 8 bytes
    uint8_t rank;         // 1 byte
    size_t shape[rank]; // 8 bytes each
};
Layout (before shape):
  • dcode : 1 byte
  • offset : 8 bytes
  • nbytes : 8 bytes
  • rank : 1 byte

Fixed size = 18 bytes + (8 * rank) for shape

Shape array:
  • Each dimension size is stored as an unsigned long long (8 bytes).
  • The number of entries equals rank.
Returns

str A format string of the form "<B Q Q B{rank}Q".

pack()

Serialize the metadata into a binary blob.

Returns

bytes Packed metadata using little-endian struct format.

Source code in pytannic/torch/tensors.py
72
73
74
75
76
77
78
79
80
81
def pack(self) -> bytes:
    """
    Serialize the metadata into a binary blob.

    Returns
    -------
    bytes
        Packed metadata using little-endian struct format.
    """
    return pack(self.format, self.dcode, self.offset, self.nbytes, self.rank, *self.shape)

unpack(data) classmethod

Deserialize a Metadata instance from binary data.

Parameters

data : bytes Binary blob containing packed metadata.

Returns

Metadata A new Metadata instance.

Source code in pytannic/torch/tensors.py
 83
 84
 85
 86
 87
 88
 89
 90
 91
 92
 93
 94
 95
 96
 97
 98
 99
100
101
102
103
104
@classmethod
def unpack(cls, data: bytes):  
    """
    Deserialize a `Metadata` instance from binary data.

    Parameters
    ----------
    data : bytes
        Binary blob containing packed metadata.

    Returns
    -------
    Metadata
        A new `Metadata` instance.
    """
    head = calcsize("<B Q Q B")
    dcode, offset, nbytes, rank = unpack("<B Q Q B", data[:head])
    if rank == 0:
        shape = ()
    else:
        shape = unpack(f"<{rank}Q", data[head:head + rank * 8])
    return cls(dcode, offset, nbytes, rank, shape) 

deserialize(data)

Deserialize a PyTorch tensor from a binary blob.

Parameters

data : bytes A binary blob produced by serialize.

Returns

torch.Tensor The reconstructed tensor with the same dtype and shape.

Raises

ValueError If the dcode in metadata is not recognized.

Notes

  • Uses torch.frombuffer, so the returned tensor shares memory with the input data buffer when possible.
Source code in pytannic/torch/tensors.py
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
def deserialize(data: bytes) -> Tensor:
    """
    Deserialize a PyTorch tensor from a binary blob.

    Parameters
    ----------
    data : bytes
        A binary blob produced by `serialize`.

    Returns
    -------
    torch.Tensor
        The reconstructed tensor with the same dtype and shape.

    Raises
    ------
    ValueError
        If the dcode in metadata is not recognized.

    Notes
    -----
    - Uses `torch.frombuffer`, so the returned tensor shares memory
      with the input `data` buffer when possible.
    """
    hsize = calcsize(Header.FORMAT) 
    head = calcsize("<B Q Q B")
    dcode, offset, nbytes, rank = unpack("<B Q Q B", data[hsize : hsize + head])

    msize = head + 8 * rank
    metadata = Metadata.unpack(data[hsize : hsize + msize]) 
    offset = hsize + msize
    buffer = bytearray(data[offset: offset + nbytes])
    return frombuffer(buffer, dtype=dtypeof(dcode)).reshape(metadata.shape) 

serialize(tensor)

Serialize a PyTorch tensor into a binary blob for network transport.

Parameters

tensor : torch.Tensor Input tensor. If the tensor is on a GPU, it will be moved to CPU.

Returns

bytes A binary blob representing the tensor, consisting of:

- Header
- Metadata
- Raw buffer (tensor bytes)

Notes

  • Only contiguous CPU tensors are supported.
  • Dtype is encoded using dcodeof.
Source code in pytannic/torch/tensors.py
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
def serialize(tensor: Tensor) -> bytes:
    """
    Serialize a PyTorch tensor into a binary blob for network transport.

    Parameters
    ----------
    tensor : torch.Tensor
        Input tensor. If the tensor is on a GPU, it will be moved to CPU.

    Returns
    -------
    bytes
        A binary blob representing the tensor, consisting of:

        - Header
        - Metadata
        - Raw buffer (tensor bytes)

    Notes
    -----
    - Only contiguous CPU tensors are supported.
    - Dtype is encoded using `dcodeof`.
    """
    if tensor.device.type != 'cpu':
        tensor = tensor.cpu()

    rank = tensor.dim()
    shape  = tensor.shape
    buffer = tensor.numpy().tobytes() 
    metadata = Metadata(dcode=dcodeof(tensor.dtype), offset=0, nbytes=len(buffer), rank=rank, shape=shape)  
    header = Header(magic=MAGIC, version=1, checksum=0xABCD , nbytes = calcsize(metadata.format) + len(buffer)) 
    return header.pack() + metadata.pack()+ buffer 

Parameter metadata serialization.

Unlike tensor payloads (which are typically transmitted over the network), parameters are usually stored in files (e.g., model checkpoints). This module defines a binary metadata structure for describing individual named tensor parameters.

The binary layout is:

[dcode (1B)][offset (8B)][nbytes (8B)][namelength (4B)][name (bytes)]

Attributes

  • dcode : dtype code (see pytannic.torch.types.dcodeof)
  • offset : byte offset to parameter data within the file
  • nbytes : number of bytes of parameter data
  • namelength : length of the name in bytes
  • name : parameter name (string)

Examples

m = Metadata(dcode=24, offset=128, nbytes=4096, namelength=4, name="fc1") blob = m.pack() Metadata.unpack(blob) Metadata(dcode=24, offset=128, nbytes=4096, namelength=4, name='')

Metadata dataclass

Metadata structure for a named parameter.

Attributes

dcode : int Encoded dtype for the parameter (see pytannic.torch.types.dcodeof). offset : int Byte offset in the file where the parameter data begins. nbytes : int Size of the parameter data in bytes. namelength : int Length of the parameter name in bytes. name : str Parameter name. Currently not packed/unpacked automatically.

Source code in pytannic/torch/parameters.py
 43
 44
 45
 46
 47
 48
 49
 50
 51
 52
 53
 54
 55
 56
 57
 58
 59
 60
 61
 62
 63
 64
 65
 66
 67
 68
 69
 70
 71
 72
 73
 74
 75
 76
 77
 78
 79
 80
 81
 82
 83
 84
 85
 86
 87
 88
 89
 90
 91
 92
 93
 94
 95
 96
 97
 98
 99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
@dataclass
class Metadata:
    """
    Metadata structure for a named parameter.

    Attributes
    ----------
    dcode : int
        Encoded dtype for the parameter (see `pytannic.torch.types.dcodeof`).
    offset : int
        Byte offset in the file where the parameter data begins.
    nbytes : int
        Size of the parameter data in bytes.
    namelength : int
        Length of the parameter name in bytes.
    name : str
        Parameter name. Currently not packed/unpacked automatically.
    """
    dcode: int
    offset: int
    nbytes: int
    namelength: int  
    name: str

    FORMAT = "<B Q Q I"   
    def pack(self) -> bytes: 
        """
        Serialize metadata into a binary blob.

        Returns
        -------
        bytes
            Packed metadata (excluding the parameter name).
        """
        return pack(self.FORMAT, self.dcode, self.offset, self.nbytes, self.namelength)

    @classmethod
    def unpack(cls, data: bytes):
        """
        Deserialize metadata from a binary blob.

        Parameters
        ----------
        data : bytes
            Raw binary data containing packed metadata.

        Returns
        -------
        Metadata
            A new `Metadata` instance. The `name` field is left empty,
            since it must be read separately from the stream.
        """
        fsize = calcsize(cls.FORMAT)
        dcode, offset, nbytes, namelength = unpack(cls.FORMAT, data[:fsize])
        return cls(dcode=dcode, offset=offset, nbytes=nbytes, namelength=namelength)

    @property
    def format(self) -> str:
        """
        Struct format string for this metadata.

        The format is `"<B Q Q I"`, which corresponds to the following
        C struct layout in little-endian order (no padding):

        .. code-block:: cpp

        struct Metadata<nn::Parameter> {
            uint8_t dcode;        // 1 byte
            size_t offset;        // 8 bytes
            size_t long nbytes;   // 8 bytes
            uint16_t namelength;  // 4 bytes 
        };

        Layout (before name):
        ---------------------
        - dcode      : 1 byte
        - offset     : 8 bytes
        - nbytes     : 8 bytes
        - namelength : 4 bytes
        --------------------------------
        Total        : 21 bytes + name

        Returns
        -------
        str
            Always `"<B Q Q I"`.
        """
        return self.FORMAT  

format property

Struct format string for this metadata.

The format is "<B Q Q I", which corresponds to the following C struct layout in little-endian order (no padding):

.. code-block:: cpp

struct Metadata { uint8_t dcode; // 1 byte size_t offset; // 8 bytes size_t long nbytes; // 8 bytes uint16_t namelength; // 4 bytes };

Layout (before name):
  • dcode : 1 byte
  • offset : 8 bytes
  • nbytes : 8 bytes
  • namelength : 4 bytes

Total : 21 bytes + name

Returns

str Always "<B Q Q I".

pack()

Serialize metadata into a binary blob.

Returns

bytes Packed metadata (excluding the parameter name).

Source code in pytannic/torch/parameters.py
68
69
70
71
72
73
74
75
76
77
def pack(self) -> bytes: 
    """
    Serialize metadata into a binary blob.

    Returns
    -------
    bytes
        Packed metadata (excluding the parameter name).
    """
    return pack(self.FORMAT, self.dcode, self.offset, self.nbytes, self.namelength)

unpack(data) classmethod

Deserialize metadata from a binary blob.

Parameters

data : bytes Raw binary data containing packed metadata.

Returns

Metadata A new Metadata instance. The name field is left empty, since it must be read separately from the stream.

Source code in pytannic/torch/parameters.py
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
@classmethod
def unpack(cls, data: bytes):
    """
    Deserialize metadata from a binary blob.

    Parameters
    ----------
    data : bytes
        Raw binary data containing packed metadata.

    Returns
    -------
    Metadata
        A new `Metadata` instance. The `name` field is left empty,
        since it must be read separately from the stream.
    """
    fsize = calcsize(cls.FORMAT)
    dcode, offset, nbytes, namelength = unpack(cls.FORMAT, data[:fsize])
    return cls(dcode=dcode, offset=offset, nbytes=nbytes, namelength=namelength)

Module serialization for checkpointing.

This module provides utilities to serialize a PyTorch torch.nn.Module into two companion binary files:

  1. Weights file (.tannic)
    Contains raw parameter tensors stored back-to-back.

Layout: [Header][TensorBytes...]

  • Header = pytannic.header.Header
  • Each parameter is written as a raw contiguous CPU buffer in the order returned by state_dict().

  • Metadata file (.metadata.tannic)
    Contains parameter descriptors (dtype, offset, name, etc.).

Layout: [Header][Metadata1][Name1][Metadata2][Name2]...

  • Header = pytannic.header.Header
  • Each Metadata = pytannic.torch.parameters.Metadata
  • Name is stored as UTF-8 with length namelength

These two files together allow reconstructing the module parameters.

Notes

  • Parameters are always written in CPU memory order.
  • Gradients are not serialized (only .data).
  • Offsets in metadata are relative to the weights file.

write(module, filename)

Serialize a torch.nn.Module into .tannic files.

Parameters

module : torch.nn.Module The PyTorch module to serialize. Only parameters from module.state_dict() are saved. filename : str Base filename for output. Two files are written:

- `<stem>.tannic` for raw parameter data
- `<stem>.metadata.tannic` for metadata

File Layout

Weights file (.tannic):

.. code-block:: cpp

struct Header header;
unsigned char buffer[];   // concatenated tensor data

Metadata file (.metadata.tannic):

.. code-block:: cpp

struct Header header;
struct Metadata<nn::Parameter> {
    uint8_t dcode;        // 1 byte
    size_t offset;        // 8 bytes
    size_t long nbytes;   // 8 bytes
    uint16_t namelength;  // 4 bytes 
};

Notes

  • The two files must always be kept together.
  • The header nbytes field includes only payload size (not counting the header itself).
Source code in pytannic/torch/modules.py
 56
 57
 58
 59
 60
 61
 62
 63
 64
 65
 66
 67
 68
 69
 70
 71
 72
 73
 74
 75
 76
 77
 78
 79
 80
 81
 82
 83
 84
 85
 86
 87
 88
 89
 90
 91
 92
 93
 94
 95
 96
 97
 98
 99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
def write(module: Module, filename: str) -> None:
    """
    Serialize a `torch.nn.Module` into `.tannic` files.

    Parameters
    ----------
    module : torch.nn.Module
        The PyTorch module to serialize. Only parameters from
        `module.state_dict()` are saved.
    filename : str
        Base filename for output. Two files are written:

        - `<stem>.tannic` for raw parameter data
        - `<stem>.metadata.tannic` for metadata

    File Layout
    -----------
    **Weights file (`.tannic`):**

    .. code-block:: cpp 

        struct Header header;
        unsigned char buffer[];   // concatenated tensor data

    **Metadata file (`.metadata.tannic`):**

    .. code-block:: cpp

        struct Header header;
        struct Metadata<nn::Parameter> {
            uint8_t dcode;        // 1 byte
            size_t offset;        // 8 bytes
            size_t long nbytes;   // 8 bytes
            uint16_t namelength;  // 4 bytes 
        }; 

    Notes
    -----
    - The two files must always be kept together.
    - The header `nbytes` field includes only payload size
      (not counting the header itself).
    """
    path = Path(filename) 
    state: dict[str, Parameter] = module.state_dict()
    metadata: list[Metadata] = []  

    with open(f'{path.stem}.tannic', 'wb') as file:   
        nbytes = sum(parameter.nbytes for parameter in module.parameters()) 
        header = Header(magic=MAGIC, version=1, checksum=0xABCD, nbytes=nbytes) 
        offset = 0
        file.write(header.pack())  
        for name, parameter in state.items():  
            metadata.append(Metadata( 
                dcode=dcodeof(parameter.dtype),
                offset=offset,
                nbytes=parameter.nbytes,   
                namelength=len(name),
                name=name,  
            )) 
            offset += parameter.nbytes   
            file.write(parameter.detach().cpu().numpy().tobytes())   

    with open(f'{path.stem}.metadata.tannic', 'wb') as file:     
        nbytes = sum(calcsize(obj.format) + obj.namelength for obj in metadata)
        header = Header(magic=MAGIC, version=1, checksum=0xABCD, nbytes=nbytes) 
        file.write(header.pack())  
        print(metadata)
        for object in metadata: 
            file.write(object.pack())       
            file.write(object.name.encode('utf-8'))