Skip to content
Permalink

Comparing changes

Choose two branches to see what’s changed or to start a new pull request. If you need to, you can also or learn more about diff comparisons.

Open a pull request

Create a new pull request by comparing changes across two branches. If you need to, you can also . Learn more about diff comparisons here.
base repository: parca-dev/parca-agent
Failed to load repositories. Confirm that selected base ref is valid, then try again.
Loading
base: main
Choose a base ref
...
head repository: parca-dev/parca-agent
Failed to load repositories. Confirm that selected head ref is valid, then try again.
Loading
compare: padata-convert
Choose a head ref
Checking mergeability… Don’t worry, you can still create the pull request.
  • 2 commits
  • 34 files changed
  • 1 contributor

Commits on Aug 28, 2026

  1. WORK IN PROGRESS

    reporter: add OTLP/profiles as a third export format
    
    Adds an OTLP/profiles exporter alongside the Arrow v1 and v2 paths,
    selected with --remote-store-format=otlp. Verified end to end against a
    Dash0 collector into ClickHouse, rendering symbolized flamegraphs.
    
    The conversion is ours rather than upstream's. Upstream's lives in
    reporter/internal/pdata, which is unimportable, and no extension point
    would have helped: pprofile.Profile carries exactly one SampleType, so a
    memory profile with four value axes needs four Profiles, and oomprof
    never passes through an origin at all; baseReporter also rejects the CUDA
    and GPU-PC origins before the encoding is reached. Owning it means every
    origin the tracer produces has a sample type, including GPU and memory.
    
    reporter/frames.go extracts frame classification into one function that
    both the Arrow writers and the pprofile builder consume, so a fork bump
    that changes libpf.Frame semantics produces one compile error rather than
    two silent divergences. The label pipeline, executable tracker, metrics
    bridge, counters and telemetry providers are extracted the same way, and
    reporter.New now takes a Config struct instead of 24 positional
    parameters.
    
    Flags:
    
      --remote-store-format=arrow-v1|arrow-v2|otlp supersedes the deprecated
      --remote-store-use-v2-schema, which still maps onto arrow-v1. All three
      ride the existing remote-store connection, so OTLP inherits its TLS,
      auth, retry and client metrics rather than opening a second one.
    
      --remote-store-compression=gzip|zstd|none applies to the OTLP export
      only; the Arrow paths compress inside their own IPC payload. gzip is the
      default because it is the only codec grpc-go ships on both sides, at
      level 1 rather than the level-6 default since the CPU is spent on a node
      being profiled. zstd is smaller and faster but some receivers reject it.
    
      --debuginfo-address points symbol upload at its own host, inheriting the
      remote store's credentials. Symbols travel the DebuginfoService protocol
      whatever encoding profiles use, so OTLP mode is not standalone.
    
      --symbolize-go enables the upstream Go interpreter, which resolves
      function, file and line from gopclntab in-agent. Previously hardcoded
      off with no way to enable it. Off by default, for backends with no
      symbolizer service.
    
    Also fixed along the way:
    
      service.name was set only from APMServiceName, so a system-wide profiler
      sent anonymous resources for nearly everything and consumers keying on
      it named every process by a hash. Now falls back to comm, then the
      executable's basename.
    
      parcaReporter.Start()'s error was discarded in main.go, and the arrow
      Start() leaked its reporting context when offline-mode MkdirAll failed.
    
      Tracer() dereferenced a nil provider despite its doc comment promising a
      no-op fallback.
    
      The vtproto codec was registered as a side effect of the remote-store
      dial. Every OTLP signal marshals through it, so it is now registered
      explicitly before any dial; otherwise an export fails at the first flush
      with an opaque marshal error.
    
    Known gaps, all recorded in the fork-strategy ADR: OTLP has no field for
    aggregation temporality, so a consumer must derive delta from the sample
    type; off-CPU is wallclock on Arrow and off_cpu on OTLP, which is still an
    open naming decision; and offline mode has no OTLP representation and is
    rejected at flag validation.
    gnurizen committed Aug 28, 2026
    Configuration menu
    Copy the full SHA
    9526c8f View commit details
    Browse the repository at this point in the history
  2. WORK IN PROGRESS

    cmd/padata-convert: arrow/OTLP round-trip and size experiment
    
    Not for main. This is the throwaway tool that answered two questions about
    the OTLP exporter, kept on a branch because rebuilding it would be wasteful
    if either question comes back.
    
    How the two encodings compare on size, over the offline-mode corpus at
    /home/tpr/parca-agent-offline (15 files, 6,713,833 samples, 146 batches):
    
      encoding          total bytes    per sample
      arrow IPC         1,283,526,472       191.2
      arrow + LZ4         537,896,704        80.1
      arrow + zstd        260,514,843        38.8
      OTLP protobuf       653,126,472        97.3
      OTLP + zstd         230,647,963        34.4
    
    Uncompressed OTLP is about half the size, because protobuf varints replace
    arrow's 64-byte buffer padding, validity bitmaps and fixed-width ListView
    offsets. Compressed the gap nearly closes, 11.5% in OTLP's favour, since
    most of arrow's bulk is the redundancy zstd eats: arrow compresses 4.93x
    against OTLP's 2.87x. On the wire the comparison is arrow+LZ4 at 80.1
    bytes/sample against OTLP+gzip at 40.2, so roughly 2x cheaper to ship.
    
    Whether the conversion is lossless, in both directions. It is:
    
      arrow -> canonical -> OTLP -> protobuf -> OTLP -> canonical -> arrow IPC
      -> canonical
    
    diffed as sorted canonical JSON at every stage, identical across all 146
    batches. The reverse encoder goes through the production SampleWriterV2, so
    it validates the format the agent actually writes.
    
    Three arrow fields have no native OTLP home and ride as attributes to make
    the round trip close: producer, temporality, and stacktrace_id. Those keys
    are an experiment artifact, not a contract; nothing should depend on them.
    Two model mismatches also surfaced: Location.MappingIndex has no presence
    flag, so table entry 0 has to be reserved as a "no mapping" sentinel, and
    fields arrow carries per sample (sample_type, period, duration) become
    per-Profile grouping keys in OTLP.
    
    reporter/arrow_v2_import.go is the only production-package addition, and it
    exists solely for this tool: the v2 location builders are unexported, so an
    importer has no way to write locations without a libpf.Frame to classify.
    It is the inverse of appendLocationV2. Delete it along with the tool if this
    branch is abandoned.
    
    Modes: inspect, sizes, compress, roundtrip, decode. The full round trip
    takes about nine minutes over the corpus, which is why it is not a test.
    gnurizen committed Aug 28, 2026
    Configuration menu
    Copy the full SHA
    8a96b34 View commit details
    Browse the repository at this point in the history
Loading