Rendered at 23:39:37 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
trashb 14 hours ago [-]
I'm missing some important parts that I would argue any (binary)format needs.
- a magic header
- a version number
- a crc or data corruption check
Additionally I would argue one would be better off writing a custom text parser instead of parsing a binary format in this case, this is similar to the debate about unixlike config vs windows regedit. I would prefer something other then json but sill readable as txt.
Even the json example provided at the bottom can be minified from 418 characters to 166 by replacing the field names with single characters and removing the spaces. Almost all of the savings in this format come from not including the the field names and having a position dependent layout. You can choose to have delimiters or arrange for a byte to indicate the type(+size) for example the following string encodes the example data almost (85bytes vs 80bytes) as efficient but is still readable and supports utf-8 interpretation.
You may optimize it further by allowing recurring entries and allowing assumed values from a defined default and only sending delta's can be dependent on the type of data you are expecting.
You should just use protobuf, far before you get to this step
ErikHuisman 14 hours ago [-]
I also want to be a real developer so i just GZIP the JSON to make it binary.
masklinn 6 hours ago [-]
If you want to be a real developer you should use raw deflate (or atleast zlib streams), they're way harder to identify so they look more really developpery.
Also gzip adds something like a dozen unnecessary bytes.
nacozarina 11 hours ago [-]
Yeah, I’m trying to imagine how this is superior to simply adding ZIP’d JSON support to your app.
entrope 2 hours ago [-]
It depends. My personal experience is with RINEX, a text file format for GNSS observation data. gzip compresses large sets of files by about 4x. Someone named Yuki HATANAKA came up with a compact text format (CRX) that shrinks it by about a factor of 4.1x. You can gzip CRX for a total ratio of almost 12x, or bzip3 it for a total ratio of 17x. But if you combine Hatanaka's technique (essentially higher-order delta coding) with some others, you get a binary format with a compression ratio of 19x that reads faster than the CRX text.
When one's uncompressed RINEX files are 200 GB/day, it's worth some effort to shrink the files, especially if it means they are faster to read.
jareklupinski 9 hours ago [-]
just Poob it
kstenerud 16 hours ago [-]
Schemas save you space, but then you lose the ability to understand the data without the schema. That's where JSON has always been handy despite its inefficiencies.
Of course you can get the same kind of thing in binary. I wrote a drop-in binary JSON replacement because it's easy to write a binary one that's twice as fast as simjson and yyjson. The important thing is to never give the drop-in replacement any extras that break roundtrip compatibility.
> Schemas save you space, but then you lose the ability to understand the data without the schema. That's where JSON has always been handy despite its inefficiencies.
Of course that’s orthogonal to using a text v binary representation, you can have a schema’d text format, and a self-describing binary format.
BSON, UBJSON, MessagePack and CBOR are binary but self-describing.
Skwid 14 hours ago [-]
One of my favourite yak shaving adventures in a previous job was writing a parse in place UBJSON decoder for ~1MB of data on a device with about as much free memory. Fast (enough) access by key, binary search with a few shortcuts for the mostly numeric payload.
It built an index of maybe 30 bytes to speed things along, but also to let it keep working on the tail of the old message whilst the new one was overwriting it's start.
Should the system design have required all this of a device with 4MB of memory? Probably not. But it worked, and I had a great time
deepsun 18 hours ago [-]
Cool. And the next enlightening step would be to generate custom parser code from message schemas. So that when de-serializer reads a next field, it does not read what name and type it will be, it expects and fails if it's not. Safer and more compiler friendly.
MiroslavPokorny 18 hours ago [-]
Congratulations, your strings dont support Unicode and you just ignored probably more than 3/4 or more of the world's languages.
zimpenfish 14 hours ago [-]
I'm probably missing something obvious but ... isn't this just (the moral equivalent of) ASN.1 (specifically BER[0])?
There already exists a "binary analogy" of JSON called Protocol Buffers or Kiwi. They are pretty simple (a parser can be made in 2 kB of code). Wouold be great if the author compared his reinvented wheel with existing wheels :D
wojciii 9 hours ago [-]
TIL about Kiwi. Thanks.
I find these blogs which show how the author finds out about someting a bit boring, as they often show the way to something obvious and making it seem like its something profound.
masklinn 6 hours ago [-]
protobuf is definitely not a binary analogue to JSON, because it's not a self-describing document format, it requires a schema.
Even the json example provided at the bottom can be minified from 418 characters to 166 by replacing the field names with single characters and removing the spaces. Almost all of the savings in this format come from not including the the field names and having a position dependent layout. You can choose to have delimiters or arrange for a byte to indicate the type(+size) for example the following string encodes the example data almost (85bytes vs 80bytes) as efficient but is still readable and supports utf-8 interpretation.
You may optimize it further by allowing recurring entries and allowing assumed values from a defined default and only sending delta's can be dependent on the type of data you are expecting. could become:You should just use protobuf, far before you get to this step
Also gzip adds something like a dozen unnecessary bytes.
When one's uncompressed RINEX files are 200 GB/day, it's worth some effort to shrink the files, especially if it means they are faster to read.
Of course you can get the same kind of thing in binary. I wrote a drop-in binary JSON replacement because it's easy to write a binary one that's twice as fast as simjson and yyjson. The important thing is to never give the drop-in replacement any extras that break roundtrip compatibility.
Was a fun little project to write, and quite useful for me: https://github.com/kstenerud/bonjson
Of course that’s orthogonal to using a text v binary representation, you can have a schema’d text format, and a self-describing binary format.
BSON, UBJSON, MessagePack and CBOR are binary but self-describing.
Should the system design have required all this of a device with 4MB of memory? Probably not. But it worked, and I had a great time
[0] https://en.wikipedia.org/wiki/X.690#BER_encoding
I find these blogs which show how the author finds out about someting a bit boring, as they often show the way to something obvious and making it seem like its something profound.
cbor is a json analogue.