Floating-point arithmetic


Series Overview

This article is part of the series. Below are links to all posts in the series:
  1. Floating-Point Arithmetic
  2. Decimal class in Python
  3. Fraction class in Python
  4. Rounding Errors in Java

What's inside this article ⌄
  • Floating point arithmetic explained
  • Float vs double precision differences
  • How to handle floating point errors in programming
  • Binary exponential notation conversion
  • Java rounding errors
  • Java float arithmetic
  • Python rounding errors
  • Python float arithmetic

Suppose you want the computer to memorize a really large or tiny number like this:

4_000_000_000_000_000

What if I’d say to you that we can leverage the exponential notation?

Looking at the number, we can represent it as 4 × 10^15. Now, you have to memorize only two small numbers: 4 and 15!

You know that computers operate in a binary numeral system. 4_000_000_000_000_000 turned out to be 100000000000000000000000000000000000000000000000 in binary. But utilizing exponential notation with base 2, it’s just 1 × 2^47.

Just to remind, that in m * b^r:

  • m is called significand/mantissa
  • b is the base
  • r is the exponent.

This type of thing, when you store mantissa and exponent only, is called floating-point arithmetic.

We have two floating-point datatypes. They are float (single-precision) and double (double-precision). Let’s look at them closer.

Single-precision (float, 32 bits):

  • Sign bit: 1 bit (indicates positive or negative)
  • Exponent: 8 bits
  • Mantissa: 23 bits

Double-precision (double, 64 bits):

  • Sign bit: 1 bit (indicates positive or negative)
  • Exponent: 11 bits
  • Mantissa: 52 bits

It’s worth mentioning that they use exponential notation, where base equals 2.

Now, let’s move on to the real numbers. In binary, they look like this:

20.625 (decimal) = 10100.101 (binary)

What we need to do now is represent it in exponential notation with the base of 2:

10100.101 (binary) = 1.0100101 × 2^4

Recognize mantissa and exponent? Here we go!

Depending on the data type you’ve chosen (float or double), you have different amounts of space allocated to store each part of the 1.0100101 × 2^4.

By the way, that is the reason for floating-point round errors when dealing with real numbers:

  • Some of the real numbers just can’t be represented in binary format accurately due to their nature
  • Some of them are big enough to fit in the space reserved for mantissa.

Therefore, in those cases, you get a number close to what you had initially. It’s the same as you trying to divide 1 by 3 and get 0.3333 (repeatable decimal). When using float and double types, you can face an inaccuracy during computations, e.g.:

a: float = 0.1 + 0.2
b: float = 0.3
print(a == b) # Output: false

Read next articles in the series if you’d like to know how to bypass this thing.

Moreover, the use of exponential notation allows float and double types to store tremendous numbers:

  • up to ~3.4 × 10^38 for float
  • up to ~1.7 × 10^308 for double

Whereas a long data type (that is intended to be used just for integers, does not use an exponential notation, and maintains 100% accuracy) can store approximately 9.2 × 10^18 as a maximum value.

Hell yeah!