Skip to article frontmatterSkip to article content
Site not loading correctly?

This may be due to an incorrect BASE_URL configuration. See the MyST Documentation for reference.

Iterator vs Generator

Generator is a special case of Iterator. Generator is easy and convenient to use but at additional cost (memory and speed). If you need performance, use plain iterator (with the help of the itertools module). If you need convenience and concise code, use generator.

Please refer to Python Generator vs Iterator for more detailed discussions.

When Not to Use Generator

  1. You need to access the data multiple times (i.e. cache the results instead of recomputing them).

     for i in outer:           # used once, okay to be a generator or return a list
         for j in inner:       # used multiple times, reusing a list is better
             ...
  2. You need random access (or any access other than forward sequential order).

     for i in reversed(data): ...     # generators aren't reversible
    
     s[i], s[j] = s[j], s[i]          # generators aren't indexable
  3. You need to join strings (which requires two passes over the data).

     s = ''.join(data)                # lists are faster than generators in this use case
  4. You are using PyPy which sometimes can’t optimize generator code as much as it can with normal function calls and list manipulations.

builtins.enumerate

0: how
1: are
2: you
10: how
11: are
12: you

builtins.all

True
False

builtins.any

True
True
False

builtins.max

3
---------------------------------------------------------------------------
ValueError                                Traceback (most recent call last)
<ipython-input-2-a48d8f8c12de> in <module>
----> 1 max([])

ValueError: max() arg is an empty sequence
0

itertools.accumulate

[1, 2, 6, 24]
[1, 3, 6, 10]

count (range)

Notice that count does NOT count the number of elements in an iterator, but instead generate an iterator of arithematic sequence of integers.

[0, 1, 2, 3, 4, 5, 6, 7, 8, 9]

cycle

['a', 'b', 'c', 'd', 'a', 'b', 'c', 'd', 'a', 'b']

itertools.chain

You can use itertools.chain to combine/join multiple iterable objects.

['a', 'b', 'c', 'd', 'e', 'f', 'g', 'h', 'i', 'j', 'k', 'l']
<itertools.chain at 0x7f713aeeb358>
[1, 2, 3, 4]
<itertools.chain at 0x7f713a4f8438>
[1, 2, 3, 4]

compress

[0, 1, 2, 5, 6, 9]

dropwhile

[5, 6, 7, 8, 9]
[5, 4, 3, 2, 1]
[1, 2, 3, 4, 5, 6, 7, 8, 9]

combinations

[(0, 1, 2), (0, 1, 3), (0, 1, 4), (0, 2, 3), (0, 2, 4), (0, 3, 4), (1, 2, 3), (1, 2, 4), (1, 3, 4), (2, 3, 4)]

combinations_with_replacement

[(0, 0, 0), (0, 0, 1), (0, 0, 2), (0, 0, 3), (0, 1, 1), (0, 1, 2), (0, 1, 3), (0, 2, 2), (0, 2, 3), (0, 3, 3), (1, 1, 1), (1, 1, 2), (1, 1, 3), (1, 2, 2), (1, 2, 3), (1, 3, 3), (2, 2, 2), (2, 2, 3), (2, 3, 3), (3, 3, 3)]

groupby

itertools.groupby is a convenient way to group elements into groups of elements. Its return type is Iterable[Tuple[KeyType, Iterable[ValueType]]]. Notice that itertools.groupby iterate through the original collection only once, so you must consume the result of itertools.groupby all at one time (convert all iterators at one time).

<itertools.groupby at 0x12621f2c0>

The following way of consuming the resutl of itertools.groupby is not right as it does not convert all iterators at one time.

[(0, <itertools._grouper at 0x10d67f550>), (1, <itertools._grouper at 0x10d67fd90>), (2, <itertools._grouper at 0x10d67f700>), (3, <itertools._grouper at 0x10d67f250>)]
(0, <itertools._grouper at 0x10d67f550>)

Got no values for the key 0 as the iterator as already been consumed the first time it.groupby(x, lambda e: e // 3) is converted to a list.

[]

Below is the right way to convert the result of itertools.groupby to a list of tuples.

[(0, [0, 1, 2]), (1, [3, 4, 5]), (2, [6, 7, 8]), (3, [9])]

Below is the right way to convert the result of itertools.groupby to a dict.

{0: [0, 1, 2], 1: [3, 4, 5], 2: [6, 7, 8], 3: [9]}

Below is a the right way to extract groups of values (without keys).

[[0, 1, 2], [3, 4, 5], [6, 7, 8], [9]]

Notice that you can also use itertools.zip_longest to group values into equal size of groups according to the iteration order of elements.

[['a', 'b', 'c'], ['d', 'e', 'f'], ['g', 'h', 'i'], ['j', None, None]]

filter

[0, 2, 4, 6, 8]

filterfalse

[1, 3, 5, 7, 9]

from_iterable

['a', 'b', 'c', 'd', 'e', 'f', 'g', 'h', 'i', 'j', 'k', 'l']

islice

['a', 'b']
[]
['c', 'd']

next

next (when used with a conditon) is equivalent to Seq.find in Scala. A default value can be passed to next.

1
144
144

slice

slice(None, 10, None)
[1, 2]
[1, 2]

map

[0, 1, 4, 9, 16, 25, 36, 49, 64, 81]

permutations

[(0, 1), (0, 2), (0, 3), (0, 4), (1, 0), (1, 2), (1, 3), (1, 4), (2, 0), (2, 1), (2, 3), (2, 4), (3, 0), (3, 1), (3, 2), (3, 4), (4, 0), (4, 1), (4, 2), (4, 3)]

product

[('a', '1'), ('a', '2'), ('a', '3'), ('a', '4'), ('b', '1'), ('b', '2'), ('b', '3'), ('b', '4'), ('c', '1'), ('c', '2'), ('c', '3'), ('c', '4'), ('d', '1'), ('d', '2'), ('d', '3'), ('d', '4')]

repeat

['abcd', 'abcd', 'abcd', 'abcd', 'abcd', 'abcd', 'abcd', 'abcd', 'abcd', 'abcd']

starmap

[0, 8, 1024]

takewhile

[0, 1, 2, 3, 4]
[0, 1, 2, 3, 4]
[0]

tee

['a', 'b', 'c', 'd']
['a', 'b', 'c', 'd']

zip_longest

[('a', 1), ('b', 2), ('c', 3), (None, 4), (None, 5)]

Empty Iterator or Not

False
True

Iterator Receipes