I recently got into Rust out of nowhere because its little crab mascot is adorable, so I started reading The Rust PL to see what all the fuss is about. Turns out Rust is a genuinely sane, pleasant language: its functional programming features and its build system make it a real joy to use.

I’ve only gotten through two ten chapters so far, but I can already smell that rich Rust aroma or is it just rust? wafting off the code.

Basically finished it now.

On a fairly lazy weekend morning, I decided to knock out the last chapter.

Created: 2021-01-09 21:10:07

Prologue

This post is mostly written for myself, to make up for my pathetic memory. Most of it is notes and code from reading The Rust PL, with a few of my own thoughts sprinkled in. Normal people can read this section and stop here.

Why I’m Learning Rust

After using it for a while, here’s what I’ve figured out about Rust-chan’s personality:

  • She tells you exactly what’s wrong
    • Rust isn’t like a girlfriend who picks fights for no reason, the kind who thinks you did something wrong but won’t tell you, so you’re stuck asking, “Did I mess up here? Did I do something wrong over there?”;
    • With a girlfriend like C, even when you’re sure you did nothing wrong, she won’t say you did, and yet somehow nothing ever works;
    • What makes Rust different is that she tells you exactly what you did to upset her though she gets upset a lot. Like, a LOT;
    • And even when Rust-chan is annoyed with you, she won’t leave you hanging. She sweetly offers all kinds of help and sticks with you until the problem’s solved though honestly, Google solves it more often than she does.
  • A whole bunch of awesome features
    • Closures, although there seem to be some restrictions on the lifetimes of the arguments used to create a closure
    • match: makes sure you never miss a single corner case?
  • The syntax is fairly flexible, but the checks are insanely strict

Toolchain, Builder and Cargo

Installing Rust is smooth and painless. The official site gives you a handy little script that installs the whole toolchain, including the compiler rustc, cargo, and so on. Dependency management in the style of cargo and python-pip is just so civilized; I wonder who first came up with that stroke of genius. For a language that can do systems programming, having Cargo basically lets Rust run circles around the old-timers of the field. On top of that, you get what feels like black-magic web-based documentation: just run cargo doc --open and your browser pops up with the docs for every dependency of the project (at the versions pinned in the toml). Seriously convenient.

When you hit some baffling, bizarre compile error, it gives you an error code, like E0061. Run rustc --explain E0061 and you get a deeper explanation of what the heck that thing actually is. The Rust devs are so worried you won’t figure it out that they’ve stopped just short of opening a portal right in front of your face to teach you Rust in person. That kind of thoughtful, caring treatment of users really makes you feel warm and fuzzy inside.

Style

Here’s a random chunk of code from Chapter 2 so you can get a feel for it!

use rand::Rng;
use std::cmp::Ordering;
use std::io;

fn main() {
    println!("Guess the number!");

    // This uses the newer API of the rand crate; in the original book, `gen_range` takes two arguments.
    // rand = "0.8.1"
    let secret_number = rand::thread_rng().gen_range(1..101);
    println!("The secret number is: {}", secret_number);

    loop {
        println!("Please input your guess");
        let mut guess = String::new();
        io::stdin()
            .read_line(&mut guess)
            .expect("Failed to read line");

        let guess: u32 = match guess.trim().parse() {
            Ok(num) => num,
            Err(_) => continue,
        };
        // expect("Please type a number");
        println!("You guessed: {}", guess);
        match guess.cmp(&secret_number) {
            Ordering::Less => println!("Too small!"),
            Ordering::Greater => println!("Too big!"),
            Ordering::Equal => {
                println!("You win!");
                break;
            }
        }
    }
}

A quick rundown

  • mut distinguishes mutable variables from immutable ones
  • A variable name can be shadowed: the new variable bound to that name can have a different type and value
  • match, that darling of functional programming, really seems to let you pull off all kinds of fancy tricks, and there are functional programming vibes all over the place
  • Function calls return Ok or Err, which feels like a bit of a hassle, but it’s definitely safer
  • There are quite a lot of ::s
  • Rust is a statically typed lang, but you can tell its type inferencer is pretty powerful

Rust’s Ground Rules

Copy

To prevent errors like double frees, for any variable whose size is unknown at compile time (as well as any variable that does not implement the Copy trait), Rust only does a shallow copy on an assignment like this (the book calls it a move). The consequence is that s can’t be used normally anywhere afterwards. To avoid the performance cost of a runtime GC, Rust completely frees a variable (clears it off the heap) as soon as it goes out of scope (roughly speaking, the region wrapped in curly braces). So if both s and s2 were usable but pointed to the same heap address, you’d get a double free when leaving the scope, which is a safety hazard. If you want a deep copy, use the clone method.

let s = SomeType;
let s2 = s;

Ownership

This oddball of a language also transfers ownership when you pass arguments. Ownership, simply put, is which function a variable belongs to. Whether the transfer is a deep copy follows the same rules as assignment in the previous subsection. This leads to some weird issues:

fn main() {
    let s = String::from("hello");
    takes_ownership(s);

    let x = 5;
    makes_copy(x);
    //To make this compile, the print below has to go
    //otherwise you get error: value borrowed here after move
    // println!("{}", s);

    println!("{}", x);

    let s = String::from("hello2");
    //
    let s = takes_and_gives_back(s);
    println!("{}", s);
}

fn takes_ownership(some_str: String) {
    println!("{}", some_str);
}

fn makes_copy(some_int: i32) {
    println!("{}", some_int);
}

fn takes_and_gives_back(some_str: String) -> String {
    println!("{}", some_str);
    //return
    some_str
}

Since s also hands over its ownership when it’s passed to takes_ownership, the ownership is gone once the function returns, and the String’s memory gets freed. In contrast, x can still be used after calling a function that takes it as an argument.

However, the following code produces the correct output:

fn main() {
    let s = "hello";

    takes_ownership(s);

    println!("{}", s);
}

fn takes_ownership(some_str: &str) {
    println!("{}", some_str);
}

That’s because s here isn’t a String but a str, i.e. a string literal that’s hardcoded into the text.

Call by reference

fn main() {
    let s1 = String::from("hello");

    let len = calculate_len(&s1);

    println!("{}, {}", s1, len);
}

fn calculate_len (s: &String) -> usize {
    s.len()
}

Unlike C, Rust doesn’t seem to require * every single time you dereference. This question on Stack Overflow partially cleared up my confusion: when calling a method, Rust automatically adds &, &mut, or * to match the method’s signature.

Multiple borrows

How do you pass an argument without giving up ownership? Pass a reference to the function. In the code above, s in the function calculate_len stores a pointer to s1. In Rust, passing arguments this way is called borrowing.

Arguments passed this way have one drawback, though: the data is immutable. You can use another kind of ref:

fn main() {
    let mut s1 = String::from("hello");

    let len = calculate_len(&mut s1);

    println!("{}, {}", s1, len);
}

fn calculate_len (s: &mut String) -> usize {
    s.push_str(" world");
    s.len()
}

However, two mutable borrows are allowed if they happen in different scopes

let mut s = String::from("hello");

{
    let r1 = &mut s;
}
let r2 = &mut s;

To prevent data races, Rust forbids creating two mutable borrows of the same thing. You can have as many immutable borrows as you like (and if a mutable reference shows up too, it still compiles as long as every use of the immutable references in their scope comes before that mutable reference). It feels like these rules are mostly there to head off bugs that can crop up under concurrency.

Slice

Just like in Python, Strings in Rust can be sliced. Slices exist to avoid the awkward situation where an index is still hanging around after the variable has been dropped.


let s = String::from("hello");
let s1 = &s[0..2];
// let s1 = &s[..2];

Here’s the earlier function reimplemented with slices. The slice is returned as a &str.

fn main() {
    let s1 = String::from("hel lo");

    let len = first_space(& s1);

    //s1.clear();
    //error because clear() will need a mutable reference to truncate the String

    println!("{}, {}", s1, len);
}

fn first_space(s: &String) -> &str {
    let bytes = s.as_bytes();
    for (i, &item) in bytes.iter().enumerate() {
        if item == b' ' {
            return &s[0..i];
        }
    }    
    
    &s
}

A function with a prototype like fn first_space(s: &str) -> &str can take a String as its argument too.

The same slice mechanism works for other kinds of arrays as well.

Struct

struct User {
    username: String,
    email: String,
    active:bool,
    sign_in_count: u64,
}

fn main() {
    let user1 = User {
        email: String::from("eg@eg.com"),
        username: String::from("user1"),
        active: false,
        sign_in_count: 1
    };
}
  • Structs can be defined and instantiated like this. Rust also provides some extra syntactic sugar to shorten some of the fields:

    fn init_user(email: String, username: String) -> User {
        User {
            //email: email,
            email,
            //username: username,
            username,
            active: false,
            sign_in_count: 1,
        }
    }
    
  • Rust also lets you create one struct from another:

        let user2 = User {
            username: String::from("user2"),
            ..user1
        };
    
  • Does the String get deep copied?

  • Tuples can also be used as Structs:

    struct Color(i32, i32, i32);
    let blk = Color(0, 0, 0);
    
  • The instantiated variable owns the Struct’s data fields.

  • A Struct can hold references to other data, but that requires specifying a lifetime.

  • You can have structs that contain no data at all (unit-like Structs)

Debug output

#[derive(Debug)]
struct Rectangle {
    width: u32,
    height: u32,
}

fn main() {
    let rec1 = Rectangle {
        width: 30,
        height: 50,
    };

    println!("rec1 is {:#?}", rec1);
}

Formatted output for Structs can be done with the help of #[derive(Debug)]. This usage is a bit like decorators in Python. In official terms, we’re deriving a Debug trait here. What exactly that means, I’ll have to read further to find out.

Method

#[derive(Debug)]
struct Rectangle {
    width: u32,
    height: u32,
}

impl Rectangle {
    fn area(&self) -> u32 {
        self.width * self.height
    }

    fn can_hold(&self, sub_rec: &Rectangle) -> bool {
        self.width > sub_rec.width && self.height > sub_rec.height
    }

    fn new_square(edge: u32) -> Rectangle {
        Rectangle {
            width: edge,
            height: edge,
        }
    }
}

fn main() {
    let rec1 = Rectangle {
        width: 30,
        height: 50,
    };
    let rec2 = Rectangle {
        width: 20,
        height: 45,
    };

    let square = Rectangle::new_square(30);

    println!("rec1 is {:#?}, area is {}", rec1, rec1.area());
    println!("rec1 can hold re2? {}", rec1.can_hold(&rec2));
}

You can implement an area method. Like methods in Python, the first parameter of a Struct’s method here is &self. Since area is implemented inside impl Rectangle {}, the type of &self doesn’t need to be specified. Methods can take ownership here as well.
Unlike C/C++, Rust doesn’t use -> to call methods or access members on an object (when it’s a pointer). Rust automatically adds &, &mut, or * to match the function’s signature. Personally I think this is super convenient, and it looks way cleaner than C. This feature is called automatic referencing and dereferencing.
Though this means you’d have to pay extra attention to what type each parameter actually is when defining functions. Maybe it only gets added automatically for the Object (a.k.a. self)? It doesn’t seem to happen for the other parameters.

Associated Functions

Methods that don’t take self as a parameter. If I remember correctly, these are called class methods in Python? In Rust they’re called associated functions. They’re usually used to return a new object. See the code above for the syntax.

Rust also supports multiple impl blocks for the same type, which seems to be designed with traits and generics in mind.

Enum

Enums in Rust can be defined like this:

enum Message {
    // unit
    Quit,
    // anonymous nested struct
    Move {x: i32, y: i32},
    // String included
    Write(String),
    // 3 * i32 included
    ChangeColor(i32, i32, i32),
}

impl Message {
    fn call(&self) {
        //
    }
}

As you can see, different variants of an enum can have different nested data types. That’s a really handy feature.

Option

To avoid the problems NULL causes in C/C++, like null pointer dereferences, Rust introduces the Option enum.

enum Option<T> {
    Some(T),
    None,
}

This is somewhat like the idea of a monad in functional programming. It lets you have fields that can be empty.

So how do you get the stuff out of an Option? That’s where the mighty match comes in

fn main() {
    let x = Some(3);
    let x = plus_one(x);
    println!("x is {:?}", x);

    let y = None;
    let y = plus_one(y);
    println!("y is {:?}", y);
}

fn plus_one(num: Option<i32>) -> Option<i32> {
    match num {
        Some(v) => Some(v+1),
        None => None,
    }
}

PS: inside a match, _ matches any value. You can also use if let to match a single case, in which case Rust won’t check whether any other cases went unmatched.

if let Some(3) = num {
    println!("Lucky!");
}

Rust Packages

  • Create a Rust lib with cargo new --lib [libname]
  • In src/lib.rs you can create multiple modules with the mod keyword
  • Modules can be nested and can contain enums, structs, and functions, all of which can be public
  • All fn, mod, enum, and struct items are private by default; a parent can’t directly access its children’s items, but the reverse works
  • You can use relative or absolute “paths” to access things inside a module
  • Code inside an impl can access the private fields of the type being impl’d
  • use is basically Python’s import, and you can also do use A as a
  • pub use can be used to lift things into a higher namespace? ref
  • There’s also use [namespace]::*, similar to Python
  • nested path Rust provides some syntactic sugar:
    // original
    use std::cmp::Ordering;
    use std::io;
    use std::io::Write;
    
    // nested
    use std::{cmp::Ordering, io};
    use std::io::{self, Write};
    

Splitting into Separate Files

Rust nests modules by putting a module’s submodules in other files/folders under the root. A Rust file’s name is the name of the module it belongs to.

Here’s the module structure all in a single file:

//#[cfg(test)]
mod back_of_house {

    // #[derive(Debug)]
    pub struct Breakfast {
        pub toast: String,
        seasonal_fruit: String,
    }

    impl Breakfast {
        pub fn summer(toast: &str) -> Breakfast {
            Breakfast {
                toast: String::from(toast),
                seasonal_fruit: String::from("peaches"),
            }
        }
    }
}
mod front_of_house {
    pub mod hosting {
        pub fn add_to_waitlist() {}

        fn seat_at_table() {}

    }

    mod serving {
        fn take_order() {}

        fn serve_order() {}

        fn take_payment() {}
    }
}

use crate::front_of_house::hosting;
// or use relative path
// use self::front_of_house::hosting;

pub fn eat_at_restaurant() {

    // abosulute path
    crate::front_of_house::hosting::add_to_waitlist();
    // relative path
    front_of_house::hosting::add_to_waitlist();
    // effective after use
    hosting::add_to_waitlist();

    let mut meal = back_of_house::Breakfast::summer("Rye");
    meal.toast = String::from("Wheat");
    println!("{} toast pls!", meal.toast);
}

Here, front_of_house can be refactored like this:

  • src/lib.rs root
mod front_of_house;
  • src/front_of_house.rs
pub mod hosting {
    pub fn add_to_waitlist() {}
}

Or you can split hosting out even further:

  • src/front_of_house.rs
pub mod hosting;
  • src/front_of_house/hosting.rs
pub fn add_to_waitlist() {}

Rust Collections

Rust provides a bunch of handy data structures.

Vector

A Vector here is basically a variable-length array, and its definition is pretty similar to vectors in the Pie language.
It can be initialized with the vec! macro or Vec::new().
Because of how vectors are implemented internally, a mutable vector may get moved somewhere else in memory when its length changes (once the vector grows, the originally malloc’d memory might not be big enough). So vectors also follow the rule that you can’t have mutable and immutable references in the same scope, even if the immutable reference points at the first few elements.

You can also get elements out of a vector with a reference or the get() method. The two are handled slightly differently:

fn main() {
    let mut v = Vec::new();

    v.push(5);
    v.push(6);

    // panic if out of bound
    let third = &v[2];

    // will get an Option<T>
    match v.get(1) {
        Some(val) => println!("get {}", val),
        None => println!("not even get!"),
    }
}

Iterate a vector

With Python-like syntax and C-like semantics, you can iterate over a vector with pointers:

    for i in &mut v {
        *i += 10;
    }

But using the iterator like this throws an error:

    for mut i in v {
        i += 10;
    }

    for j in v {
        println!("{}", j);
    }

The error:

error[E0382]: use of moved value: `v`
   --> src/main.rs:12:14
    |
2   |     let mut v = Vec::new();
    |         ----- move occurs because `v` has type `Vec<i32>`, which does not implement the `Copy` trait
...
8   |     for mut i in v {
    |                  -
    |                  |
    |                  `v` moved due to this implicit call to `.into_iter()`
    |                  help: consider borrowing to avoid moving into the for loop: `&v`
...
12  |     for j in v {
    |              ^ value used here after move
    |
note: this function consumes the receiver `self` by taking ownership of it, which moves `v`

That’s because handing v to the iterator transfers ownership, so v is invalid for the rest of the scope. Looks like iterators should borrow unless you really need them not to.

If you want to store elements of different types in the same vector, you can store an enum in it, with the enum’s variants holding the different types.

String

Like the APIs we’ve seen in earlier code, you can create a String with methods like ::new(), ::from(), and .to_string().
The String API also includes push_str(), push(), and an overloaded + operator. Note that these operations don’t take ownership. There’s also the format!() macro:

let s1 = String::from("tic");

let s = format!("{}-tok", s1);

s is a String too.

In Rust, a String is really a Vec<u8>. But since it supports UTF-8, Rust can’t index into a String directly the way C can. You can use slicing to grab sub-elements instead, but you’ll get an error if the slice lands in the middle of a char boundary.

You can use .chars() or .bytes() to produce an iterable: the former splits the string into individual Unicode chars, while the latter gives you u8s.

Hashmap

Look at the code below: unlike vectors, HashMap doesn’t have a macro for initialization, but it can be built from two vectors.

use std::collections::HashMap;

fn main() {
    let mut scoreboard = HashMap::new();

    scoreboard.insert(String::from("A"), 20);
    scoreboard.insert(String::from("B"), 30);

    println!("{:?}", scoreboard);

    let teams = vec![String::from("A"), String::from("B")];
    let scores = vec![20, 30];

    // must specify the type here
    let mut scoreboard: HashMap<_, _> = teams.into_iter().zip(scores.into_iter()).collect();
    println!("{:?}", scoreboard);

    // get an Option<T> here
    let team = String::from("A");
    let s = scoreboard.get(&team);
    // euivalent to
    let s = scoreboard.get("A");
    println!("{:?}", s);

    for (k, v) in &scoreboard {
        println!("{}: {}", k, v);
    }

    // insert if the Key has no value
    scoreboard.entry(String::from("A")).or_insert(50);
    scoreboard.entry(String::from("C")).or_insert(50);

    // update the data
    let score = scoreboard.entry(String::from("A")).or_insert(0);
    *score += 100;
    println!("{:?}", scoreboard);
}

Note that HashMap keys can’t be duplicated; inserting an existing key updates the value for that key. So you can use entry().or_insert to check whether a key exists and insert only if it doesn’t.
entry() is a pretty magical thing: it returns an Entry. If the key exists it returns a mutable reference, and if not, it creates the entry and returns a mutable reference to it.
You can also use entry() to get a mutable reference to a slot in the HashMap and then update the data in it.

Error

Unlike other languages, Rust has no exception handling mechanism. It splits errors into recoverable and unrecoverable ones. The former are handled with Result<T, E>, while the latter happen when the panic!() macro gets called.
You can call this macro yourself to make the program panic, and you can set RUST_BACKTRACE=1 at runtime to get a backtrace when the program panics.

Result is defined as follows:

enum Result<T, E> {
    Ok(T),
    Err(E),
}

Handling a result with match is the standard Rust move. Any fancier tricks? Sure there are: unwrap_or_else(), but I’ll keep you in suspense on that one for now. If you can’t be bothered with match, you can just unwrap(). If there’s no error, it returns whatever’s wrapped inside Ok; otherwise it panics. You can also use expect() and pass it a description for a friendlier error message; apart from taking an argument, expect() is exactly the same as unwrap().

use std::fs::File;
use std::io::ErrorKind;

fn main() {
    let f = File::open("./src/hello.txt");
    let f = match f {
        Ok(file) => file,
        Err(err) => match err.kind() {
            ErrorKind::NotFound => match File::create("hello.txt") {
                Ok(f) => f,
                Err(err) => panic!("create file failed {:?}", err),
            },
            other => {
                panic!("unexpected error: {:?}", other)
            }
        },
    };
    println!("open file successfully! {:?}", f);
}

Propagating Errors

use std::fs::File;
use std::io::{self, Read};

fn read_username_from_file() -> Result<String, io::Error> {
    let f = File::open("hello.txt");

    let mut f = match f {
        Ok(file) => file,
        Err(e) => return Err(e),
    };

    let mut s = String::new();

    match f.read_to_string(&mut s) {
        Ok(_) => Ok(s),
        Err(e) => Err(e),
    }
}

This simple code returns a Result. And Rust provides an even simpler operator, ????. With this operator, the function can be written as:

    let mut f = File::open("hello.txt")?;

    let mut s = String::new();

    f.read_to_string(&mut s)?;

    Ok(s)

What ? does is almost exactly the same as the match we wrote earlier, except that it calls the From trait from the standard library to convert the error into the error type in the function’s declared return type. (Does match not do a type conversion on return?) But there’s one more change here: f is now declared as mutable. This completely baffled me. Why did the immutable f in the earlier code work just fine???? read_to_string is defined as fn read_to_string(&mut self, buf: &mut String) -> Result<usize>, which clearly needs a mutable self. But the damn thing actually runs??? I was totally stumped as to why.

Mystery solved: the first f got shadowed… I didn’t notice f was defined twice!!!

But there’s an even simpler way to do the task above:

    fs::read_to_string("hello.txt")

This presumably uses something like a class method, so we can call it without creating a File first. But its signature looks like this: pub fn read_to_string<P: AsRef<Path>>(path: P) -> io::Result<String>, so the return value isn’t the Result<T, E> we saw before. For some reason I can’t jump to its definition directly, but my guess is it wraps the latter.

Generic Types, Traits, and Lifetimes

Generics and related stuff. This chapter explains what code like this is actually saying:

use std::fmt::Display;

fn longest_with<'a, T>(x: &'a str, y: &'a str, ann: T) -> &'a str
where
    T: Display,
{
    println!("announcement {}", ann);
    if x.len() > y.len() {
        x
    } else {
        y
    }
}

Generic Types

Struct

struct Point<T> {
    x: T,
    y: T,
}

impl<T> Point<T> {
    fn x(&self) -> &T {
        &self.x
    }
}

impl Point<f32> {
    fn distance(&self) -> f32 {
        (self.x.powi(2) + self.y.powi(2)).sqrt()
    }
}

struct PowerfulPoint<T, U> {
    x: T,
    y: U,
}

impl<T, U> PowerfulPoint<T, U> {
    fn mixup<V, W>(self, other: PowerfulPoint<V, W>) -> PowerfulPoint<T, W> {
        PowerfulPoint {
            x: self.x,
            y: other.y,
        }
    }
}

This implements generics in a struct. Generics in enums work similarly. A few things worth pointing out:

  • To implement functions for a generic struct, you need to put the generic parameter <T> after impl
  • You can also implement functions for the struct non-generically, i.e. for one specific type only
  • The implemented functions can have generic parameters that differ from the generic struct’s, like in the PowerfulPoint implementation
  • Note that x() returns a reference here. If it didn’t, you’d get an error:
    • cannot move out of self.x which is behind a shared reference
    • There should actually be a way around this if T implements the Copy trait. But Rust has no idea whether it does, so for now we’ll just return a reference

Function

Let’s make a generic version of this function:

fn largest(list: &[i32]) -> &i32 {
    let mut largest = &list[0];

    for n in list {
        if n > largest {
            largest = n;
        }
    }
    largest
}

But just slapping on a generic parameter doesn’t make it compile:

fn largest<T>(list: &[T]) -> &T {
    let mut largest = &list[0];

    for n in list {
        if n > largest {
            largest = n;
        }
    }
    largest
}

That’s because not every type can be compared. This T has to implement a partial-ordering trait.
Rewriting it like this makes it compile:

fn largest<T: PartialOrd + Copy>(list: &[T]) -> T {
    let mut largest = list[0];

    for &n in list {
        if n > largest {
            largest = n;
        }
    }
    largest
}

The syntax details here are discussed further below.

Compilation

Rust implements generics by compiling out a separate copy for every concrete type that’s actually used, so there’s very little performance overhead. For example, if foo<T>() gets called with an i32 and a u8, it’s turned into two functions at compile time, foo_i32() and foo_u8(), which cuts down the overhead that generics would otherwise cause.

Trait

Simply put, a trait tells the compiler that certain types share some common functionality. Rust says it’s pretty similar to interfaces.

Define a trait

Here we define a trait called Summary. Note that the function signature ends with ;. A trait can have multiple functions, and any type that wants to implement the trait has to provide its own implementation of every function in it.

// src/lib.rs
pub trait Summary {
    fn summarize(&self) -> String;
}

pub struct News {
    pub headline: String,
}

impl Summary for News {
    fn summarize(&self) -> String {
        format!("{}", self.headline)
    }
}

pub struct Tweet {
    pub username: String,
}

impl Summary for Tweet {
    fn summarize(&self) -> String {
        format!("{}", self.username)
    }
}

// src/main.rs
mod lib;
use lib::Summary;

fn main() {
    println!("aaa");
    let n = lib::News{
        headline: String::from("headline_A"),
    };

    println!("{}", n.summarize());
}

Rust traits follow the coherence property (orphan rule): a trait implementation has to live either with the trait or with the type (i.e. the crate of either the type or the trait must be local). Conversely, say that somewhere random (e.g. your own lib) you make Vec<T> implement Display: both of those come from the stdlib and are extern to your lib, so that’s not allowed. This keeps a trait from being implemented twice. There’s also a security angle: if these extern traits and types could be re-implemented, any crate used in the project could hijack those implementations.

To use the type and trait we just wrote, go over to main.rs, first mod lib, then use lib::Summary. Note that lib here can actually be any module name (file name), as long as it matches the name of an rs source file in the src directory. And if you want to call summarize(), the function from the Summary trait, you have to use the trait.

Just like you can with interfaces, you can give the trait a default implementation:

pub trait Summary {
    fn summarize(&self) -> String {
        String::from("Unimplemented Summmary")
    }
}

pub struct News {
    pub headline: String,
}

impl Summary for News {
}

Now calling news.summarize() goes straight to the implementation in the trait definition. But the call only works if you impl Summary for News, even if the impl block is empty. This is also a bit like overriding (virtual) functions in a class.

mod lib;
use lib::Summary;

// also works here if fn notify(item: &Summary)
fn notify(item: &impl Summary) {
    println!("Notification: {}", item.summarize());
}

fn main() {
    println!("aaa");
    let n = lib::News{
        headline: String::from("headline_A"),
    };

    notify(&n);
}
  • A function defined in a trait can call other functions in the same trait, even if those aren’t implemented in the trait
  • You can define a function whose parameter just has to implement a trait, without specifying a concrete type
    • The impl in the function signature above is actually syntactic sugar
    • The desugared form is: fn notify<T: Summary>(item: &T)
    • Can you require a parameter to implement multiple traits at once? Yes, use &(impl Trait1 + Trait2)
  • You can also require the return value to implement a trait
  • When a generic function has a lot of trait bounds, you can use the where keyword to make it cleaner
fn some_fn<T, U>(t: &T, u: &U) -> impl Summary
    where T: Summary + Display,
          U: Summary
{
    lib::News{
        headline: String::from("headline_A"),
    }
}

Right now some_fn can only return one type. That is, it can’t return either a News or a News2, even though both implement Summary. This has to do with how it gets compiled. We’ll see later how to make it return values of different types. My guess is that if it returned multiple types, the compiler wouldn’t know how the code receiving the object should call its methods, because it seems traits in Rust aren’t implemented like virtual functions.

The largest function from earlier can also do without traits like Copy or Clone. Here it can just return a slice:

fn largest<T: PartialOrd>(list: &[T]) -> &T {
    let mut largest = &list[0];

    for n in list {
        if *n > *largest {
            largest = &n;
        }
    }
    largest
}

Conditionally Implemented Method

We can also pull this off: implement a method for a type only when its generic parameter implements certain trait(s). The syntax is:

impl<T: Display + PartialOrd> Some_type<T> {
    fn some_fn(&self) {
        // do something
    }
}

Here, some_fn is implemented for Some_type<T> if and only if its T implements both the Display and PartialOrd traits.

Lifetime

The thing that most sets Rust apart from other languages is probably lifetimes. Lifetime here means a variable’s lifetime. Since by default every variable is destroyed as soon as it leaves its scope, sometimes you need certain variables to live longer.

fn main() {
    let r;
    {
        let x = 5;
        r = &x;
    }
    println!("{}", r);
}

Try to compile this code and you’ll get an error:

error[E0597]: `x` does not live long enough  
 --> src/main.rs:5:13  
  |  
5 |         r = &x;  
  |             ^^ borrowed value does not live long enough  
6 |     }  
  |     - `x` dropped here while still borrowed  
7 |     println!("{}", r);  
  |                    - borrow later used here  

When writing functions that deal with references, sometimes you also need to annotate lifetimes explicitly:

fn main() {
    let s1 = String::from("str1");
    {
    let s2 = "s2";
    let result = longest(&s1, s2);
    println!("{}", result);
    }

}

// fn longest (s1: &str, s2: &str) -> &str
// rasie error: expected named lifetime parameter
fn longest<'a>(s1: &'a str, s2: &'a str) -> &'a str {
    if s1.len() > s2.len() {
        s1
    }
    else {
        s2
    }
}

Here, since longest’s parameters s1 and s2 may have different lifetimes, the compiler has no clue what to do after the return (it doesn’t know when the variables should be destroyed). So we have to annotate their lifetimes explicitly. What the annotation means is that the overlap of s1 and s2 (or, the shorter of the two?) is the lifetime of the return value. With the annotation, even a call like the following is ok:

fn main() {
    let s1 = String::from("str1");
    {
        let s2 = "s2";
        let result = longest(&s1, s2);
        println!("{}", result);
    }
}

Therefore, a call like this will error:

fn main() {
    let s1 = String::from("str1");
    let result;
    {
        let s2 = String::from("s2");
        result = longest(&s1, &s2);
    }

    println!("{}", result);
}

Lifetimes are similar to generic parameters and are usually written as ' + a lowercase letter. Like types, lifetimes can’t be changed.

You can annotate lifetimes on structs and methods too:

struct Strstr<'a> {
    part: &'a str,
}

impl<'a> Strstr<'a> {
    fn level(&self) -> i32 {
        42
    }
}

In practice, we don’t need nearly that many annotations when actually writing code, because the Rust compiler has a few rules that annotate lifetimes automatically

  1. Every parameter gets its own distinct lifetime
  2. If there’s only one parameter, the returned lifetime is the same as that parameter’s
  3. If the method has a parameter like self, the returned lifetime is the same as self’s

'static

Besides the annotations above, there’s a special lifetime annotation: 'static. It means the value lives until the program ends. All string literals have this lifetime.

Test

Rust comes with full-fledged testing support. For a library, a test module is generated in lib.rs when you create it. Something like:

#[cfg(test)]
mod tests {
    #[test]
    fn exploration() {
        assert_eq!(2 + 2, 4);
        assert!(true);
    }

    #[test]
    fn another() {
        panic!("Fail!")
    }
}

Put #[test] on top of the functions you want as tests, then run cargo test and the tests run automatically. And marking a module with #[cfg(test)] gives you unit tests for that module.

On top of that, Rust can also catch panics and print more detailed test info.

Functional Rust

Finally, we’re fast-forwarding to magic Rust. Time to get a taste of Rust’s powerful functional features.

Closure

Functions in Rust can be first-class citizens, which means functions can be passed as arguments and returned as values. Here’s an example using a closure:

use std::thread;
use std::time::Duration;

fn generate_workout(intensity: u32, random_number: u32) {
    let expensive_closure = |num| {
        println!("calulating slowly...");
        thread::sleep(Duration::from_secs(2));
        num
    };

    if intensity < 25 {
        println!("Do {} pushups!", expensive_closure(intensity));
        println!("Do {} situps!", expensive_closure(intensity));
    } else {
        if random_number == 3 {
            println!("Take a break");
        } else {
            println!("Run for {} mins!", expensive_closure(intensity));
        }
    }
}

Here expensive_closure is a closure. |num| declares the anonymous function’s parameter, and the function returns num at the end. The closure can be called later on.

You can also annotate its types:

    let expensive_closure = |num: u32| -> u32 {
        println!("calulating slowly...");
        thread::sleep(Duration::from_secs(2));
        num
    };

Memoization

Notice that expensive_closure gets called twice inside the outer if, which is actually inefficient. We can use memoization to get lazy evaluation.

To implement memoization, we build a Cacher struct. It stores a closure, plus the last argument passed to it and the result. Note the use of the Fn trait here: where T: Fn(u32) -> u32. This Cacher implementation isn’t perfect.

struct Cacher<T> where T: Fn(u32) -> u32 {
    calc: T,
    val: Option<u32>,
    arg: Option<u32>,
}

impl<T> Cacher<T> where T: Fn(u32) -> u32 {
    fn new(calc: T) -> Cacher<T> {
        Cacher {
            calc,
            val: None,
            arg: None,
        }
    }

    fn value(&mut self, argument: u32) -> u32 {
        match self.val {
            Some(val) => val,
            None => {
                let v = (self.calc)(argument);
                self.val = Some(v);
                v
            }
        }
    }
}    

Capturing the Environment

Closures in Rust can use variables that are valid in the current scope, like this:

fn main() {
    let x = 4;
    let eqx = |y| y == x;
    eqx(4);
}

Since a closure trades stuff with the scope (environment) it lives in, it runs into the same ownership issues as functions. Rust provides three corresponding traits:

  • FnOnce: takes ownership, and can only be called once (since ownership can only be taken once)
  • FnMut: passes mutable borrows, so it can change the environment
  • Fn: borrows values immutably
fn main() {
    let x = vec![1, 2, 3];

    let eqx = move |z: Vec<i32>| z == x;

    let y = vec![1, 2, 3];
    let z = vec![1, 2, 3];

    eqx(y);
    eqx(z);
    eqx(y);
    println!("{:?}", x);
}

Code like this throws a pile of compile errors:

  • Ownership of x is handed over to eqx, so it can’t be used afterwards
  • Ownership of y is also handed over to eqx when it’s used
  • z can still be used as an argument

Iterator

The Iterator trait in Rust is defined as follows:

pub trait Iterator {
    type Item;

    fn next(&mut self) -> Option<Self::Item>;
}

type Item here is an associated type, which seems to be covered in detail later. Simply put, a Rust iterator needs to implement a next function, and iter() returns an Iterator. Personally, I feel like you could think of this Item in terms of dependent type theory, i.e. it says Item is a type.

This opens the door to lots of functions that are staples of functional programming, like map and filter. Usage:

fn main() {
    let v1 = vec![1, 2, 3, 4, 5];

    let it = v1.iter();

    let v2: Vec<_> = it.map(|x| x + 1).collect();

    let it = v1.iter();

    let v2: Vec<_> = it.filter(|x| **x > 3).collect();
}

Simply put, map takes a function as an argument and applies it; filter filters out every element for which the function returns false. Both produce yet another iterator.

Note:

  • You need collect() to actually evaluate anything, because an Iterator is lazy
  • Why does filter need two dereferences?

Implementing the Iterator trait

Here’s a simple example, tweaked from the one in the book:

struct Counter {
    count: i32,
}

impl Counter {
    fn new(count: i32) -> Self {
        Counter {
            count: count,
        }
    }
}

impl Iterator for Counter {
    type Item = i32;

    fn next(&mut self) -> Option<Self::Item> {
        self.count += 1;
        Some(self.count)
    }
}

fn main() {
    let mut counter = Counter::new(10);

    for i in (0..10) {
        println!("{}", counter.next().unwrap());
    }

}

In a sense, it seems like we could build a stream this way.

Smart Pointer

Smart-ass pointers Some smart pointers can count references automatically and clean things up once the count hits 0.

Box<T>

Box allocates your data on the heap for you. This avoids copying when ownership is transferred. Box’s syntax is as simple as this:

fn main() {
    let b = Box::new(5);
    println!("b = {}", b);
}

Is there anything only Box can do? Yes. Sometimes we don’t know how much space to allocate statically on the stack. A cons list is a recursive type commonly used in functional programming, and if you define it like this, it fails to compile:

enum List {
    Cons(i32, List),
    Nil,
}

Rust even kindly points out:

 --> src/main.rs:1:1
  |
1 | enum List {
  | ^^^^^^^^^ recursive type has infinite size
2 |     Cons(i32, List),
  |               ---- recursive without indirection
  |
help: insert some indirection (e.g., a `Box`, `Rc`, or `&`) to make `List` representable
  |
2 |     Cons(i32, Box<List>),
  |               ^^^^    ^

So you can build a cons list like this:

enum List {
    Cons(i32, Box<List>),
    Nil,
}

use crate::List::{Cons, Nil};

fn main() {
    let ls = Cons(1, Box::new(Cons(2, Box::new(Cons(3, Box::new(Nil))))));
}

Deref and Drop Traits

Smart pointers in Rust are smart because they implement certain traits, which gives them a lot of room for clever tricks when they’re dereferenced or go out of scope. That way, smart pointers can behave like ordinary references.

Deref

The Deref trait lets you customize what happens when you use *, kind of like overloading the dereference operator. First, notice that Box behaves basically the same as an ordinary reference when dereferenced:

fn main() {
    let x = 5;
    let y = &x;
    let y_box = Box::new(x);

    assert_eq!(*y, 5);
    assert_eq!(*y_box, 5);
}

Leaving out the heap allocation part, Box is implemented roughly like this:

use std::ops::Deref;

struct MyBox<T>(T);

impl<T> MyBox<T> {
    fn new(val: T) -> MyBox<T> {
        MyBox(val)
    }
}

impl<T> Deref for MyBox<T> {
    type Target = T;

    fn deref(&self) -> &T {
        &self.0
    }
}

fn main() {
    let x = 5;
    let y_mybox = MyBox::new(x);

    assert_eq!(*y_mybox, 5);
    assert_eq!(*(y_mybox.deref()), 5);
}

Dereferencing a MyBox actually gets turned into something like *(y_mybox.deref()) automatically. Worth noting: deref doesn’t return self.0 directly, but a reference to it. The reason it doesn’t return the value directly is so that ownership isn’t transferred.

To make code easier to write, Rust also introduces a mechanism called Deref Coercions.

fn hello(name: &str) {
    println!("hello, {}", name);
}

fn main() {
    let m = MyBox::new(String::from("f***"));
    hello(&m);
}

This code runs just fine. Here’s roughly how the book (the Chinese edition I was reading) explains it:

We call hello with &m, which is a reference to a MyBox<String> value. Since MyBox<T> implements the Deref trait (back in Listing 15-10), Rust can call deref to turn &MyBox<String> into &String. The standard library implements Deref for String as well, returning a string slice (you can see this in the Deref API docs). So Rust calls deref one more time to turn &String into &str, which finally matches what hello expects.

Without this mechanism, the code would have to look like this:

fn main() {
    let m = MyBox::new(String::from("Rust"));
    hello(&(*m)[..]);
}

Likewise, dereferencing a mutable reference can be overloaded via DerefMut. Also, Deref can dereference a mutable reference into an immutable one.

Drop

You can implement the Drop trait yourself:

struct CustomSmartPointer {
    data: String,
}

impl Drop for CustomSmartPointer {
    // why &mut here?
    fn drop(&mut self) {
        println!("Dropping... {}", self.data);
    }
}

fn main() {
    let c = CustomSmartPointer { data: String::from("first pointer")};
    let d = CustomSmartPointer { data: String::from("second pointer")};

    println!("created pointers");
}

The output is:

created pointers Dropping… second pointer Dropping… first pointer

  • Drop happens in reverse order
  • Rust doesn’t let you call drop directly yourself
  • You can drop things manually with std::mem::drop

Rc<T>

Rc is short for reference count(er). It lets multiple variables own the same pointer. Internally it maintains a reference counter that goes up by one when a new owner is created (cloned) and down by one when one is dropped. Usage:

enum List {
    Cons(i32, Rc<List>),
    Nil,
}

fn main() { 
    let mut a = Rc::new(Cons(5, Rc::new(Cons(10, Rc::new(Nil)))));

    println!("count after creating a: {}", Rc::strong_count(&a));

    // *a = Cons(10, Rc::new(Nil));

    let b = Cons(3, Rc::clone(&a));
    let c = Cons(2, Rc::clone(&a));

    println!("count after creating c: {}", Rc::strong_count(&a));

    println!("{:?}", b);
}

But Rc doesn’t allow the pointer to point to a mutable reference.

RefCell<T>

When it’s an element of some struct, that element can still be changed even if the struct wasn’t instantiated as mutable. Borrows through this pointer get the borrowing rules checked at runtime. A simple example shows how to use it:

pub trait Messenger {
    fn send(&self, msg: &str);
}

pub struct LimitTracker<'a, T:Messenger> {
    messenger: &'a T,
    value: usize,
    max: usize,
}

impl<'a, T> LimitTracker<'a, T> where T:Messenger {
    pub fn new(messenger: &T, max: usize) -> LimitTracker<T> {
        LimitTracker {
            messenger,
            value: 0,
            max,
        }
    }

    pub fn set_value(&mut self, value: usize) {
        self.value = value;

        let percentage = self.value as f64 / self.max as f64;

        if percentage >= 1.0 {
            self.messenger.send("Error: over quota");
        } else if percentage >= 0.9 {
            self.messenger.send("Urgent Warning: over 90% quota");
        } else if percentage >= 0.75 {
            self.messenger.send("Warning: over 75% quota");
        }
    }
}

#[cfg(test)]
mod tests{
    use super::*;
    use std::cell::RefCell;

    struct MockMessenger {
        sent_messages: RefCell<Vec<String>>,
    }

    impl MockMessenger {
        fn new() -> MockMessenger {
            MockMessenger {
                sent_messages: RefCell::new(vec![]),
            }
        }
    }

    impl Messenger for MockMessenger {
        fn send(&self, message: &str) {
            self.sent_messages.borrow_mut().push(String::from(message));
            
            // This panics at runtime
            let b1 = self.sent_messages.borrow_mut();
            let b2 = self.sent_messages.borrow_mut();
        }
    }

    #[test]
    fn sends_over_75() {
        let mockMessenger = MockMessenger::new();
        let mut LimitTracker = LimitTracker::new(&mockMessenger, 100);

        LimitTracker.set_value(80);

        assert_eq!(mockMessenger.sent_messages.borrow().len(), 1);
    }
}

RefCell’s rules are checked at runtime. borrow_mut() returns a mutable borrow, and more than one mutable borrow causes a runtime panic.

This example also shows that the data inside can be modified. Note that the Lists created here are all immutable:

use std::rc::Rc;
use std::cell::RefCell;

#[derive(Debug)]
enum List {
    Cons(Rc<RefCell<i32>>, Rc<List>),
    Nil,
}

use crate::List::{Cons, Nil};

fn main() { 
    let val = Rc::new(RefCell::new(5));

    let a = Rc::new(Cons(Rc::clone(&val), Rc::new(Nil)));
    
    let b = Cons(Rc::new(RefCell::new(3)), Rc::clone(&a));
    let c = Cons(Rc::new(RefCell::new(4)), Rc::clone(&a));

    println!("a = {:?}", a);

    *val.borrow_mut() += 10;
    
    println!("a = {:?}", a);
    println!("b = {:?}", b);
    println!("c = {:?}", c);
}

Whereas this code creates a pointer cycle:

use std::rc::Rc;
use std::cell::RefCell;

#[derive(Debug)]
enum List {
    Cons(i32, RefCell<Rc<List>>),
    Nil,
}

impl List {
    // why &RefCell in Option
    fn tail(&self) -> Option<&RefCell<Rc<List>>> {
        match self {
            Cons(_, item) => Some(item),
            Nil => None,
        }
    }
}

use crate::List::{Cons, Nil};

fn main() { 
    let a = Rc::new(Cons(5, RefCell::new(Rc::new(Nil))));

    println!("a initial rc = {}", Rc::strong_count(&a));
    println!("a next: {:?}", a.tail());

    let b = Rc::new(Cons(10, RefCell::new(Rc::clone(&a))));
    println!("a rc after b = {}", Rc::strong_count(&a));
    println!("b initial rc = {}", Rc::strong_count(&b));
    println!("b next: {:?}", b.tail());

    if let Some(link) = a.tail() {
        *link.borrow_mut() = Rc::clone(&b);
    }

    println!("a rc after b = {}", Rc::strong_count(&a));
    println!("b rc after b = {}", Rc::strong_count(&b));

    // infinite output loop here
    // println!("a next item = {:?}", a.tail());
    // println!("a next item = {:?}", a);
}

At the end, if you print a or a.tail(), it goes into the RefCell and keeps dereferencing in a loop until it overflows. A reference cycle like this means Rust can’t clean up the smart pointers automatically, so you get a memory leak.

To avoid this kind of reference cycle, we can use Rc::downgrade to turn a strong reference into a weak reference (Weak<T>). Rc cleanup is driven by strong_count, and weak_count doesn’t keep it from being dropped.

use std::cell::RefCell;
use std::rc::{Rc, Weak};

#[derive(Debug)]
struct Node {
    value: i32,
    parent: RefCell<Weak<Node>>,
    children: RefCell<Vec<Rc<Node>>>,
}

fn main() {
    let leaf = Rc::new(Node { 
        value: 3,
        parent: RefCell::new(Weak::new()),
        children: RefCell::new(vec![])
    });

    println!("leaf strong = {}, weak = {}", Rc::strong_count(&leaf), Rc::weak_count(&leaf));
    println!("leaf parent: {:?}", leaf.parent.borrow().upgrade());

    {
        let branch = Rc::new(Node {
            value: 5,
            parent: RefCell::new(Weak::new()),
            children: RefCell::new(vec![Rc::clone(&leaf)]),
        });
    
        *leaf.parent.borrow_mut() = Rc::downgrade(&branch);

        println!("branch strong = {}, weak = {}", Rc::strong_count(&branch), Rc::weak_count(&branch));
        println!("leaf strong = {}, weak = {}", Rc::strong_count(&leaf), Rc::weak_count(&leaf));
    }

    println!("leaf parent: {:?}", leaf.parent.borrow().upgrade());
    println!("leaf strong = {}, weak = {}", Rc::strong_count(&leaf), Rc::weak_count(&leaf));
}
}

The code above implements a tree Node. This way, a child node can point back to its parent without creating a reference cycle, because the pointer is Weak.
The code is still a little strange, though: upgrade() doesn’t seem to turn the Weak into an Rc, but just returns an Rc directly.

Project: minigrep

Chapter 12 of the book walks through building a small project in Rust.

// main.rs
use std::env;
use std::process;

use minigrep::Config;


fn main() {
    let args: Vec<String> = env::args().collect();

    let config = Config::new(&args).unwrap_or_else(|err| {
        eprintln!("Problem parsing arguments: {}", err);
        process::exit(1)
    });

    if let Err(e) = minigrep::run(config) {
        eprintln!("Runtime error: {}", e);
        process::exit(1);
    }

}

//lib.rs
use std::error::Error;
use std::fs;
use std::env;

pub struct Config {
    pub query: String,
    pub filename: String,
    pub case_sensitive: bool,
}

impl Config {
    pub fn new(args: &[String]) -> Result<Config, &'static str> {
        if args.len() < 3 {
            return Err("Not enough args");
        }

        let query = args[1].clone();
        let filename = args[2].clone();
        let case_sensitive = env::var("CASE_INSENSITIVE").is_err();
        Ok(Config { query, filename, case_sensitive })
    }
}

pub fn run(config: Config) -> Result<(), Box<dyn Error>>{
    let contents = fs::read_to_string(config.filename)?;

    let result = if config.case_sensitive {
        search(&config.query, &contents)
    } else {
        search_case_insensitive(&config.query, &contents)
    };

    for line in result {
        println!("{}", line);
    }

    Ok(())
}

pub fn search<'a>(query: &str, contents: &'a str) -> Vec<&'a str> {
    let mut result = Vec::new();

    for line in contents.lines() {
        if line.contains(query) {
            result.push(line);
        }
    }

    result
}

pub fn search_case_insensitive<'a>(query: &str, contents: &'a str) -> Vec<&'a str> {
    let query = query.to_lowercase();
    let mut result = Vec::new();

    for line in contents.lines() {
        if line.to_lowercase().contains(&query) {
            result.push(line);
        }
    }

    result
}


#[cfg(test)]
mod tests {
    use super::*;
    
    #[test]
    fn one_result() {
        let query = "duct";
        let contents = "\
Rust:
safe, fast, productive.
Pick three.";

        assert_eq!(
            vec!["safe, fast, productive."], 
            search(query, contents));
    }

    #[test]
    fn case_insensitive() {
        let query = "rUsT";
        let contents = "\
Rust:
safe, fast, productive.
Pick three.
Trust me.";

        assert_eq!(
            vec!["Rust:", "Trust me."], 
            search_case_insensitive(query, contents));
    }
}

I have to call out how shamelessly the official book toots its own horn in the test cases, but this little project is still pretty fun. I won’t go into the implementation details, but here are a few interesting bits:

  • Using a closure in unwrap_or_else
  • Reading environment variables and CLI arguments: env::var() and env::args()
  • Having functions return errors and letting main handle them
  • Lifetime annotations in functions

After learning Rust’s functional features, this code can be rewritten a bit:

impl Config {
    pub fn new(mut args: env::Args) -> Result<Config, &'static str> {
        args.next();

        let query = match args.next() {
            Some(arg) => arg,
            None => return Err("Didn't get query string"),
        };

        let filename = match args.next() {
            Some(arg) => arg,
            None => return Err("Didn't get filename"),
        };

        let case_sensitive = env::var("CASE_INSENSITIVE").is_err();

        Ok(Config { query, filename, case_sensitive })
    }
}

pub fn search<'a>(query: &str, contents: &'a str) -> Vec<&'a str> {
    contents
        .lines()
        .filter(|x| x.contains(query))
        .collect()
}

fn main() {
    let config = Config::new(env::args()).unwrap_or_else(|err| {
        eprintln!("Problem parsing arguments: {}", err);
        process::exit(1)
    });

    if let Err(e) = minigrep::run(config) {
        eprintln!("Runtime error: {}", e);
        process::exit(1);
    }

}

I also tried to simplify the logic of Config::new with a closure:

        let getarg = |argname| {
            match args.next() {
                Some(arg) => arg,
                None => return Err("Didn't get string"),
            }
        };
        
        let query = getarg("query");

But the return here can’t return out of the enclosing function from inside the closure. So I suspect this kind of rewrite can only be done with a macro.

Concurrency

Rust aims for safe multithreading, which it achieves through features like ownership and type enforcement. And since Rust needs only a tiny runtime, its threading is fairly low-level. A simple example:

use std::thread;
use std::time::Duration;

fn main() {
    let handle = thread::spawn(|| {
        for i in 1..10 {
            println!("number {}, spawn", i);
            thread::sleep(Duration::from_millis(1));
        }
    });

    for i in 1..5 {
        println!("number {}, main", i);
        thread::sleep(Duration::from_millis(1));
    }

    handle.join();
}

Thanks to the ownership design, you can easily pass ownership of a variable between threads with move:

use std::thread;
use std::time::Duration;

fn main() {
    let v = vec![1, 2, 3];

    let clos = || println!("vector: {:?}", v);

    clos();

    let handle = thread::spawn(move || {
        println!("vector: {:?}", v);
    });

    handle.join().unwrap();
}

The closure inside spawn() can capture v and use it in the new thread. But if you remove move, you get a compile error. Note that just using a closure directly won’t error, because closures in main always run in a well-defined order. In a multithreaded setting, though, there’s no guarantee v is still valid by the time println! runs in the new thread (it might have been dropped after going out of scope), so ownership has to be transferred.

Channel

Once you have multiple threads, you inevitably need them to talk to each other. The book quotes a Go proverb:

Do not communicate by sharing memory; instead, share memory by communicating.

Rust’s communication mechanism is called a channel, which can carry objects of one particular type.

use std::thread;
use std::sync::mpsc;

fn main() {
    let (tx, rx) = mpsc::channel();

    thread::spawn(move || {
        let val = String::from("hi");
        tx.send(val).unwrap();
    });

    let received = rx.recv().unwrap();
    println!("Got: {}", received);
}   

Note that send takes ownership here.

In this model, you can also have multiple producers:

use std::thread;
use std::sync::mpsc;
use std::time::Duration;

fn main() {
    let (tx, rx) = mpsc::channel();
    let tx1 = tx.clone();

    thread::spawn(move || {
        let vals = vec![
            String::from("Hello"),
            String::from("from"),
            String::from("the"),
            String::from("thread"),
        ];
        
        for val in vals {
            tx1.send(val).unwrap();
            thread::sleep(Duration::from_secs(1));
        }
        // println!("val is {}", val);
    });

    thread::spawn(move || {
        let vals = vec![
            String::from("more"),
            String::from("messages"),
            String::from("for"),
            String::from("you"),
        ];

        for val in vals {
            tx.send(val).unwrap();
            thread::sleep(Duration::from_secs(1));
        }
    });
    
    for recvied in rx {
        println!("Got: {}", recvied);
    }
}   

Shared-state

Besides the smart pointers mentioned earlier that allow multiple owners, Rust also wraps up mutexes:

use std::sync::Mutex;

fn main() {
    let m = Mutex::new(5);

    {
        let mut n = m.lock().unwrap();
        *n += 1;
    }

    println!("m = {:?}", m);
}   

lock() here is really just acquiring the mutex in a blocking way. Notice that the code never unlocks. That’s because of how the smart pointer MutexGuard in Mutex<T> is implemented: its Drop calls unlock() when it goes out of scope.

Using it with multiple threads:

use std::sync::{Mutex, Arc};

fn main() {
    let counter = Arc::new(Mutex::new(0));

    let mut handles = vec![];

    for _ in 0..10 {
        let counter = counter.clone();
        let handle = thread::spawn(move || {
            let mut num = counter.lock().unwrap();
            *num += 1;
        });
        handles.push(handle);
    }

    for handle in handles {
        handle.join().unwrap();
    }

    println!("Result: {}", counter.lock().unwrap());
}   

This code uses Arc instead of Rc. The A stands for atomic; you can think of it as the Rc for multithreaded scenarios, since Arc implements the Send trait and Rc doesn’t. Wrapping the Mutex in an Arc lets it be shared across threads. There’s another neat trick here: let counter = counter.clone(); creates a copy of counter and moves that copy into the closure. Also, counter is actually immutable when it’s created, yet we let mut num and get a mutable reference out of counter. That’s because Mutex, like Cell, provides interior mutability. By the way, counter could also be written as (*counter) in the code

Also, Mutex can lead to deadlocks. Rust can’t stop you from making logic errors.

Sync and Send

Rust provides two traits, Sync and Send, for dealing with concurrency. Implementing Send means ownership can be transferred between threads, while implementing Sync means it can be accessed from multiple threads.

OOP (Object-Oriented Programming)

Rust wasn’t designed strictly around OOP, but it can do everything OOP can. Rust has no concept of inheritance, but you can get something similar by requiring a function’s generic parameters to implement a certain trait (bounded parametric polymorphism).

Beyond that, other OOP features like encapsulating functions (methods) and hiding implementation details/data structures all exist in Rust.

Trait Objects

In Rust, traits are the thing closest to objects, because they enable function reuse across multiple different types. In that sense, they’re closer to the concept of an object than a struct’s or enum’s impl. You can define a dynamic vector of trait objects like this:

//lib.rs
pub trait Draw {
    fn draw(&self);
}

pub struct Screen {
    pub components: Vec<Box<dyn Draw>>,
}

pub struct Button {
    pub width: u32,
    pub height: u32,
    pub label: String,
}

impl Draw for Button {
    fn draw(&self) {
        //do something
    }
}

impl Screen {
    pub fn run(&self) {
        for component in self.components.iter() {
            component.draw();
        }
    }
}

//main.rs
use rust_test::Draw;

struct SelectBox {
    width: u32,
    height: u32,
    options: Vec<String>,
}

impl Draw for SelectBox {
    fn draw(&self) {
        // code
    }
}

use rust_test::{Button, Screen};

fn main() {
    let screen = Screen {
        components: vec![
            Box::new(SelectBox {
                width: 75,
                height: 10,
                options: vec![
                    String::from("Yes"),
                    String::from("No"),
                ],
            }),
            Box::new(Button {
                width: 50,
                height: 10,
                label: String::from("OK"),
            })
        ],
    };
}

The abstraction in the code above is about managing GUI components in a uniform way: users can add their own GUI widgets, as long as they implement the Draw trait. Note the dyn keyword, which roughly marks a type that can only be determined at runtime, so extra runtime code gets added at compile time to do the checking.

In Rust, creating a trait object requires a reference or a smart pointer (Box<T>; my guess is that’s because the size has to be known to allocate memory). The difference between this and generics is that the components vector can hold values of different types, whereas with a generic parameter T it could only hold one type.

Object-safe

Rust only allows object-safe traits to be used as trait objects. The rules are:

  • The return type can’t be Self
  • No generic parameters

For the reasoning, see the Chinese edition of the Rust PL, which puts it roughly like this:

With a trait object, the concrete type gets erased, so there’s no way to know what type should go in for a generic type parameter.

OOP Example

Here Rust PL gives two reference OOP implementations of the same task, using the example of publishing a post. A post starts out as a draft and has to be reviewed before it can be published.

// lib.rs
pub struct Post {
    state: Option<Box<dyn State>>,
    // state: Box<dyn State>,
    content: String,
}

impl Post {
    pub fn new() -> Post {
        Post {
            state: Some(Box::new(Draft {})),
            // state: Box::new(Draft {}),
            content: String::new(),
        }
    }

    pub fn add_text(&mut self, text: &str) {
        self.content.push_str(text);
    }

    pub fn content(&self) -> &str {
        self.state.as_ref().unwrap().content(self)
    }

    pub fn request_review(&mut self) {
        // `take()` takes ownership and set to None temporarily
        if let Some(s) = self.state.take() {
            self.state = Some(s.request_review());
        }
        // self.state = self.state.request_review()
    }

    pub fn approve(&mut self) {
        if let Some(s) = self.state.take() {
            self.state = Some(s.approve());
        }
    }
}

trait State {
    // self: Box<Self> means the caller of this is holds a `Box` for Self type
    // note self will take the ownership, invalidating the current state
    fn request_review(self: Box<Self>) -> Box<dyn State>;
    fn approve(self: Box<Self>) -> Box<dyn State>;
    
    fn content<'a>(&self, post: &'a Post) -> &'a str {
        ""
    }
}

struct Draft {}

impl State for Draft {
    fn request_review(self: Box<Self>) -> Box<dyn State> {
        Box::new(PendingReview {})
    }

    fn approve(self: Box<Self>) -> Box<dyn State> {
        self
    }
}

struct PendingReview {}

impl State for PendingReview {
    fn request_review(self: Box<Self>) -> Box<dyn State> {
        self
    }

    fn approve(self: Box<Self>) -> Box<dyn State> {
        Box::new( Published {})
    }
}

struct Published {}

impl State for Published {
    fn request_review(self: Box<Self>) -> Box<dyn State> {
        self
    }

    fn approve(self: Box<Self>) -> Box<dyn State> {
        self
    }

    fn content<'a>(&self, post: &'a Post) -> &'a str {
        &post.content
    }
}

// main.rs
use rust_test::Post;

fn main() {
    let mut post = Post::new();

    post.add_text("some text??");
    assert_eq!("", post.content());

    post.request_review();
    assert_eq!("", post.content());

    post.approve();
    assert_eq!("some text??", post.content());
}
  • Post contains a trait object, State
  • State handles the state transitions
  • self.state.take() takes ownership to invalidate the previous state
  • Personally I think you could also make State an enum and do the state transitions with match.
  • Using state: Box::new(Draft {}) and self.state = self.state.request_review() fails, presumably because you can’t take ownership
    • cannot move out of self.state which is behind a mutable reference

    • move occurs because self.state has type Box<dyn State>, which does not implement the Copy trait

Besides the implementation above, the book also gives one where the structs themselves are the states, i.e. converting between different structs is the state transition.

The advantage of the implementation shown here is that you can add states freely. You could actually simplify the code by giving the State trait default implementations that return self, but that would break the object-safety rules, and then you couldn’t use trait objects.

Pattern & Matching

Only when I got to this chapter did I realize Rust had been using pattern matching all along, and the sneakiest part is that it never told you that even something as basic as assignment is done via pattern matching. Suddenly I feel kind of gaslit? So I wanted to know how this is implemented differently from other PLs (e.g. C). Instead I stumbled onto the blog of the legendary Yin Wang, found out the guy had actually dropped out of my school at one point, and ended up binge-reading his blog for an entire day and got “scammed” out of five bucks :). In the end, not only did I still not get how it’s actually implemented, I also ate into my Rust study time -_-.

In Rust, match, if let, while let, destructuring, and even let and function argument passing all use pattern matching under the hood. You can even write code like this:

fn main() {
    let color: Option<&str> = None;
    let is_thursday = false;
    let age: Result<u8, _> = "34".parse();

    if let Some(color) = color {
        println!("using color {}", color);
    } else if is_thursday {
        println!("Today is thursday");
    } else if let Ok(age) = age {
        if age > 30 {
            println!("old boy");
        } else {
            println!("young boy");
        }
    } else {
        println!("Nothing happened");
    }
}

It mixes if let with regular conditional branches. Personally I don’t think combining code like this is especially logical, and in a way it’s pretty likely to cause confusion.

Refutability

However, pattern matching in Rust comes in different flavors. Consider the following code and error:

    let Some(x) = Some(1);

The error:

error[E0005]: refutable pattern in local binding: `None` not covered
   --> src/main.rs:2:9
    |
2   |     let Some(x) = Some(1);
    |         ^^^^^^^ pattern `None` not covered
    | 
   ::: /home/ya0guang/.rustup/toolchains/stable-x86_64-unknown-linux-gnu/lib/rustlib/src/rust/library/core/src/option.rs:165:5
    |
165 |     None,
    |     ---- not covered
    |
    = note: `let` bindings require an "irrefutable pattern", like a `struct` or an `enum` with only one variant
    = note: for more information, visit https://doc.rust-lang.org/book/ch18-02-refutability.html
    = note: the matched value is of type `Option<i32>`
help: you might want to use `if let` to ignore the variant that isn't matched

Here Rustc is telling us that let needs an irrefutable pattern, i.e. a pattern that can’t be rejected, meaning the variable binding can never fail. Something like if let, on the other hand, can accept a refutable pattern.

Matching Syntaxes

match

_ has a special meaning in matching:

fn main() {
    let s = Some(String::from("hello"));

    if let Some(x) = s {
        println!("found string");
    }

    println!("{:?}", s);
}

Matching with either Some(_x) or Some(x) gives an error: borrow of partially moved value: s. But matching with Some(_) doesn’t cause any error, because it doesn’t bind anything. By the way, you could also use &s here to avoid the move.

  • A match creates a new scope, and may shadow variables from outside
  • You can use | as a logical OR to match multiple patterns
  • ..= matches a range, e.g. 1..=5; only numbers/char are allowed
  • _ is a wildcard for the parts you want to “ignore”, and it creates no binding. As a symbol, its semantics are different from any other variable name!

Destructuring

This code is a bit odd to write:

struct Point {
    x: i32,
    y: i32,
}

fn main() {
    let p = Point { x: 0, y: 7};

    let Point {x: a, y: b} = p;
    // in case: let Point {x: a, y: b} = p;
    // var x, y will be created
    assert_eq!(0, a);
    assert_eq!(7, b);
    
    match p {
        Point {x, y: 0} => println!("On x axis at {}", x),
        Point {x: 0, y} => println!("On y axis at {}", y),
        Point {x, y} => println!("Not on axis at {}, {}", x, y),
    }
}

More examples:

struct Point {
    x: i32,
    y: i32,
    z: i32,
}

fn main() {
    let origin = Point {x: 0, y: 0, z: 0};

    match origin {
        Point {x, ..} => println!("on x axis"),
    };
}
  • Destructuring also works on structs/tuples/enums
  • Likewise, it works on nested types
  • .. matches the fields you can’t be bothered with, and you can drop .. in the middle as long as it’s unambiguous

Match Guard

You can add an if expression after a pattern so it only matches when the expression is true, which makes patterns more expressive:

fn main() {
    let x = Some(10);

    match x {
        Some(x) if x <= 10 => println!("x is {}", x),
        _ => println!("Nothing"),
    }
}

A match guard has lower precedence than |.

Rust also provides something that looks like syntactic sugar to me, @, which lets you test a value against a range and save it in the same pattern.

// from Rust PL book
enum Message {
    Hello { id: i32 },
}

let msg = Message::Hello { id: 5 };

match msg {
    Message::Hello { id: id_variable @ 3..=7 } => {
        println!("Found an id in range: {}", id_variable)
    },
    Message::Hello { id: 10..=12 } => {
        println!("Found an id in another range")
    },
    Message::Hello { id } => {
        println!("Found some other id: {}", id)
    },
}

Advanced Features

Finally, the last chapter on language features rubs hands excitedly.

unsafe

Since computers are unsafe by nature, Rust provides some unsafe “superpowers” for system-level programming.

Dereferencing Raw Pointers

fn main() {
    let mut num = 42;
    
    let r1 = &num as * const i32;
    let r2 = &mut num as * mut i32;

    unsafe {
        println!("r1: {}", *r1);
        println!("r2: {}", *r2);
    }

    let address = 0x12345usize;
    let r = address as *const i32;
    println!("addr: {:?}",r);
    // seg fault
    unsafe {
        println!("deref a invalid pointer: {:?}", *r);
    }
}

Rust’s raw pointers come in two forms (types): * const T and * mut T. Creating a raw pointer isn’t illegal in itself, but dereferencing one takes superpowers. We can even make one straight out of an address.

Calling Unsafe Functions/Methods

Here’s an example:

use std::slice;

fn split_at_mut(slice: &mut [i32], mid: usize) -> (&mut [i32], &mut [i32]) {
    let len = slice.len();
    let ptr = slice.as_mut_ptr();

    assert!(len >= mid);

    unsafe {
        (
            slice::from_raw_parts_mut(ptr, mid),
            slice::from_raw_parts_mut(ptr.add(mid), len - mid),
        )
        
    }
    // (&mut slice[..mid], &mut slice[mid..])


}

fn main() {
    let mut v = vec![1, 2, 3, 4, 5, 6];

    let r = &mut v;

    let (a, b) = split_at_mut(r, 3);

    assert_eq!(a, &mut [1, 2, 3]);
    assert_eq!(b, &mut [4, 5, 6]);
}

If you skip the stuff inside unsafe and use the commented-out code instead, it won’t compile: Rust doesn’t allow two mutable references to the same variable. So we need the unsafe functions provided by std::slice.

extern

Lets Rust use external functions, or export a function written in Rust to other languages:

extern "C" {
    fn abs(input: i32) -> i32;
}

#[no_mangle]
pub extern "C" fn call_from_c() {
    println!("Hello Rust");
}

fn main() {
    unsafe {
        println!("abs fo -3: {}", abs(-3));
    }
}

Here, no_mangle tells the compiler not to mangle the function name.

Mutable Static Variables

static variables in Rust live at a fixed location in memory. They’re not supposed to be mutable, but with this trick:

static mut s: &str = "hello";

fn push(ins: &'static str) {
    unsafe {
        s = ins;
    }
}

fn main() {
    push("Rust");
    
    unsafe { println!("unsafe string: {}", s);}
}

Pretty hacky. The original book uses a u32, but I wanted to see if I could change the value of a &str. I still don’t know how to force a String’s value into this &str, but I have a feeling it’d take a painful fight with the compiler.

Implementing Unsafe Traits

unsafe trait Foo {
    // do something
}

unsafe impl Foo for i32 {
    // do something else
}

Accessing Fields of a Union

Advanced Traits

Associated Type

You can use associated types in traits:

pub trait Iterator {
    type Item;

    fn next(&mut self) -> Option<Self::Item>;
}

impl Iterator for Counter {
    type Item = u32;

    fn next(&mut self) -> Option<Self::Item> {
        // code
    }
}

// generics signature
pub trait Iterator<T> {
    fn next(&mut self) -> Option<T>;
}

Of course, you could do something similar with generics. An associated type limits us to a single implementation of the trait for a type, while generics let us provide different implementations for multiple types.

Default Generic Type Para & Overloading

You can give generic parameters default values and overload operators:

use std::ops::Add;

#[derive(Debug, PartialEq)]
struct Point {
    x: i32,
    y: i32,
}

impl Add for Point {
    type Output = Point;

    fn add(self, other: Point) -> Point {
        Point { 
            x: self.x + other.x,
            y: self.y + other.y,
        }
    }
}

fn main() {
    assert_eq!(
        Point { x: 1, y: 2 } + Point { x: 2, y: 1},
        Point { x: 3, y: 3}
    )
}

// `Add` trait
trait Add<Rhs=Self> {
    type Output;

    fn add(self, rhs: Self) -> Self::Output;
}

Disambiguation

If methods with the same name are defined in different traits, and some type then implements all of those traits, you can call them like this: Trait::method(&var). The more general syntax is: <Type as Trait>::function().

Supertrait

When one trait relies on another trait for its functionality, the latter is a supertrait.

use std::fmt;

trait OutlinePrint: fmt::Display {
    fn outline_print(&self) {
        let output = self.to_string();
        let len = output.len();

        println!("{}", "*".repeat(len + 4));
        println!("* {} *", output);
        println!("{}", "*".repeat(len + 4));
    }
}

impl OutlinePrint for Point {}

impl fmt::Display for Point {
    fn fmt(&self, fmt: &mut fmt::Formatter) -> fmt::Result {
        write!(fmt, "({}, {})", self.x, self.y)
    }
}

Advanced Types

// nothing different from i32 for Kilometers
type Kilometers = i32;

fn main() {
    let km: Kilometers = 42;
    let m: i32 = 44;
    // can still add km and m here 
    let x = km + m;
}

fn generic<T: ?Sized>(t: &T);
  • You can create a type alias with code like type Kilometers = i32;
  • Rust has a weird ! type:
    • fn bar() -> ! means the function never returns
    • panic!() has type !
    • continue has type !
    • ! can be treated as any type, which lets every arm of a match have the same return type
  • Because of static compilation/DST (dynamically sized type) concerns, Rust has the Sized trait
    • It determines whether a type’s size is known at compile time
    • All generic parameters must implement Sized
    • Following on from that: unless you use <T: ?Sized>, in which case you can use &T in the function

Advanced Functions/Closures

You can pass a function to another function as an argument like this:

fn add_one(x: i32) -> i32 {
    x + 1
}

fn do_twice(f: fn(i32) -> i32, arg: i32) -> i32 {
    f(arg) + f(arg)
}

fn main() {
    let answer = do_twice(add_one, 10);

}

Gotta say, this style reminds me of Typed Racket.
The type of f here is a function pointer. Note that fn is a type that implements all of the closure traits; don’t mix it up with the closure trait Fn. So in theory, anywhere that accepts a closure should also accept an fn.

You can also initialize an array with a slick trick like this:

enum Status {
    Value(u32),
    Stop,
}

fn main() {
    let list: Vec<Status> = (0u32..20).map(Status::Value).collect();
}

And of course, a function can also return a closure:

// fn return_closure() -> dyn Fn(i32) -> i32 {
//     |x| x + 1
// }

fn return_closure() -> Box<dyn Fn(i32) -> i32> {
    Box::new(|x| x + 1)
}

Since a closure’s size isn’t known at compile time (the Sized issue), it has to be wrapped in a pointer.

Macros

Macros are a form of metaprogramming. Simply put, they let Rust write Rust for you.

Declarative Macro

For example, here’s how a simplified vec! is implemented:

#[macro_export]
macro_rules! vec {
    ( $( $x:expr ),* ) => {
        {
            let mut temp_vec = Vec::new();
            $(
                temp_vec.push($x);
            )*
            temp_vec
        }
    };
}

This rule is actually a bit like Racket’s define-syntax, and it’s basically how macros get defined.

  • ( $( $x:expr ),* ) is the pattern to match
  • In $(),*, the , may or may not appear, and * means match any number of times
  • => means the same thing as in match
  • $x:expr matches a Rust expression and binds it to $x
  • #[macro_export] makes it available to callers

Procedural Macro

More like a function: it takes code as input and outputs code.

use proc_macro;

#[some_attribute]
pub fn some_name(input: TokenStream) -> TokenStream {}

Custom Derive

Create a rust_test_derive lib inside the rust_test directory like so:

extern crate proc_macro;

use proc_macro::TokenStream;
use quote::quote;
use syn;

#[proc_macro_derive(HelloMarco)]
pub fn hello_macro_derive(input: TokenStream) -> TokenStream {
    let ast = syn::parse(input).unwrap();

    impl_hello_macro(&ast)
}

fn impl_hello_macro(ast: &syn::DeriveInput) -> TokenStream {
    let name = &ast.ident;
    let gen = quote! {
        impl HelloMarco for #name {
            fn hello_macro() {
                println!("Hello!, my name is {}!", stringify!(#name));
            }
        }
    };
    gen.into()
}

And add this to that project’s config file:

[lib]
proc-macro = true

[dependencies]
syn = "1"
quote = "1"

The code in the outer rust_test project looks like this:

// lib.rs
pub trait HelloMarco {
    fn hello_macro();
}

// main.rs
use rust_test::HelloMarco;
use rust_test_derive::HelloMarco;

#[derive(HelloMarco)]
struct Pancakes;

fn main() {
    Pancakes::hello_macro();
}

Remember to add rust_test_derive to the config: rust_test_derive = {path = "./rust_test_derive"}.

Simply put, this code works by having the Rust compiler build an AST, grabbing the type-name metadata from the AST, generating a new impl from that metadata, and inserting it back into the original code.

  • Rust can’t know type names at runtime, but you can get them via macros
  • An annoying Rust rule means macro definitions have to live in their own crate
  • proc_macro is the compiler’s API for reading code
  • syn is the crate for building the AST, and quote generates the code

Besides that, there are also attribute-like macros and function-like macros. The former are like the decorators a Python web server uses to map URL paths to handler functions, while the latter are similar to macro_rules! but rely on #[proc_macro].

Final Project: Multithreaded Web Server

These few lines of code are enough for a tiny webserver.

use std::net::TcpListener;
use std::net::TcpStream;
use std::io::prelude::*;
use std::fs;

fn main() {
    let listener = TcpListener::bind("127.0.0.1:7878").unwrap();

    for stream in listener.incoming() {
        let stream = stream.unwrap();

        println!("connection established!");
        handle_connection(stream);
    }
}

fn handle_connection(mut stream: TcpStream) {
    let mut buffer = [0; 1024];
    stream.read(&mut buffer).unwrap();

    let get = b"GET / HTTP/1.1\r\n";

    let (status_line, filename) = if buffer.starts_with(get) {
        ("HTTP/1.1 200 OK\r\n\r\n", "hello.html")
    } else {
        ("HTTP/1.1 404 NOT FOUND\r\n\r\n", "404.html")
    };

    println!("Request from client: {}", String::from_utf8_lossy(&buffer));

    let content = fs::read_to_string(filename).unwrap();

    let response = format!("{} \r\nContent-Length: {}\r\n\r\n{}", status_line, content.len(), content);
    stream.write(response.as_bytes()).unwrap();
    stream.flush().unwrap();
}

But if the work in handle_connection is heavy, it can’t serve multiple users efficiently. Below, we improve this code with a thread pool.

// file: src/bin/main.rs
use std::net::TcpListener;
use std::net::TcpStream;
use std::io::prelude::*;
use std::fs;
use std::time::Duration;
use std::thread;
use server::ThreadPool;

fn main() {
    let listener = TcpListener::bind("127.0.0.1:7878").unwrap();
    let pool = ThreadPool::new(4);


    //or stream in listener.incoming().take(2) {
    for stream in listener.incoming() {
        let stream = stream.unwrap();

        println!("[Server:] connection established!");
        pool.execute(|| {
            handle_connection(stream);
        });
    }
}

fn handle_connection(mut stream: TcpStream) {
    let mut buffer = [0; 1024];
    stream.read(&mut buffer).unwrap();

    let get = b"GET / HTTP/1.1\r\n";
    let sleep = b"GET /sleep HTTP/1.1\r\n";

    let (status_line, filename) = if buffer.starts_with(get) {
        ("HTTP/1.1 200 OK\r\n\r\n", "hello.html")
    } else if buffer.starts_with(sleep) {
        thread::sleep(Duration::from_secs(5));
        ("HTTP/1.1 200 OK\r\n\r\n", "hello.html")
    } else {
        ("HTTP/1.1 404 NOT FOUND\r\n\r\n", "404.html")
    };

    println!("Request from client: {}", String::from_utf8_lossy(&buffer));

    let content = fs::read_to_string(filename).unwrap();

    let response = format!("{} \r\nContent-Length: {}\r\n\r\n{}", status_line, content.len(), content);
    stream.write(response.as_bytes()).unwrap();
    stream.flush().unwrap();
}

// file: src/lib.rs
use std::thread;
use std::sync::mpsc;
use std::sync::Arc;
use std::sync::Mutex;

pub struct ThreadPool {
    workers: Vec<Worker>,
    sender: mpsc::Sender<Message>,
}

type Job = Box<dyn FnOnce() + Send + 'static>;

enum Message {
    NewJob(Job),
    Terminate,
}

struct Worker {
    id: usize,
    job: Option<thread::JoinHandle<()>>,
}

impl Worker {
    fn new(id: usize, receiver: Arc<Mutex<mpsc::Receiver<Message>>>) -> Worker {
        let job = thread::spawn(move || loop {
            let message = receiver.lock().unwrap().recv().unwrap();

            match message {
                Message::NewJob(job) => {
                    println!("Worker {} got a job; exec...", id);
                    job();
                }
                Message::Terminate => {
                    println!("Worker {} told me to stop work", id);
                    break;
                }
            }

        });

        Worker {
            id,
            job: Some(job),
        }
    }
}

impl ThreadPool {
    /// Creates a new thread pool
    /// 
    /// # Panics
    /// 
    /// size cannot be zero or negetive
    pub fn new(size: usize) -> ThreadPool {
        assert!(size > 0);

        let (sender, receiver) = mpsc::channel();
        let receiver = Arc::new(Mutex::new(receiver));

        let mut workers = Vec::with_capacity(size);

        for i in 0..size {
            workers.push(Worker::new(i, Arc::clone(&receiver)))
        }

        ThreadPool {
            workers,
            sender,
        }
    }

    pub fn execute<F>(&self, f: F)
    where F: FnOnce() + Send + 'static {
        let job = Box::new(f);

        self.sender.send(Message::NewJob(job)).unwrap();
    }
}

impl Drop for ThreadPool {
    fn drop(&mut self) {
        println!("Sending terminate msg to workers");

        for _ in &self.workers {
            self.sender.send(Message::Terminate).unwrap();
        }

        for worker in &mut self.workers {
            println!("Shutting down worker {}", worker.id);
            
            if let Some(thread) = worker.job.take() {
                thread.join().unwrap();
            }
        }
    }
}
  • This program implements a thread-pool web server; the idea is basically the same as a C++ implementation
  • It shuts down gracefully, but doesn’t handle shutdown via a Signal mechanism
  • It uses advanced features like the channels, Arc, Mutex, etc. that we saw earlier.

Epilogue

That about wraps up my study of the Rust PL book, but learning programming languages seems to be a never-ending story. Rust really is a super cool language: it manages to be concise and very “verbose” at the same time, which is a really contradictory combo. I guess verbosity is the price of safety. To save your fingers from all that typing, let me plug the tabnine extension for VSCode. It’ll save you a ton of keystrokes!

But at what cost?

Other References