Image classification: from a scalar model to a tested experiment contract
Derive loss and softmax, plan controlled Fashion-MNIST experiments, and follow complete local TensorFlow references for training, callbacks and save/load.
Lesson preparation & details
Level: intermediate
By the end, you should be able to
- Derive the scalar MSE gradients and distinguish softmax from argmax
- Trace image batches, Dense parameters and the sparse-label loss contract
- Design validation-selected capacity and normalization experiments
- Follow complete local-data training and serialization references with explicit runtime limits
Bring with you
- Python, NumPy arrays and basic algebra
Editorial review: · What review means
In this article · 12 sections
Two experiments, one learning loop
Start with a relationship whose answer you know: the six pairs below obey y=3x+1. Then replace a scalar with an image and a numeric target with a category. The loop stays the same—forward computation, loss, parameter update, held-out evaluation—but the shapes and loss contract change. This is a corrected companion to Google's image-classification course, not proof that a cloud lab was completed.
By the end you should be able to explain Dense, Flatten, ReLU, softmax, sparse cross-entropy, batching, callbacks and model serialization; run a small independent arithmetic check; and follow a complete local-data TensorFlow reference without mistaking training accuracy for generalization.
Environment and the optional Google Cloud route
The executed example below uses only NumPy in the existing CPU runtime. The two TensorFlow programs are complete local-data references, syntax checked but not executed here. TensorFlow, TensorFlow Datasets, Workbench, Cloud Logging and Fashion-MNIST training were not run or downloaded. TensorFlow/Keras behaviour is tied to the official documentation snapshots accessed 8 October 2026; before using a different installed release, record its versions and check its save/load and callback APIs.
For an already authorized Google Cloud lab, use the lab's temporary account and project, open the instructed Vertex AI Workbench instance, choose its Python kernel, and verify the project/region and kernel interpreter before executing notebook cells. Do not copy guest credentials into source. Outside a supervised lab, provisioning needs a budget, IAM review, region/storage decision and cleanup plan. The captured setup page now uses Agent Platform Workbench branding; these older course notes use Vertex AI Workbench terminology. Bind setup steps to the page/version you actually use. Workbench is the notebook host, not the learning algorithm. Local Python is sufficient for the scalar experiment.
Keep environment setup separate from Python: pip invocations belong in a terminal or an explicitly marked notebook magic, not in a Python program. Avoid upgrading the active interpreter's dependencies halfway through a notebook. Use a separately approved environment and record tensorflow, keras, numpy, tensorflow-datasets and google-cloud-logging versions if those packages are actually used. The original notes mixed installation commands with executable Python, omitted a NumPy import, configured logging twice and trained the scalar model twice; neither duplication is needed.
Notebook readiness check: restart the selected kernel, execute imports first, display the interpreter/version, then run the scalar model top to bottom. A successful import establishes package availability, not access to a cloud project. Run local assertions before enabling optional remote logging.
Lab 1: learn the line, not a screenshot
Traditional software specifies a mapping directly; supervised learning chooses parameters using examples and an objective. Neither eliminates human design: the feature representation, labels, model family and loss are specified by the practitioner. A difficult activity-recognition problem is not impossible to program, and learning from examples does not guarantee a correct rule outside those examples.
A one-unit Dense layer on a scalar computes y-hat=wx+b. Its two trainable numbers are w and b. For the inputs −1,0,1,2,3,4, the targets −2,1,4,7,10,13 obey w=3,b=1 exactly. At x=10 the reference function gives31. A trained prediction near31 is an extrapolation check on this invented relationship, not evidence that all future data obey it.
For n examples the mean squared error and derivatives are
At w=b=0, the mean target is5.5, mean x is1.5 and mean x-squared is31/6. Therefore the gradients are−34 and−11. With learning rate0.01 the first update is w=.34,b=.11. This is a computed gradient step, not a random guess of another line. A larger learning rate can overshoot; lower loss is checked, not assumed. Scientific notation such as 1e-4 means0.0001.
Executed CPU oracle: gradients and class probabilities
This program uses the lesson's exact six points. It verifies the analytic gradients by finite differences, solves the line independently by least squares, and tests the classification formulas needed for Lab2. Its probabilities are synthetic numbers, not Fashion-MNIST accuracy.
import numpy as np
x = np.array([-1., 0., 1., 2., 3., 4.])
y = np.array([-2., 1., 4., 7., 10., 13.])
def mse(w, b):
return np.mean((w*x+b-y)**2)
r = -y
grad = np.array([2*np.mean(r*x), 2*np.mean(r)])
np.testing.assert_allclose(grad, [-34., -11.])
eps = 1e-5
numeric = np.array([(mse(eps, 0)-mse(-eps, 0))/(2*eps),
(mse(0, eps)-mse(0, -eps))/(2*eps)])
np.testing.assert_allclose(grad, numeric, atol=1e-8)
assert mse(.34, .11) < mse(0, 0)
w, b = np.linalg.lstsq(np.column_stack([x, np.ones(len(x))]), y, rcond=None)[0]
np.testing.assert_allclose([w, b, 10*w+b], [3., 1., 31.])
def softmax(z):
z = np.asarray(z, dtype=np.float64)
e = np.exp(z-z.max(axis=-1, keepdims=True))
return e/e.sum(axis=-1, keepdims=True)
p = softmax([0., 1., 2.])
np.testing.assert_allclose(p.sum(), 1.)
assert (p > 0).all() and (p < 1).all() and p.argmax() == 2
np.testing.assert_allclose(p, softmax([1000., 1001., 1002.]))
assert not np.array_equal(p, [0., 0., 1.]) # softmax is not argmax/one-hot
loss = -np.log(p[2])
np.testing.assert_allclose(loss, np.log(np.exp(-2)+np.exp(-1)+1))
# Passing probabilities to a logits loss applies softmax twice: different loss.
wrong_loss = -np.log(softmax(p)[2])
assert abs(wrong_loss-loss) > .1
pixels = np.array([0, 127, 255], dtype=np.uint8).astype(np.float32)/255.
assert pixels.min() == 0 and pixels.max() == 1
assert 784*64+64 + 64*10+10 == 50890
assert 784*128+128 + 128*10+10 == 101770
assert 784*64+64 + 64*64+64 + 64*10+10 == 55050
print('line, gradients, softmax/loss and capacity controls passed')Local TensorFlow scalar reference — not run
Save this as a separate script. It imports everything it uses, sets a seed, states the input shape explicitly and creates one model. SGD minimizes MSE; neither optimizer nor loss is an accuracy metric. This deterministic mathematical dataset needs no external fetch. Exact training values still depend on framework/runtime numerics.
import numpy as np
import tensorflow as tf
tf.keras.utils.set_random_seed(646)
x = np.array([-1., 0., 1., 2., 3., 4.], dtype=np.float32).reshape(-1, 1)
y = 3*x+1
model = tf.keras.Sequential([
tf.keras.Input(shape=(1,)),
tf.keras.layers.Dense(1),
])
model.compile(optimizer=tf.keras.optimizers.SGD(learning_rate=.01),
loss=tf.keras.losses.MeanSquaredError())
history = model.fit(x, y, epochs=500, batch_size=6, shuffle=False, verbose=0)
pred = model.predict(np.array([[10.]], dtype=np.float32), verbose=0)
print('first/last training loss:', history.history['loss'][0], history.history['loss'][-1])
print('weights and bias:', model.get_weights(), 'prediction at10:', pred[0, 0])
assert pred.shape == (1, 1) and np.isfinite(pred).all()The loss history tests whether optimization improved the fitted objective. The least-squares oracle establishes the reference parameters independently. To investigate a poor fit, inspect learning rate, loss reduction, input shape, dtype and number of updates; do not assume that500 epochs guarantees a specified tolerance.
Lab 2: Fashion-MNIST is a ten-class task
The dataset authors describe60,000 training and10,000 test examples, each a28×28 grayscale image with one of ten labels. The labels are T-shirt/top, Trouser, Pullover, Dress, Coat, Sandal, Shirt, Sneaker, Bag and Ankle boot, in that integer order0–9. These are clothing images, not a general-purpose object-recognition benchmark. Dataset identity and label order must travel with a saved model.
An individual TFDS example is an image/label pair, not a batch: naming it image_batch before calling .batch() was misleading. TFDS images commonly include a final single-channel dimension, whereas a local NumPy array may be (N,28,28). In the reference below we deliberately require the latter. Inspect the whole image's minimum/maximum, not only its first row. Divide integer pixel values by255 after conversion to float32; this fixed transform learns no statistics. If you later standardize using a fitted mean/variance, fit those statistics on training only.
Shapes, parameters and nonlinearities
For a batch of B images, Flatten maps (B,28,28) to (B,784) without learned weights. Dense64 computes (B,784) @ (784,64) + (64,); ReLU applies max(0,z) elementwise. Dense10 maps those features to ten logits. The parameter count is (784×64+64)+(64×10+10)=50,890. Flatten preserves pixel values but discards explicit spatial layout; a CNN supplies a different inductive bias.
ReLU is not a probabilistic neuron switch: positive values remain their original magnitude. Two linear Dense layers without a nonlinearity still compose to an affine map, so the hidden layer alone does not create a nonlinear classifier. Softmax converts logits into a normalized distribution; it never turns finite logits into a one-hot class decision. For logits0,1,2 it gives approximately .0900,.2447,.6652. argmax selects index2 separately. Softmax scores need not be calibrated confidence.
Integer labels use sparse categorical cross-entropy. With logits, use SparseCategoricalCrossentropy(from_logits=True); with a softmax output, use from_logits=False. Both valid configurations represent the same mathematical likelihood. Applying softmax and then using a logits loss is a different calculation, demonstrated by the negative control above. A ten-class output is required even if a mini-batch happens to contain fewer classes.
Complete local-data training, callback, evaluation and reload reference
This reference accepts an existing, authorized local fashion_mnist.npz with keys x_train, y_train, x_test, y_test. It does not call load_data, tfds.load, a network API or an installer. Prepare that artifact separately if you choose to run the experiment. Record its SHA-256, provenance, dataset licence and framework versions. An official TFDS adapter is an alternative ingestion path; its default download behaviour is not part of this local-only recipe.
The program reserves6,000 official training rows for validation using a fixed seed, selects checkpoints on validation loss, and evaluates the official test split only after the recipe is fixed. It shows a safe custom callback for the original84% training-accuracy exercise, but that optional stop is deliberately disabled for the default validation-selected run.
from pathlib import Path
import hashlib
import numpy as np
import tensorflow as tf
path = Path('fashion_mnist.npz')
if not path.is_file():
raise FileNotFoundError('Supply an authorized local Fashion-MNIST NPZ; no download is performed')
print('dataset sha256:', hashlib.sha256(path.read_bytes()).hexdigest())
print('versions:', tf.__version__, np.__version__)
tf.keras.utils.set_random_seed(646)
with np.load(path, allow_pickle=False) as data:
x, y = data['x_train'], data['y_train']
xt, yt = data['x_test'], data['y_test']
assert x.shape == (60000, 28, 28) and xt.shape == (10000, 28, 28)
assert y.shape == (60000,) and yt.shape == (10000,)
assert x.dtype == np.uint8 and xt.dtype == np.uint8
assert np.issubdtype(y.dtype, np.integer) and np.issubdtype(yt.dtype, np.integer)
assert y.min() >= 0 and y.max() < 10 and yt.min() >= 0 and yt.max() < 10
order = np.random.default_rng(646).permutation(len(x))
tr, va = order[6000:], order[:6000]
assert not set(tr) & set(va)
print('first full image raw min/max:', x[0].min(), x[0].max())
x = x.astype(np.float32)/255.
xt = xt.astype(np.float32)/255.
print('first full image scaled min/max:', x[0].min(), x[0].max())
class StopOnTrainingAccuracy(tf.keras.callbacks.Callback):
def __init__(self, threshold=.84):
super().__init__()
self.threshold = threshold
def on_epoch_end(self, epoch, logs=None):
value = (logs or {}).get('sparse_categorical_accuracy')
if value is not None and value > self.threshold:
print('Training threshold exceeded; this is not test accuracy')
self.model.stop_training = True
model = tf.keras.Sequential([
tf.keras.Input(shape=(28, 28)),
tf.keras.layers.Flatten(),
tf.keras.layers.Dense(64, activation='relu'),
tf.keras.layers.Dense(10),
])
model.compile(optimizer=tf.keras.optimizers.Adam(),
loss=tf.keras.losses.SparseCategoricalCrossentropy(from_logits=True),
metrics=[tf.keras.metrics.SparseCategoricalAccuracy()])
assert model.count_params() == 50890
callbacks = [tf.keras.callbacks.EarlyStopping(
monitor='val_loss', patience=3, restore_best_weights=True)]
# Optional separate teaching experiment, not the default selection rule:
# callbacks = [StopOnTrainingAccuracy(.84)]
history = model.fit(x[tr], y[tr], validation_data=(x[va], y[va]),
batch_size=32, epochs=20, callbacks=callbacks, verbose=2)
model.summary()
# Freeze all model/threshold choices before this call.
print('test:', model.evaluate(xt, yt, batch_size=32, return_dict=True, verbose=0))
before = model.predict(xt[:8], verbose=0)
model.save('fashion_classifier.keras')
loaded = tf.keras.models.load_model('fashion_classifier.keras')
after = loaded.predict(xt[:8], verbose=0)
np.testing.assert_allclose(before, after, rtol=1e-5, atol=1e-6)
probs = tf.nn.softmax(after, axis=-1).numpy()
print('predicted label IDs:', probs.argmax(axis=-1))
assert probs.shape == (8, 10)
np.testing.assert_allclose(probs.sum(axis=-1), 1., atol=1e-6)Batch32 means up to32 examples per optimizer update, not32 epochs. The final partial batch can be smaller. Epoch20 is an upper bound in this reference: early stopping may stop sooner. Validation observations influence selection, so they are not independent final-test evidence. Restoring best weights is enough for evaluating that candidate; it does not necessarily rewind the optimizer's complete training state to the best epoch for exact continuation.
The .keras format stores architecture, weights and training configuration/optimizer state where available. This example uses built-in activations, so the original custom_objects={'softmax_v2': ...} workaround is unnecessary. Custom layers require their own serialization contract. Only load trusted artifacts. Output parity on8 images checks that round-trip, not general correctness, calibration, portability across arbitrary versions or equal training resumption. Saving the same model to two filenames adds no evaluation evidence.
Optional Cloud Logging adapter
Cloud logging is observability, not a mathematical dependency. In an authorized environment, use Application Default Credentials and the least-privilege logging role for the intended project. Do not put credentials in a notebook. The official Python integration documents Client().setup_logging(); configure it once per process rather than repeatedly adding handlers, which can duplicate messages. This adapter is unexecuted and makes remote writes if used:
# Optional remote-service reference; NOT executed in this review.
import logging
import google.cloud.logging
client = google.cloud.logging.Client()
client.setup_logging()
logging.info('Scalar model experiment completed; inspect locally recorded metrics')Log experiment IDs, versions, aggregate metrics and errors without dumping images, credentials or sensitive labels. A successful log entry proves delivery to a logging service, not successful training. Keep an ordinary local logging.StreamHandler path when remote observability is unnecessary. At the end of a cloud exercise, follow the lab's resource cleanup instructions and verify instance/disk lifecycle and billing separately; shutting a notebook tab is not a cleanup operation.
Controlled experiments: capacity, depth and input scale
Keep training/validation partitions fixed and instantiate a fresh seeded model for each candidate. Change one factor at a time and select using validation—not repeated official-test evaluation.
- 64→128 hidden units: parameter count increases50,890→101,770. This increases arithmetic/storage at fixed batch shape; it does not guarantee slower wall time on every device or better accuracy. Compare validation loss, accuracy, training duration and variability over several seeds.
- Add a second64-unit ReLU layer: the count becomes55,050. The extra nonlinearity changes the function family; compare it with the wider single-layer model rather than treating depth as universally superior.
- Remove division by255: keep labels and splits unchanged and record actual input maxima, loss, finite gradients and validation scores. Input scale changes preactivations and optimization behaviour. It does not prove a particular accuracy drop; a learning-rate change could alter the comparison, so name whether that factor was retuned.
- Use the84% callback: inspect the exact metric key, handle absent logs, and identify whether the threshold watches training or validation. Training accuracy crossing a threshold does not establish held-out quality. The default20-epoch recipe and optional callback are alternative experiments, not a promise of reaching84%.
Report errors by class as well as overall accuracy. Shirt versus T-shirt confusion may dominate while easy categories hide it. A majority-class baseline, confusion matrix, per-class counts and inspection of mistaken examples give context. A screenshot of one run is not a controlled comparison; record the code/configuration, seed, split and actual measurements.
Recall and worked answers
- At w=b=0, what is the first SGD update for the scalar example at rate.01? Answer: gradients−34,−11 give w=.34,b=.11; x=10's true reference remains31, not the new model's prediction.
- Why is softmax not the class decision? Answer: it returns all ten normalized probabilities; argmax is a separate selection and neither operation proves calibration.
- How many parameters does Flatten have? Answer: zero; Dense64 and Dense10 together have50,890 in the reference.
- Why not stop when test accuracy exceeds84%? Answer: that uses test outcomes to select a model. Use training for the callback demonstration or validation for model selection and reserve test for the frozen choice.
- A
.kerasfile reloads successfully but gives different logits. What next? Answer: compare preprocessing, label mapping, inference mode, custom-object serialization, versions and tolerances before treating it as the same model. - Why is a first-row pixel maximum an incomplete normalization check? Answer: an unbatched example is a full image; indexing
[0]selects only its first row. Inspect the complete tensor and dtype before/after the fixed transform. - Does a wider network's lower training loss settle the exercise? Answer: no; compare validation, compute and variation, then report an independent final score after selection.
Sources and scope of the evidence
Primary references: TensorFlow image classification, Keras activations, callbacks, saving/loading, Fashion-MNIST authors, Workbench setup, and Cloud Logging integration.
Only the NumPy line/gradient/probability/capacity controls were executed. No Fashion-MNIST score, TensorFlow training, save/load execution, Workbench provisioning, logging delivery or cloud bill is asserted. The complete reference programs and experiment plans retain those learning objectives without inventing outcomes.
Pause / Recall / Apply
Can you explain it without the page?
Close the example. Reconstruct the core idea, then change one assumption. Mark complete when you’re ready; you can always undo it.
Stored in this browser only. No account, no sync. Clearing browser data removes your record.