Keyboard, Pointer, Wheel, and Composite Actions API: Guided Hands-On Workflow
Operate one disposable Actions fixture end to end: record provenance, exercise keyboard modifiers and shortcuts, hover and pointer-hold movement, scroll with the wheel source, inspect event/focus state after each step, and clean up virtual input state.
Learning objectives
- Create a reproducible loopback fixture that records keyboard, pointer, and scroll events without real accounts or external services.
- Use platform-aware Command/Control modifiers and verify focus plus application shortcut state.
- Exercise hover, click-and-hold, pointer movement, release, and wheel scrolling through semantic element anchors.
- Build one composite sequence incrementally and distinguish queued action state from performed browser state.
- Capture session/capability, DOM, event-log, screenshot, and timing evidence after meaningful state changes.
- Release virtual input state and quit the browser deterministically.
1. Scenario, assumptions, and state boundaries
The fixture is a single static local page. It exposes a text field and shortcut, hover menu, pointer-hold target, long viewport, and event log. All data is synthetic. The browser profile is the temporary automation profile owned by the WebDriver session; there are no credentials, cookies, external APIs, downloads, or production systems.
http://127.0.0.1:8771 only. If the URL is not loopback,
the workflow aborts before creating browser-side mutations.
Version assumptions: Selenium Python 4.47.0, Python 3.10+, one supported local Chromium-family browser. Selenium Manager may resolve the driver. Record the actual browser version and session ID returned at runtime rather than assuming them.
2. Create the disposable project and fixture
The following example makes the Create the disposable project and fixture behavior concrete. Read it with the stated assumptions, then compare its observable output or state changes with the explanation that follows.
mkdir selenium-ch07-actions
cd selenium-ch07-actions
python -m venv .venv
# Linux/macOS:
. .venv/bin/activate
# Windows PowerShell:
# .\.venv\Scripts\Activate.ps1
python -m pip install --upgrade pip
python -m pip install "selenium==4.47.0"
mkdir site evidence
The following example makes the Create the disposable project and fixture behavior concrete. Read it with the stated assumptions, then compare its observable output or state changes with the explanation that follows.
<!doctype html>
<html lang="en">
<head><meta charset="utf-8"><title>Chapter 07 Actions Lab</title>
<style>
body{font:16px system-ui;max-width:840px;margin:30px auto;padding:0 20px}.panel{border:1px solid #888;padding:16px;margin:18px 0}.menu{display:inline-block;position:relative;padding:10px;border:1px solid #777}.submenu{display:none;position:absolute;top:100%;left:0;background:white;color:black;border:1px solid #777;padding:8px}.menu.hovered .submenu{display:block}.token{display:inline-block;padding:12px;background:#def;border:1px solid #478;user-select:none}.drop{min-height:60px;border:2px dashed #777;padding:14px;margin-top:12px}.drop.active{border-style:solid}.spacer{height:780px;background:linear-gradient(#fff,#eee)}#event-log{font:13px ui-monospace,monospace;max-height:220px;overflow:auto}button,input{font:inherit;padding:8px}
</style></head>
<body data-build="ch07-actions-v1">
<h1>Actions laboratory</h1>
<div class="panel"><label>Notes <input id="notes" autocomplete="off"></label><button id="commit">Commit</button><p id="keyboard-status">Keyboard idle</p></div>
<div class="panel"><div id="menu" class="menu" tabindex="0">Hover menu<div id="submenu" class="submenu"><button id="menu-action">Menu action</button></div></div><p id="hover-status">Not hovered</p></div>
<div class="panel"><div id="token" class="token" tabindex="0">Synthetic token</div><div id="dropzone" class="drop" tabindex="0">Drop zone</div><p id="drag-status">Not moved</p></div>
<div class="spacer"></div>
<div class="panel" id="wheel-marker" tabindex="0"><h2>Wheel destination</h2><p id="scroll-status">No wheel/scroll evidence yet</p></div>
<div class="panel"><h2>Event evidence</h2><ol id="event-log"></ol></div>
<script>
const log=(m)=>{const li=document.createElement('li');li.textContent=`${performance.now().toFixed(1)} ${m}`;document.querySelector('#event-log').append(li)};
const notes=document.querySelector('#notes'); const ks=document.querySelector('#keyboard-status');
document.addEventListener('keydown',e=>{log(`keydown:${e.key}:ctrl=${e.ctrlKey}:meta=${e.metaKey}:shift=${e.shiftKey}`);if((e.ctrlKey||e.metaKey)&&e.key.toLowerCase()==='k'){e.preventDefault();ks.textContent='Shortcut accepted';notes.focus();}});
document.addEventListener('keyup',e=>log(`keyup:${e.key}`));
document.querySelector('#commit').addEventListener('click',()=>{ks.textContent=`Committed:${notes.value}`;log(`commit:${notes.value}`)});
const menu=document.querySelector('#menu');menu.addEventListener('pointerenter',()=>{menu.classList.add('hovered');document.querySelector('#hover-status').textContent='Hovered';log('pointerenter:menu')});menu.addEventListener('pointerleave',()=>{menu.classList.remove('hovered');log('pointerleave:menu')});
const token=document.querySelector('#token'), drop=document.querySelector('#dropzone');let held=false;
token.addEventListener('pointerdown',()=>{held=true;log('pointerdown:token')});drop.addEventListener('pointerenter',()=>{if(held)drop.classList.add('active');log('pointerenter:drop')});drop.addEventListener('pointerup',()=>{if(held){document.querySelector('#drag-status').textContent='Moved by pointer sequence';log('pointerup:drop');}held=false;drop.classList.remove('active')});document.addEventListener('pointerup',()=>held=false);
window.addEventListener('scroll',()=>{document.querySelector('#scroll-status').textContent=`scrollY:${Math.round(scrollY)}`;log(`scroll:${Math.round(scrollY)}`)},{passive:true});
</script></body></html>
Save the HTML as site/index.html. Start it in a second
terminal:
python -m http.server 8771 --bind 127.0.0.1 --directory site
Preflight with a browser or
curl http://127.0.0.1:8771/. The page title should be
Chapter 07 Actions Lab.
3. Run the complete observable workflow
The following example makes the Run the complete observable workflow behavior concrete. Read it with the stated assumptions, then compare its observable output or state changes with the explanation that follows.
import json
import sys
from pathlib import Path
from urllib.parse import urlparse
import selenium
from selenium import webdriver
from selenium.webdriver.common.action_chains import ActionChains
from selenium.webdriver.common.by import By
from selenium.webdriver.common.keys import Keys
BASE = "http://127.0.0.1:8771/"
EVIDENCE = Path("evidence")
EVIDENCE.mkdir(exist_ok=True)
def require_loopback(url: str) -> None:
host = (urlparse(url).hostname or "").lower()
if host not in {"127.0.0.1", "localhost", "::1"}:
raise RuntimeError(f"Refusing non-loopback target: {host}")
def event_text(driver):
return [x.text for x in driver.find_elements(By.CSS_SELECTOR, "#event-log li")]
require_loopback(BASE)
driver = webdriver.Chrome()
actions = ActionChains(driver)
try:
driver.get(BASE)
assert driver.title == "Chapter 07 Actions Lab"
caps = driver.capabilities
print("selenium", selenium.__version__)
print("session", driver.session_id)
print("browser", caps.get("browserName"), caps.get("browserVersion"))
print("platform", caps.get("platformName"))
notes = driver.find_element(By.ID, "notes")
print("before focus", driver.switch_to.active_element.get_attribute("id"))
# 1) Keyboard shortcut: Command on macOS, Control elsewhere.
modifier = Keys.COMMAND if sys.platform == "darwin" else Keys.CONTROL
ActionChains(driver).key_down(modifier).send_keys("k").key_up(modifier).perform()
assert driver.find_element(By.ID, "keyboard-status").text == "Shortcut accepted"
assert driver.switch_to.active_element.get_attribute("id") == "notes"
# 2) Send text to the focused field and commit through a pointer click.
ActionChains(driver).send_keys("actions-demo").perform()
driver.find_element(By.ID, "commit").click()
assert driver.find_element(By.ID, "keyboard-status").text == "Committed:actions-demo"
# 3) Hover by moving to an element semantic anchor.
menu = driver.find_element(By.ID, "menu")
ActionChains(driver).move_to_element(menu).perform()
assert driver.find_element(By.ID, "hover-status").text == "Hovered"
assert driver.find_element(By.ID, "submenu").is_displayed()
# 4) Click-and-hold, move to target, release.
token = driver.find_element(By.ID, "token")
drop = driver.find_element(By.ID, "dropzone")
ActionChains(driver).click_and_hold(token).move_to_element(drop).release().perform()
assert driver.find_element(By.ID, "drag-status").text == "Moved by pointer sequence"
# 5) Wheel source: scroll to a semantic element, not an OS screen coordinate.
marker = driver.find_element(By.ID, "wheel-marker")
ActionChains(driver).scroll_to_element(marker).perform()
assert driver.find_element(By.ID, "scroll-status").text.startswith("scrollY:")
# 6) One composite queue: focus notes, hold Shift while typing, release, then click Commit.
composite = ActionChains(driver)
composite.click(notes)
composite.key_down(Keys.SHIFT)
composite.send_keys("abc")
composite.key_up(Keys.SHIFT)
composite.click(driver.find_element(By.ID, "commit"))
composite.perform()
assert notes.get_property("value").endswith("ABC")
packet = {
"selenium": selenium.__version__,
"session_id": driver.session_id,
"browserName": caps.get("browserName"),
"browserVersion": caps.get("browserVersion"),
"platformName": caps.get("platformName"),
"active_element": driver.switch_to.active_element.get_attribute("id"),
"keyboard_status": driver.find_element(By.ID, "keyboard-status").text,
"drag_status": driver.find_element(By.ID, "drag-status").text,
"scroll_status": driver.find_element(By.ID, "scroll-status").text,
"events": event_text(driver)[-30:],
}
(EVIDENCE / "actions-state.json").write_text(json.dumps(packet, indent=2), encoding="utf-8")
driver.save_screenshot(str(EVIDENCE / "actions-final.png"))
finally:
try:
actions.reset_actions()
finally:
driver.quit()
Run with python workflow.py. Each block has a different
state boundary:
| Action | Reads | Changes | Independent verification |
|---|---|---|---|
| Shortcut | current context + virtual key state | modifier state, keyboard events, focus | status text + active element + event log |
| Hover | menu element geometry + pointer state | pointer position, hover CSS/app state | submenu displayed + pointerenter evidence |
| Hold/move/release | source/target geometry + pointer button state | pointer-down state and application drag marker | drag status + pointer event log |
| Wheel | viewport + destination geometry | wheel/scroll state and viewport position | scroll-status + event evidence |
| Composite | focused input + commit button + key/pointer sources | text value, focus, modifier state, click outcome | input property + commit status + event log |
The assertions do not claim that event micro-order is identical on every browser. They verify stable application outcomes and key evidence points that matter to the fixture contract.
4. Build composite input incrementally
Do not begin debugging with a ten-step chain. Prove one state transition at a time. First make focus observable, then add the modifier, then text, then pointer work.
from selenium.webdriver.common.action_chains import ActionChains
from selenium.webdriver.common.keys import Keys
notes = driver.find_element(By.ID, "notes")
commit = driver.find_element(By.ID, "commit")
# Step A: establish focus.
ActionChains(driver).click(notes).perform()
assert driver.switch_to.active_element.get_attribute("id") == "notes"
# Step B: add balanced modifier state.
ActionChains(driver).key_down(Keys.SHIFT).send_keys("abc").key_up(Keys.SHIFT).perform()
assert notes.get_property("value").endswith("ABC")
# Step C: only now combine pointer + keyboard intent.
ActionChains(driver).click(notes).send_keys("-ok").click(commit).perform()
This sequence makes failure attribution sharper. If Step A fails, the problem is not a modifier. If Step B fails, the click target was already proven. Composite tests should preserve this debugging decomposition even if the final production helper composes the actions.
5. Platform-aware modifier keys are part of the test contract
Copy/paste and many shortcuts use Command on macOS and Control on other desktop platforms. Selenium’s current keyboard documentation demonstrates exactly this distinction. A test that hard-codes Control and then labels macOS failure “flaky” has encoded the wrong platform contract.
import sys
from selenium.webdriver.common.keys import Keys
PRIMARY_MODIFIER = Keys.COMMAND if sys.platform == "darwin" else Keys.CONTROL
print("primary modifier", "COMMAND" if sys.platform == "darwin" else "CONTROL")
For remote Grid execution, the test-runner operating system and the browser node platform can differ. Later distributed-testing chapters will make the node capability the stronger source of truth. In this mandatory local exercise runner and browser share the host.
6. Inspect the evidence packet
evidence/actions-state.json should contain the Selenium
version, session identity, browser/platform capabilities,
active-element identity, application status values, and the tail of
the fixture event log. actions-final.png captures the
final visible state.
Do not log real keystrokes from real accounts in production suites. Event logging here is safe because the fixture contains only synthetic text. Keyboard/network logs can become sensitive evidence in real systems and require redaction/retention controls.
7. Challenge: choose the smallest correct control
Without copying a provided sequence, decide how you would automate each requirement:
- Activate a plain button once.
- Open a submenu that appears only after pointer hover.
- Type
ABCby keeping Shift depressed. - Bring an off-screen marker into view using the wheel input source.
- Verify a keyboard shortcut changed focus.
A good answer uses element.click() for the first
requirement, Actions for hover/modifier/wheel behavior, and an
active-element/application assertion for the final requirement. The
selection criterion is behavioral intent, not API novelty.
8. Cleanup and rollback
The following example makes the Cleanup and rollback behavior concrete. Read it with the stated assumptions, then compare its observable output or state changes with the explanation that follows.
# Stop the loopback http.server with Ctrl+C in its terminal.
# After reviewing the synthetic evidence:
cd ..
rm -rf selenium-ch07-actions
On Windows PowerShell, remove only the known disposable directory
with
Remove-Item -Recurse -Force .\selenium-ch07-actions
after changing to its parent. Do not delete normal browser profiles
or Selenium Manager caches as “cleanup”; this lab did not create
them as course-owned resources.
9. Summary and next step
You have now exercised all three practical input sources and proved their outcomes through focus, DOM state, application status, event evidence, and a screenshot. The next design problem is deciding when this power improves realism and when it creates avoidable layout/platform coupling.
Knowledge check
Why does the workflow choose Command on macOS and Control elsewhere?
The user-level shortcut modifier is platform-specific; Selenium’s own keyboard examples make that distinction. Hard-coding one modifier creates an incorrect cross-platform test contract.
What does scroll_to_element() improve compared
with a hard-coded desktop coordinate?
The destination is a WebElement semantic anchor resolved inside the browser viewport rather than an operating-system screen location.
Why inspect driver.switch_to.active_element after
the shortcut?
The intended shortcut outcome includes focus transfer; merely observing that key commands were sent does not prove the user-visible behavior occurred.
Why is the event log safe here but potentially sensitive in a real suite?
The fixture contains synthetic input only. Real keystroke/network/event evidence may contain credentials, personal data, or business content and needs minimization/redaction.
A hover submenu never appears. What should you check before retrying?
Confirm the current context, target identity/geometry/display state, pointer event evidence, browser/platform provenance, and whether the application actually received pointerenter. Retrying alone does not identify the failing layer.
Official references and version notes
- Selenium 4.47 release notes — current stable release baseline used for this chapter.
- Selenium downloads — current stable binding and Grid versions.
- Selenium Actions API — virtualized key, pointer, and wheel input-source model.
- Keyboard actions — key down/up and modifier behavior, including platform-specific Command versus Control examples.
- Mouse actions — pointer move, click-and-hold, release, and drag-like interaction patterns.
-
Selenium Python 4.47 ActionChains API
— queued actions,
perform(),reset_actions(), wheel methods, and pointer methods. - W3C WebDriver 2 Actions — input sources, input state, ticks, Perform Actions, and Release Actions protocol semantics.
Version-sensitive behavior was rechecked against current primary documentation on 2026-08-28. Mandatory examples pin Selenium Python 4.47.0, require Python 3.10+, use a supported local Chromium-family browser with the actual browser/driver/session provenance recorded at runtime, and use only loopback fixtures. Wheel and detailed pointer behavior can differ at browser/platform edges; assertions therefore target application/focus/event state rather than incidental pixel coordinates. The examples intentionally avoid low-level private ActionBuilder internals except where public documentation is referenced; normal teaching uses public ActionChains conveniences.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.