# Workroom Productions > Exercises, toys, ideas and consultancy from James Lyndsay / Workroom Productions / @workroomprds Public Ghost content for AI and LLM tooling. This file includes a bounded export of public pages first, then recent public posts. Append `.md` to any post or page URL to get the content in Markdown (for example, `/example-post.md`). ## Pages ### All About Everything URL: https://www.workroom-productions.com/about/ Last updated: 2025-06-12T09:28:16.000Z Workroom Productions is a small, London based consultancy, typically providing services to help organisations get better at software testing – revealing your options, firming up your strategies, polishing your people. There's only me, James Lyndsay: when you arrange for something to be done, you'll always get the principal consultant. ### James Lyndsay - brief bio I had my first job, and my first job testing, in 1986: working at IBM in a small team on an x86 graphics processor chip. I was handy at testing, and way out of my depth at everything else. I set up Workroom Productions in 1994\. I've worked for myself, as a systems tester, since then. I've resisted specialising in a particular technology or market, and I've tried to change with my industry. My clients include retail and telecommunications, banking and social media, rapidly-evolving internet start-ups, traditional large-scale enterprise, and government bodies. I've worked as an innovator (loved it), and as a standardiser (less so). I work as a consultant with an eye for strategy, quality and motivation. I've stepped into roles as manager, mentor, advocate and implementer – with a particular focus on tooled-up exploratory testing. I've worked internationally as a consultant and teacher since 2000, and I've been trusted to deliver keynotes and hands-on workshops at the largest and most interesting conferences. I've won prizes for several papers, and was the 2015 recipient of [EuroSTAR's Testing Excellence](https://conference.eurostarsoftwaretesting.com/eurostar-testing-excellence-award/) award. I have an MA in Physics from Trinity College, Oxford. I've studied experiential teaching with [Jerry Weinberg](https://en.wikipedia.org/wiki/Gerald%5FWeinberg), Impro with [Keith Johnstone](https://www.keithjohnstone.com), and writing with [Joel Morris](https://substack.com/@joelmorris?utm%5Fsource=about-page). I sing with the [London Bulgarian Choir](https://www.londonbulgarianchoir.co.uk), and co-produced all three of their albums. You can see less about me at [LinkedIn](http://uk.linkedin.com/in/jameslyndsay/). Contact me for a CV, if you must. ![James Lyndsay headshot](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/2021/06/James-headshot.jpg) ### Workroom Productions - Services **Consultancy:** Facilitated discussions leading to test strategy, analysis of existing methods and approaches, process improvement, tools acquisition, analysis of key bugs, expert perspective. **Knowledge transfer:** A range of courses in Exploratory Testing, Experimentation and Diagnosis, Session-based Test Management, and Test Data.Talks, workshops and seminars in a wide-ranging group of test-related subjects. Direct coaching in test techniques and test management from within a team. **Testing:** Test management, hands-on exploration of your product or system, test script design and automated testing, test-focussed work on agile projects, great bug logs. ### Endorsements You'll need to take these with a grain of salt - but give me a call, or meet me at a conference, and make up your own mind. **From lectures and conferences** > Workshop and performance were top-notch, with an insightful session, a well-structured agenda, and impressive projects. ***Conference participant, interactive workshop*** > Well prepared, detailed instructions and background info on a website - for being used in the workshop but also later on to continue practicing. Unmatchable presentation style. Need more of this! ***Conference participant, interactive workshop*** > very insightful and super well-organised. ***Conference participant, interactive workshop*** > Really well structured and executed workshop. ***Conference participant, interactive workshop*** > exceeded my high expectations! The clear instructions were helpful and the exercises were fun. I learned a lot. ***Conference participant, interactive workshop*** > "James exudes subject matter credibility, and was well received by our Testing community. His Agile session was particularly popular, and appealed to both the less and the more experienced members of the audience." ***David Moore, Test Practice Lead, Aviva Life*** > "Outstanding presentation... easily one of the best at conference." ***Conference participant, talk*** > "Great presenter, good stories, excellent presenting skills and timing, good humor." ***Conference participant, talk*** **From people who have come to my classes** > "I attended James Lyndsay's 2-day exploratory testing tutorial and it was very useful and fun. If you expect loads of theory to be thrown at you - don't bother. But if you want lots of hands-on exercises, eye-openers and discussions, combined with a refreshing view on testing, this is the course for you. On top of that, James is a very engaging 'teacher', and an allround nice guy :-)" ***Zeger Van Hese - Senior Test Manager*** > "Since real life projects push us to optimise testing to do as much as possible in the same amount of time, Exploratory Testing seems a natural step forward. James Lyndsay's class gave me plenty of practical tips and tricks about how to manage and execute Exploratory Testing. And, most importantly, how to measure improvements and catch the 'stuck moments'. No optimization is good unless it really gives results. Thank you for all the shared experiences and practical knowledge that allowed me to implement ET right after I took the class ­ and for the tips when to and when not to use ET. Taking part in your class was a very pleasant and useful experience!" ***Mira Razman, Senior Test Lead*** > "The interactive tools really demonstrated the different approaches and considerations required for each of the techniques we were taught. Without the practical examples, application of the theory in the future would have been far more difficult." ***ET class participant*** **From consulting clients** > "...innovative and flexible approach is combined with strong testing and management skills to provide a complete solution, allowing us to concentrate on core development. Workroom Productions has been an effective and valuable resource to our company." ***Derek Clark, CEO, Chasseral Ltd.*** > "...provided a cost effective service that has added real value... significantly strengthened the quality of our delivered solutions... a highly professional consulting company with an impressive depth of technical knowledge and experience" **Jim Sutton, Operations Manager, Three X Communication Ltd*.* > "...epitomises a professional's approach to testing. He has an ability to take subjective assessments and where possible make them objective and measurable" **Stuart Ayling, CIO, Securicor Omega Logistics** **About my exercises** > "James Lyndsay's testing machines are one of the best tools available to teach the actual process of testing software. There are too many touch points to list for applied critical thinking, test planning, test execution and observation, and reporting. > > I couldn't recommend them enough for training software testers how to think about their craft, the skills required to do it well, and hands on problem solving." **Keith Klain*, Director – Quality Engineering, KPMG.*** ***About mentoring*** > Your guidance has been instrumental in my professional growth, and I truly appreciate the time and effort you invested in helping me develop my skills and confidence. From the very beginning, you instilled in me the confidence to take on various tasks, ensuring I had the knowledge and support needed to perform effectively. Your expertise in testing activities, from test documentation creation to execution strategies and stakeholder engagement, has been invaluable in shaping my approach to problem-solving and decision-making. The weekly check-ins and knowledge-sharing sessions you organised provided a strong foundation for continuous learning and growth. \[...\] You have continued to be a source of support and accountability, helping me stay focused on my professional development goals. Your adaptability and willingness to offer guidance, no matter the challenges I brought to you, have made a significant impact on my journey. \[...\] Your mentorship has been an essential part of my development. **Melvin Kamau, SDET** > *Working with James has been an invaluable experience. His enthusiasm and expertise in software testing have shone throughout our engagement. James is a master at creating space for honest, productive discussions. He listens carefully and asks the tough questions that lead straight to the heart of the problem—all in a thoughtful and systematic way. On top of that, his positive energy makes solving complex issues actually fun. Highly recommended for anyone seeking expert guidance.* **Giorgos *Siamantas, Software Test Engineer*** ### Clients **Consulting and customised training:** QA Consulting, Bank of England, Games Workshop, Channel 4, Qentinel, Nokia, BCS/ISEB, Google, T-Mobile, Thomson Financial, Transition Consulting, GE Mobile Communications (3X), Legal Services Commission, Disney, Securicor Omega Logistics, Exel Logistics, WorkShare, Test Partners, Harrods, Technology Evaluation Centres, ProfessionalSpirit, Chasseral, JobPartners. **Participants in public training courses have included staff from:** BBC, TCL, Marks&Spencer, Thoughtworks, GE Capital, 3X, Expedia, Microsoft, Research Machines, Mozilla Foundation. **Conferences and talks:** Agile Testing Days, Ministry of Test, Let's Test, StarEast, StarWest, EuroStar, UKSTAR, Google TechTalks, London SIGiST, Nordic Testing Days, TestingCup, QBIT, Aviva/Norwich Union, Ordina, Delta, Avenir, DANSK-IT, IIR, FAST, Tampere University of Technology, Unicom, SoftTest, ScotTest, SAST, AsiaSTAR, ProcessWorks Singapore, Quality Week, QW Europe. ### Community Testing is a young discipline, in a fast-moving world. I've learnt a lot from my peers and colleagues, and I try to put something back in when I can. My involvement - sometimes paid, sometimes free, sometimes contributing money as well as time - has included: - Paper 'Four exercises for teaching Exploratory Testing', with associated software and teaching notes, placed into the **public domain** for use by the community. - Most papers put into the **public domain**. Recordings of most associated talks available on this website. - Contributing participant in **peer conferences**: LEWT, WOPR/SWOPR, WHET, ExTRS, WTST, LAWST, AA-FTT. - Facilitator / convenor / founder of **LEWT**. - **CAST** sponsor: 2006, 2008 - Speaker at community conferences: **SIGiST**, **SoftTest**, **ScotTest** \- contributing talks, tutorials, and special sessions. - **Open-source** tools for working with session-based testing available through this website. - **ISEB** Examiner / question designer for practitioner certificate in Software Testing. Question designer for foundation certificate in Software Testing. Contributor to Test Analyst syllabus for Software Testing Practitioner replacement. Member of ISEB Software Test Steering panel. *Note: having worked with ISEB/BCS from 2002-2007, I've finished this role.* ### Public information for Workroom Productions Established 22 July 1994, Company no. 2951389\. Registered for VAT with number GB653962802\. Registered office is currently 36 Appleshaw House, London SE5 8DW. Director: James Lyndsay ### Consultancy URL: https://www.workroom-productions.com/consultancy/ Last updated: 2025-02-26T23:26:24.000Z You need someone to advise, analyse, advocate and act. I can do that. I'm a Test Strategist. I see the big picture, and how the details matter. I'll help you and your team to understand your shared values, so that you can work in concert towards an achievable future. I believe that good testing provides vital information to those building and using a system. Good testing can provide that information swiftly, in a way that aims for relevance and which can demonstrate evidence. You'll get the most use from me if you want the truth about the systems you're delivering. You'd get someone recognised by the testing industry, with hands-on experience in banking, telecoms, utilities, retail and more. I'm good with tools, good with people, good with crowds – and I'm as happy talking with the board as I am talking with new hires. I've worked with startups, vast institutions, blue-chips and unicorns. Your organisation and situation are unique, and I've seen enough uniqueness that I can cope. I'm expensive and valuable, expert and discreet, available swiftly and only as much as needed. [Find a time to chat (Savvycal)](https://savvycal.com/workroomprds/chat) --- ### How you might use me to help your Organisation - You could ask me to check over particular products, projects and teams, to judge how they are working from a testing point of view. - You could share my time with your suppliers and clients, with the aim of finding ways to help everyone's testing serve the overall needs. - You could resolve sticking points by engaging me to work with a disparate group to clarify how each views their testing needs and constraints. - You could ask me to help your decision-makers to understand what's happening in testing, the capabilities of the testers, how to translate organisational values into testing actions, and to identify how the information flows from testing into key decisions. - You could ask me to work with your senior testing team on their structures, strategies, tooling and tactics, seeking to achieve coherence in their approach to software quality. ### How you might use me to help your Relationships - You could demonstrate your commitment to your client by bringing me in as a temporary expert resource, especially around exploratory testing. - You could engage more closely with your supplier's people by share a training course, or resolve difficulties by asking me to provide an independent assessment. - You can grow connections with peer teams in your ourganisation with a shared testing workshop. - You can build new connections with teams in your wider organisation by sending out a general invitation to a lunch-and-learn on test strategy, test perspectives or bug diagnosis. ### How you might use me to help your Product - You could ask me to work with your product owners and business analysts to discover ways that they can use testing work and budget to get to higher quality, or swifter time to market. - You could ask me to work with your key users to identify what matters to them about product quality and functionality, and how they can help your developers to deliver those characteristics. ### How you might use me to help your Project - You could ask me to judge whether the testing actions on the project match the values of those planning testing work, and whether both together help the project to deliver value on schedule, or (perhaps) what they hinder. - You could ask me to coach or teach key staff on the project in testing skills, aligned with overall project values. ### How you might use me to help your Team - You could ask me to teach your team to explore, investigate and diagnose. - You could ask me to work with a high-performing team to see what makes them that way, what be successfully transferred, and how to help the transferred approaches and skills to stick, and how to assess if it's working. - You could ask me to assist a troubled team to identify the most-acheiveable ways to adjust their work to better fit their situation. ### How you might ask me to help your Individuals - You could ask me to help your new test manager to get to grips with the role. - You could ask me to work one-to-one with your testers as they explore. ### Mentoring URL: https://www.workroom-productions.com/mentoring/ Last updated: 2025-11-04T11:48:17.000Z You need someone to guide your testers and test managers. I can help. ### Teaching URL: https://www.workroom-productions.com/teaching/ Last updated: 2025-11-04T11:51:01.000Z You need to build skills in your test team. Skills like Exploratory Testing, Software Diagnosis, Bulk Testing, Use of Tools, Building Test Data. I can help. I have built and delivered interactive testing workshops since 2001\. I build exercises and tools for other trainers. My teaching clients include Google, Oracle, the BBC, Cassidian, Reuters, Adobe, Cisco, GE, Aviva. I've delivered paid full-day tutorials at EuroSTAR, Agile Testing Days, Nordic Testing Days, Let's Test, STAREast, STARWest and more. My workshops are hands-on, and based on my experiences as a tester and as a consultant – they keep people engaged, and use practice to introduce core transferrable testing concepts. My core workshop is *Insights into Exploratory Testing* – I offer the workshop in different sizes and with different content to suit your needs, and can deliver online or in person. I have material from several other workshops; on data, on bulk testing, on management, tool use, collaboration and much more. [Find a time to chat (Savvycal)](https://savvycal.com/workroomprds/chat) ## Short Sessions Book me for an hour with your team; I'll adjust my materials to your people, deliver an interactive and engaging online workshop, and leave you with something to play with. **£490 / €590** for one 40-45 minute online workshop with 15-20 minutes Q&A, for up to 25 people. You and your team will have access to a page for your workshop with all the materials in one place. On the day, I typically use Zoom or gather.town to allow small groups to work together, and enable shared workspaces with Miro. ## Asynch Sessions My workshops tend to be experiential, and you might not need me to be involved. If you have a team who want to learn by themselves, then I'll give them all the materials they need. They'll spend time playing and exploring together at times that suit them, over a few weeks. I'll run a 1-hour facilitated discussion when your team are ready. **£1190 / €1390** for materials to cover 2-3 hours self-directed learning followed by an hour's facilitated conversation online. ## Deeper Content – online or on-site I'll teach your team online, working with them for a series of sessions over weeks, timed to fit around their work, and allowing the team to develop ideas and practices of their own, and to review them with me as we progress. I provide videos and other materials for study before a session, and follow-up materials after the session for those who wish to go further. Class size 5-15 people. **£2750** for six one-hour sessions (or a day on-site teaching). **£4900** for 12 x 1-hour group sessions (or two days onsite). ## On-site Workshops Bring your team together in person for a day or two's teaching, with similar hands-on exercises, and even more opportunities for conversation and connection. Workshop prices match online costs, with an add on for time away. ## Content I frequently customise my materials to suit my client's interests. Here are partial lists of conference / public [talks](http://oldsite.workroom-productions.com/talks.html) and [workshops](http://oldsite.workroom-productions.com/training.html) – fuller details to follow. ❕ ****Prices exclude VAT (and travel if applicable).** ## More questions... Unfold answers - mainly older answers for on-site workshops... --- - *Do you still do in-person workshops?* > Sure. Do **you** do onsite workshops? Circumstances may mean we can't plan as easily as we might. Let's talk. - *Are your workshops just for testers?* > My exercises have different things to find for people with different skillsets. A more diverse group allows participants to take away more surprising ideas. In a broad group, we’d have managers, coders, systems architects, ops, business owners and key customers alongside the people labelled as testers. I often have a broader group when I go back to a site for a second visit. - *Can we have people from outside our organisation?* > Some clients like to share my workshops with suppliers and contractors. That’s fine. Others want to teach a few of their staff, and sell seats to the public. That’s generally fine, too. - *What does your workbook look like?* > The printed one is pretty, with lots of space for notes. I'll send you a sample on request. Online, we use this site, Miro or similar, Zoom or similar, and my exercises. - *Can we teach internally with your exercises?* > I offer four exercises under Creative Commons for non-commercial teaching. If you want to teach internally, we typically allow use of those, and negotiate paid licenses for others. - *Do you travel?* > I’m in London. I’ll try get to your office and back in a day. If you’re close to a mainline station or airport and not much further than Manchester, Bristol or Amsterdam, that’ll work. Otherwise, I tend to travel out the day before, and try to get back to London in the evening of the last day of the workshop. If you want me further afield, that’s fine – we’ll sort out logistics. And, if you’re flexible with dates, I may be able to visit your office from somewhere closer than London. I'll pass on travel and accomodation at cost, and if you need me to be on the road a lot, I'll add a per-diem. - *What are your typical travel expenses?* > I generally waive travel if you’re in Central London. For workshops in the South-East, travel typically costs under £70 per day. For longer trips, I often need to stay a night, and have to take a cab as well as a train, so travel is around £150, and accommodation around £130 / night. For workshops in Europe, travel is typically around £250, and accommodation around £150/night. - What is your per-diem? > Depends on distance, time away, location – and on my own costs for being away on the dates you need. Give me specifics and I'll quote for those. - *What size of group?* > The workshops go best with five or more. If we’re doing a coached workshop, where I get to work with everyone directly, I keep numbers to 12 or fewer. If you need me to teach more, I can adjust many exercises for small groups. For conferences, I’ll often have 40-80, which brings different challenges. I have a few exercises that I’ve run with more than 100. - *What are the preparatory exercises?* > Short videos on a particular topic, with a bunch of open questions. Participants may want to work through these on their own, or they might be used in a team gathering to provoke thought and stimulate engagement. - *How do you customise?* > We talk, you tell me what your organisation and team wants out of the workshop, and we discuss potential subjects and emphasis. I gauge your interest, propose topics and exercises, and link them together. I’ve got 8-10 days of materials to choose from. - *Can you customise to use our own software?* > Yes. I’ll need to get to understand what your target system is used for, and then I’ll need to explore and test it for a couple of days. I’ll charge for my time. I’ll share any bugs, adapt some of my exercises and may build new ones. Sometimes the bugs alone are worth the cost of customisation. Clients find this particuarly useful if they want to introduce greater variety, to compare themselves with my appraoches, or to create a closer engagement between the workshop and their customer’s needs. - *Do we own the exercises?* > If you pay me for an exclusive exercise, yes. If you don’t, no. I do license exercises for teaching in-house, and can “train the trainer” if needed. - *Will you teach us about penetration testing and security* > Not directly – it’s technology-specific. I find that groups do better teaching themselves by looking into the typical faults and attack vectors in the technologies that their systems use, and trying to discover ways that those can be used in their dicovery work. - *Will you teach us about tool X* > Let’s be clear: **Exploratory testing without tools is weak and slow**, but this workshop is about how to do exploratory testing, rather than how to pick up and use any specific tool. We’ll talk about types and purposes of tools, we’ll probably have a tools workshop, and we’ll be resourceful in considering and using the tools we have to hand. - *What do you need in the room?* > A flipchart or whiteboard (preferably both), a projector, internet access. Powersockets for everyone. - *What arrangement of tables?* > I strongly prefer to have the group arranged so that everyone can see each other – in a U-shape is great, round a big table is fine, rows of classroom desks are rubbish. - *What do you need beforehand?* > I want to have at least one conversation about what the organisation wants from the workshop. I put that information into a proposal for a linked set of exercises and topics, and we’ll work together until the proposal satisfies you and I. Not long before the workshop, I survey participants to find out what they want, so that I can tune the content and emphasis for the individual participants. It’s good to have everyone’s names. - *What equipment do our participants need?* > At least one laptop between two. Handheld devices aren’t great for testing, and desktop machines get in the way of conversations. Most of my software exercises run in the browser – for compatibility, rather than technology. For some legacy exercises, participant machines need to be able to run Flash (!). Laptops need access to the internet. I try to avoid anything that installs software on participant laptops. - *How do you adjust to teach people who aren’t native English speakers?* > Most of my workshops are in Europe. I’ve learned to speak more slowly, to use clear and simple language, and to listen carefully to my workshop participants. - *Will you sign our NDA?* > Generally. - *What about cancallation?* > Cancellation terms are in my proposal. - *Can we record your workshop?* > No. ### Hiring URL: https://www.workroom-productions.com/hiring/ Last updated: 2026-07-31T21:43:56.000Z Workroom Productions is a one-man company, so you'll always be dealing with me, James Lyndsay. When you book me, you get me. Speaking and teaching engagements tend to be booked several months in advance, and are often outside the UK. In some circumstances, I can extend trips - tell me where you are, and I'll let you know when I'm likely to be nearby. **Consultancy**: flexible arrangements, from a half-day at a time, to several days a week. Price dependent on work. Booking direct will typically be a faction of the price that a large consultancy would charge for my time. Fixed priced may be offered for some work. **Testing**: I take on occasional interesting short-term projects, at rates equivalent to a good contractor, and share my tricks with your team. Limited availability, rarely more than 2 days a week. **Training**: An online session is a 40-45 minute workshop with 15-20 minutes Q&A, for up to 25 people. - Short online workshop: **£490 / €590**, - Linked online sessions: **£1450** for three, **£2750** for six, **£4900** for 12. - Asynch online sessions: **£1190 / €1390** for materials to cover 2-3 hours self-directed learning followed by an hour's facilitated conversation online. - In-person 2-day workshops for a class of 4-12 people start at **£4900**, three days **£7000**. I offer one-day workshops from **£2750,** half-day classes from **£1500**. See [courses](https://workroom-productions.com/training.html). **Mentoring**: Six months unlimited phone access for one person: **£6000**. Very limited availability. More structured team mentoring is available. **Speaking**: One hour including Q+A, **£1500** for a room of up to 300 people. See [talks](https://workroom-productions.com/talks.html) for content. How do you know you'll get value for money? Jerry Weinberg (as ever) [puts it best](http://secretsofconsulting.blogspot.com/2008/04/consultants-money-back-guarantee.html). Here's my summary: **If you don't get value for money, we'll adjust our relationship until you do. I don't take on work where I can't add more value than my fee - if I make that mistake, you won't be paying for it.** Note: Prices exclude VAT and expenses. I add travel at cost. For workshops that need me to travel outside London, add £100\. If I need to travel on a different day, then travel + accommodation + £200 per travel day. ### LEWT URL: https://www.workroom-productions.com/lewt/ Last updated: 2025-11-04T11:51:23.000Z **The London Exploratory Workshop in Testing** LEWT is an exploratory peer workshop. We take the view that discussions are more interesting than lectures. We enjoy diverse ideas, and limit some activities in order to work with more ideas. Currently, the workshop is structured as a series of short talks, each followed by a longer discussion. The workshop is one day long. Most participants will make a short presentation, and talks and discussions are time-limited. New talks can be added at any time; participants prioritise the talks as the day goes on. Attendance is by application and invitation. People who have been to the previous LEWT have first claim on seats. Attendees are expected, but not required, to have a brief talk. We have space for two people with less than two year’s experience – they’re not expected to have a talk, but are otherwise full participants. We share the costs of room and food, and no one charges or is paid for their time or expenses. LEWT is run along approximately the same lines as LAWST™, particularly regarding intellectual property and publication. However, LEWT is not LAWST™. A number of the ‘basic format’ guidelines in the introduction to the LAWST™ handbook are superseded by LEWT’s local guidelines. The LAWST™ handbook is currently here: . [AST](http://www.associationforsoftwaretesting.org/) have a [LAWST™ page](http://www.associationforsoftwaretesting.org/drupal/lawst). Major differences between LEWT and LAWST™: - LEWT has many talks, and limits discussion in order to move to the next talk. - LEWT is an exploratory workshop. We aim to improve our understanding by sharing and discussing our experiences. The workshop does not necessarily share LAWST™’s aim to ‘crystallise conclusions, rules or techniques’. - Most LEWT attendees present a talk and answer questions. - We don't have a content owner - the group owns the content. - Some spaces are reserved for testers with less than two years experience. **What's the process?** LEWT has attracted interest within the testing community. This is a brief summary of the preparation for the workshop, and how the day is run. **Before the workshop:** - Everyone submits a title for a ten-minute talk. - We have a project site, where we can discuss ideas before and after the event. - It’s helpful to have abstracts for talks, and a brief biography. **During the workshop:** - All the talks we haven’t yet heard are stuck on a wall. Participants can add new talks at any time. - Everyone gets a limited number of sticky dots. These are votes – you vote for the talks you would most like to hear. Votes may be cast throughout the day. - The day is split into 90-minute sessions. Before each session, and with attention to the vote and the flow of the day, the facilitator choses a group of three talks to be covered in the 90 minutes. - A talk gets thirty minutes. The speaker present his or her ideas for ten minutes at most, preferably less. The rest of the time is spent on questions. When time is up, we move to the next talk. - During the talk, focus any questions on clarification. Leave most questions until the discussion. - The facilitator will handle the question queue during discussions, and keep track of time. **After the workshop:** - We’ll post recordings etc. on the project site. - Papers using ideas from the workshop should acknowledge the workshop and list participants. **Refinements:** - Times may change – typically reducing. - We will go to the next item if there are no more questions. - On request, we can move to the next item before time, or can add five minutes to the end of the questions. These decisions are taken collectively (action and majority are left to the facilitator’s discretion). - You can vote more than once for a talk. - You can vote for your own talk. **LEWT 01 - 25 June 2005, on Exploratory Testing** Robert Sabourin: *How the EAR model can be used to get testing ideas* Alan Richardson: *How to get unstuck and never be stuck again* Jonathan Bach: *Open-Book Testing* Antony Marcano: *A Test Driven Approach to tracking bugs found in Exploratory Testing* Neil Thompson: *Managing Exploratory alongside Scripted: Whether, and How* Marta Gonzalez: *Buggy night: an experiment on introducing ET to a test team* Juha Itkonen: *Exploratory testing case study* Julian Harty: *Exploratory Security Testing* Maaret Pyhajarvi: *Experiences in selling ET to project management* James Lyndsay: *Creativity and Software Testing* Mitchell Goldman: *Hybrid approach: Exploring while checking off Requirements* Richard Durham: *The joys and pains of testing network software* Niel vanEeden: *Risk Based Testing and the risk associated with applying this in real life* In the room: Alan Richardson, Antony Marcano, Dessislava Stefanova, James Lyndsay, Jonathan Bach, Juha Itkonen, Juha-Matti Parmonen, Julian Harty, Keith Olohan, Maaret Pyhajarvi, Marta Gonzalez, Mitchell Goldman, Neil Thompson, Niel vanEeden, Richard Durham, Robert Sabourin, Steve Green **LEWT 02 - 10 December 2005, on Exploratory Testing** Alan Richardson: *Learning how to attack* Juha-Matti Parmonen: *Mindmaps in Exploratory Testing* Marta Gonzalez: *Pair Exploratory Testing: does the back seat driver cause more crashes?* Neil Thompson: *Some personal experience & expl test “structure”* Mitchell Goldman: *New ET Metrics and beyond* Robert Sabourin: *The Taxonomizer* Scott Barber: *Applying Exploratory Testing Techniques to Performance Investigation* Roundtable: *Mind Mapping tools roundtable* Wayne Mallinson: *Exploratory thinking from Geology + Chemistry* Antony Marcano: *XPloratory Testing – XP & ET - natural partners* Jonathan Towler: *Using test automation to create useful system state starting points* James Lyndsay: *Bug rates* In the room: Alan Richardson, Antony Marcano, Danielle Novak, James Lyndsay, Jonathan Towler, Juha-Matti Parmonen, Julian Harty, Keith Olohan, Marta Gonzalez, Mitchell Goldman, Neil Thompson, Richard Durham, Robert Sabourin, Scott Barber, Wayne Mallinson **LEWT 03 - 24 June 2006, on Bugs** Alan Richardson: *Scientific Study of Anomalous Phenomena* Antony Marcano: *To bug, or not to bug, but is that the question?* Jonathan Towler: *On ‘The Wisdom Of Crowds’* Elisabeth Hendrickson: *Bugs I’ve Known* Alan Richardson: *“Bug” – what’s that in your head? – an exercise* Richard Durham: *Are all bugs created equal? (aka what is the value of a bug?)* Neil Thompson: *Bwg oration factors in Bwg Persistence/procreation Networks* Mitchell Goldman: *U-BAD (Universal Bug Analysis Database)* Robert Sabourin: *Finding Bugs That Matter: Bug Quadrants, Scenario Testing and Children’s Books* Marta Gonzalez: *What d’ya mean ‘Catastrophic’?* Mark Garnett: *Compassion Fatigue, Sainsbury’s Patisserie, Constructivism, Bugs and Me…* James Lyndsay: *You find more bugs in a dirty lab* Julian Harty: *Bug Portrait* In the room: Alan Richardson, Antony Marcano, Elisabeth Hendrickson, James Lyndsay, Jonathan Towler, Juha-Matti Parmonen, Julian Harty, Mark Garnett, Marta Gonzalez, Mitchell Goldman, Neil Thompson, Neill McCarthy, Paul Woolston, Richard Durham, Roy Madron, Robert Sabourin **LEWT 04 - 28 July 2007, on Metrics** Mitchell Goldman: *Other Team's Metrics-What To Ask For* Graham Thomas: *New ways of Measuring testing* Alan Richardson: *Metric Modeling Madness - Memories and Misadventures of an ex-Methodology Monster* Richard Durham: *Lessons learned in DDP* James Lyndsay: *Experience report: working without metrics or measurement* Neil Thompson: *Dashboard, Tachometer & Diagnostics* David Fulcher : *Erik Simmons's S Curve Assumptions* Marta Gonzalez: *Comparatively speaking…* In the room: Alan Richardson, David Fulcher, Graham Thomas, James Lyndsay, Julian Harty, Marta Gonzalez, Mitchell Goldman, Neil Thompson, Nick Gregory, Richard Durham **LEWT 05 - 15 December 2007, on Diagnosis** Neil Thompson: *Knowledge of Body, Body of Knowledge* Alan Richardson: *Diagnosis (as a noun) considered dangerous in the testers dictionary* Elisabeth Hendrickson: *Diagnosing Intermittent Bugs* Kevin Shannon: *Diagnosis and the Inexperienced* Robert Sabourin: *Diagnosis - What question should I ask next?* James Lyndsay: *Diagnostic Exercise* Marta Gonzalez: *Diagnosis and credibility* Mitchell Goldman: *Diagnosis for training* Graham Thomas: *Diagnosis follows Failure* Alan Richardson: *ePrime diagnosis* *order / content may not be reliable* In the room: Alan Richardson, Elisabeth Hendrickson, Graham Thomas, James Lyndsay, Juha-Matti Parmonen, Julian Harty, Kevin Shannon, Marta Gonzalez, Mitchell Goldman, Neil Thompson, Robert Sabourin **LEWT 06 - 18 May 2008, open theme** Neil Thompson: *The Schools of Testing: Belief systems? and/or Adaptable?* Robert Sabourin: *Outside a Tester’s Comfort Zone* Richard Durham: *How do testers add value?* Michael Bolton: *"Why didn't you find that bug?"* James Lyndsay: *Exploring a performance dataset with DataGraph (live demo)* Mitchell Goldman: *Endgames and Testing Games* Marta Gonzalez: *Treasure Hunt, Connect-4 and Black Monday* Alan Richardson: *Why I use the term ‘requisite variety’ and how I teach it to my testers* Robert Sabourin: *Sharing Examples of Session-Based Exploratory Testing - a True Story* Neil Thompson: *It’s Risk, Jim, but not as we know it: the game theory of software handovers* Scott Barber: *Using the Set game to teach people how to illustrate multi-dimensional data* In the room: Alan Richardson, Amal Mohammadi, Dessislava Stefanova, James Lyndsay, Maria Szypluk, Marta Gonzalez, Michael Bolton, Mitchell Goldman, Neil Thompson, Richard Durham, Robert Sabourin, Scott Barber **LEWT 07 - 07 December 2008, on *Systems*** Gwen Stewart: *Is testing an open or closed system (where does the energy come from)? And do we act in a way consistent with the dynamics of the system?* Julian Harty: *Possible generic models for testing* Alan Richardson: *A Predictive Teleological Weak Signal Processing Ideomoter Feedback Game (Hellstromism for beginners)* Paul Gerrard : *Using Soft Systems Methodology as a vehicle for constructing/agreeing test strategies* James Lyndsay: *Testing is a “Wicked Problem”* Neil Thompson: *Vicious loops, and tipping points to virtue - a case study* Mitchell Goldman: *Loop Diagram of Session-Based Exploratory Testing process* Graham Thomas: *Viewing Systems as a Whole* Neil Thompson: *Promoting Systems Thinking through the ranks: Private, Corporal, kernel to General* Alan Richardson: *Cybernetic Test Management* In the room: Adrian Prestidge, Alan Richardson, Graham Thomas, Gwen Stewart, James Lyndsay, Julian Harty, Marta Gonzalez, Mitch Goldman, Neil Thompson, Paul Gerrard **LEWT 08 - 26 June 2010, on *Nature*** Talks to follow In the room: Fiona Charles, James Lyndsay, Julian Harty, Marta Gonzalez, Michael Davis, Mitchell Goldman, Neil Thompson, (Noah Goldman), Robert Sabourin, Richard Durham **LEWT 09 - 17 April 2011, on *Large*** In the room: Adrian Rapan, Alan Richardson, Aliaksandr Ikhelis, Anna Baik, Binayak Prasad Silwal, Gwen Stewart, James Lyndsay, Julian Harty, Maaret Pyhajarvi, Markus Deibel, Mitchell Goldman, Nathan Bain, Neil Thompson, Paul Gerrard, Richard Durham, Tony Bruce, Vernon Richards Talks to follow **Logistics** LEWT is currently organised and facilitated by James Lyndsay. We don't lay any claim to originalty in format; if you'd like to organise your own exploratory workshop, you can contact James for advice, but you certainly don't need to ask permission. We'd ask you to call your workshop by a different name. WOPR used the format for three days of pre-WOPR SWOPR events. LEWTs currently happen once or twice a year, in London. ### Test Strategy URL: https://www.workroom-productions.com/strategy/ Last updated: 2025-11-04T11:48:25.000Z To be rationalised... **What is a strategy? Why does testing need one?** A strategy outlines what to plan, and how to plan it. A successful strategy is your guide through change, and provides a firm foundation for ongoing improvement. Unlike a plan, which is obsolete from the point of creation, a strategy reflects the values of an organisation - and remains current and useful. When an organisation tests its products or its tools, it tries to compare them against its expectations and values. By its nature, testing introduces change as problems are identified and resolved. A test strategy is necessary to allow these two impulses to work together. Furthermore, testing can never be said to be 'complete', and a core skill in testing is the justified management of conflicting demands; without a strategy, these judgements will be inconsistent to the point of failure. Software development is a creative process. A test strategy is a vital enabler to this process - keeping focus on core values and consistent decision-making to help achieve desired goals with best use of resource. A good strategy stands as a clear counter to reactive, counter-productive test approaches. **Examples and Templates** A test strategy is not a document. It is a framework for making decisions about value, and has strong links to the unique values of an organisation. It is part of the creative process. Although templates exist, most organisations and projects are poorly served by a one-size-fits-all approach to their specific goals. You may find templates useful on projects where the product to be tested can be created and marketed simply by following templates - but on other projects, they're dangerous. I try to avoid using Test Strategy templates. Instead, I use my skills and experience to rapidly help your team arrive at an understanding of their shared goals and potential conflicts. From this, a strategy will be both obvious, and shared. If you have an existing, problematic strategy, I will suggest areas where it could be slimmed down, and identify aspects that have been missed. Test strategies can cover a wide range of testing and business issues. While not a checklist, you might expect to see some of the following in your own strategy: - values and decision-making framework - approaches to risk assessment, costs and quality through the organisation - test techniques, test data, test scope and test planning - completion criteria and analysis - test management, metrics and improvement - skills, staffing, team structure and training - test environment, change control and release strategy - defect control, tracking and the approach to fixes - re-tests and regression tests - profiling and analysis for non-functional testing - test automation and test tool assessment **Testing as a genuinely strategic part of software development** A strategic decision is one that opens up opportunities which are otherwise unapproachable - a strategic resource is the mechanism by which these opportunities are exploited. Design, coding, and testing are the three key parts of software development. These are activities, not phases; they operate in parallel and are closely coupled. On a not-too-turbulent project, testing and defect resolution typically account for 25-35% of the final cost of development. Information produced by software testing is clearly an input to the overall organisation's strategic decisions about marketability and release - but testing can provide much more than this for the significant spend it absorbs. To explore ways that testing can realise its strategic value to your organisation and your project, email jdl\[at\]workroom-productions.com, or [call me on Skype](skype:workroomprds?call). ### Exploratory Testing URL: https://www.workroom-productions.com/exploratory_testing_older/ Last updated: 2025-11-04T11:52:37.000Z We build our code into relatively-simple things which work in a complex environment. The systems that result – mixing code, configuration, data, events, and other software / hardware / people systems – can surprise us. We can choose to search for those surprises *before* they bite. I've been thinking and writing about exploratory testing for years; I reckon that it's a core and necessary testing practice which doesn't yet have a firm basis in collaborative software engineering, and is often neglected in strategy and management. This page was written as a temporary holder for some of the things I've written - it's become a resource for some people, so I'm leaving it here as a page for now. I intend to revisit all these dashed-off articles. I'll replace this page with a collection of more-recent articles as I reflect on recent work. For more-polished thought-on-paper, I recommend that you check out [Elisabeth Hendrickson](https://twitter.com/testobsessed)'s fine [Explore It!](https://pragprog.com/titles/ehxta/explore-it/), and [Maaret Pyhäjärvi](https://twitter.com/maaretp)'s [Contemporary Exploratory Testing](https://leanpub.com/exploratorytesting). ## Management - [Adventures in Session-based Testing.](https://workroom-productions.com/papers/AiSBTv1.2.pdf) - ET Notes - [Why Exploration has a place in any Strategy](../why-exploration-has-a-place-in-any-strategy/) - ET Notes - [Scripting and Exploring](https://workroom-productions.com/papers/ET%20Script%20or%20Explore.pdf) ## Techniques and support - [A Positive View of Negative Testing.](https://workroom-productions.com/papers/PVoNT%5Fpaper.pdf) - ET Notes - [The Importance of being Judgemental](https://workroom-productions.com/papers/Judging.pdf) - ET Notes - [What to Record](https://workroom-productions.com/papers/Record.pdf) - ET Notes - [Software testing diagram 1 and variants](https://workroom-productions.com/papers/SWT%20diag%201.pdf) --- I don't intend this page to be a collection of ET resources – others do that better. For instance, here's [something from Ministry of Testing](https://www.ministryoftesting.com/dojo/lessons/a-really-useful-list-for-exploratory-testers). ### In-House Workshop URL: https://www.workroom-productions.com/in-house-workshop/ Last updated: 2021-06-26T16:47:54.000Z ![](../images/microscope-light.svg) # Insights into Exploratory Testing ### Help your team to find useful information more swiftly [Set a date](mailto:jdl@workroom-productions.com?subject=In-house%20Exploratory%20Testing%20workshop) - or read on for more details ## Hands-on workshop ### to enthuse and teach your testers Your team will explore custom-built systems to reveal **key techniques** of exploratory testing, try out viable approaches to **managing exploration**, and build their **collective understanding** of systems testing. We will **design tests**, **find problems**, and **identify differences** in testing style. Working together will let us find **shared values** and **co-operative strengths**. An in-house workshop has a **lower overall price** than a public workshop, with **closer focus** on what matters, and **less disruption** to ongoing work. ![](../images/user-astronaut-light.svg) ## Train your team ### Disciplined, accountable and diverse approaches Use this workshop to help your team: - Choose the right approach and tool for the situation - Explain their test approaches and share their test results - Absorb the disciplines that enable effective exploration - Improve the ways they share testing ideas - Apply existing skills in new ways Workshop for 5-12 people. Suitable for mixed role groups – involve your whole team. During the workshop, we will build a set of achieveable and personalised actions to help the team make concrete and measurable adjustments to their work. Your team will return to their work refreshed and enthused. ![](../images/lightbulb-light.svg) ## Learning support ### Helping your team get the most from the workshop Fit your context by **picking content**. Prepare for the workshop with **short videos** and open questions. Keep focus and resolve ideas with **ongoing contact**. **Practical examples** of real-world testing from a career tester. **Tried and tested** exercises lets you trust the trainer. **Closed workshop** lets you freely share sensitive information. Small workshop allows a coached approach to suit beginners, experienced testers and business experts. ![](../images/users-class-light.svg) ## Workshop contents ![](../images/wide-workshop3.jpg) explore record test experiment manage - Go beyond expectations. Practice working without explicit requirements. Give structure to your activity by choosing a clear purpose. Recognise when your purpose needs to change. - Track data and decisions for sharing and review. Build a collective understanding of exploratory work within your team. - Search for surprises. Develop judgement against internal sources, external specifications, and cultural expectations. Explain your actions and decisions. Give swift, relevant, true feedback. - Investigate reports, building deeper understanding of symptoms, interactions, faults and triggers. Focus bulk tests to trigger and observe patterns in system behaviour. Visualise data to analyse results. - Plan, measure and report exploratory work. Adjust work to suit context. Integrate exploratory testing with organisational goals and activities. Prioritise actions and set aside unproductive paths. Encourage and enable exploration. ## Customers Workroom Productions has built and delivered exploratory testing training at Google, Oracle, Adobe, EADS, GE, Nokia, the BBC and many more organisations. Workshop targets include callcentres, critical national infrastructure, medical devices, hand-helds, retail and finance. We write the exercises that other trainers use. Feedback from the team was overwhelmingly positive and all attendees said they had learned new test approaches that they would apply to their QA & Test work. His enthusiasm and knowledge of the subject was great to see and it meant that the workshop was a thoroughly enjoyable experience for all involved. Stuart Gillies James' training gained very positive feedback from our consultants. He took an excellent approach which was engaging, entertaining and useful. Stewart Noakes Among the best teachers of Exploratory Testing in the world. Scott Barber A first class test strategist & consultant, with a huge passion for testing along with a natural ability to motivate and enthuse those around him Andrew Coggins ![](../images/mind-share-light.svg) ## James Lyndsay ### @workroomprds James is a teacher, speaker and consultant test strategist. Testing since 1996, he has helped dozens of organisations to find the surprises in their systems. He’s taught exploratory testing to corporate clients from 2002, in wide combinations of agile, waterfall and whatever else. James kicked off the TestLab and the London Exploratory Workshop in Testing, built the BlackBox Puzzles and received the 2015 Tester Excellence Award. He studied experiential learing with Jerry Weinberg, improvisation with Keith Johnstone, and Physics at Trinity College, Oxford. [@workrooomprds](http://twitter.com/workroomprds). ![](../images/James3.jpg) | Coached workshop, customised to fit and delivered on-site. Pre-workshop video exercises, post-workshop followup. 5-12 people. | | | ----------------------------------------------------------------------------------------------------------------------------- | ----------- | | | | | 1 Day | £ 1950 . 00 | | Core workshop, skipping some techniques and processes | | | | | | 2 Days | £ 3300 . 00 | | Core workshop with greater depth. | | | | | | 3 Days | £ 4650 . 00 | | 2 days techniques + 1 management. | | | | | | Corporate workshop prices *exclude* VAT and expenses. | | [Set a date / set up a chat ](mailto:jdl@workroom-productions.com?subject=In-house%20Exploratory%20Testing%20workshop) - can't hurt, can it? - [Info](../faq/) - [Blog](http://workroomprds.blogspot.com) - [WPL ![Workroom Productions](../images/biglogo semitrans.png) ](http://workroom-productions.com/) - [ ](http://eepurl.com/beyuRn) - [ ](https://twitter.com/intent/user?user%5Fid=20820279) Text © 2021 Workroom Productions. [Privacy](../privacy/) ### Team Mentoring URL: https://www.workroom-productions.com/team-mentoring/ Last updated: 2024-04-17T14:24:01.000Z *Team Mentoring: Improve together with individual coaching and group consolidation.* Your test team needs keep its skills and its approaches fresh. Team Leads and Managers can help by giving the team time to change, and experiences to absorb. But it can be hard to find the time to improve incrementally while also working on the wider business obligations. Team Mentoring will give your testers the space and the means to focus on improvement. I’ll work individually with everyone in your team, and will consolidate with regular group sessions. I'll help your people to find the right lessons, help those changes to stick, and help your team to learn together. Your team will draw immediately-useful lessons from their day to day work, and grow to fit your organisation’s changing context. [Find a time to chat (Savvycal)](https://savvycal.com/workroomprds/chat) ## How does it Work? *Regular remote contact over months* This is a remote service. I'll be involved regularly and occasionally over months. Any engagement will be for at least 4 months / 17 weeks. I'll read in to your project and technology stack to gain relevant knowledge, and will sign NDAs to suit. I'll fit with your tools of choice, or supply a dedicated (iOS / FaceTime) tablet if needed. Typically I'll talk with your people over Zoom or Teams, and run group sessions in Zoom or Gather. ### Weekly Over a typical week you'll see the following interactions: - Each team member gets weekly one-on-one remote sessions - video-call mentoring, shared-screen testing and more. I'll work with your people on their testing. - One member of your senior staff gets a weekly 30-minute call to review individual and team progress, spotting opportunities and worries, and sharing goals and approaches for the next week. ### Monthly In addition, each month: - Two group sessions for the team. Typically, one would focus inwards on sharing and learning, potentially involving demos, round-tables, games and discussions. Another might turn outwards to dependent / supporting teams, to build a shared culture. ## Prices and Availability I currently have capacity from one team of 4-10 people. Cost is £1050+VAT per person, per month. ### Example Workshop Materials URL: https://www.workroom-productions.com/two_hour_exploratory_interactive/ Last updated: 2025-09-10T15:56:29.000Z _This page is for paying subscribers only._ ### home URL: https://www.workroom-productions.com/home/ Last updated: 2022-09-28T21:39:16.000Z You'll find plenty of activities here to help you get better at software testing. It's all free to use. Subscribe for more depth, and select a paid subscription for more interaction. [Scroll down](#lowFetched) if you want me to help you directly with testing. ### Stuff to Take In URL: https://www.workroom-productions.com/consumables/ Last updated: 2022-01-11T14:00:25.000Z I have masses of materials; just a fraction is currently linked here. The rest is all where it has been for ages. As we go, I'll bring it here. ## Blog / Articles Here are three recent things to read: - Not Testing but Drowning: [Part 1](../not-testing-but-drowning-1/) » [Part 2 – Wrangling](../not-testing-but-drowning-2/) » [Part 3 – Debugging](../not-testing-but-drowning-3/) » [Part 4 – Action](../not-testing-but-drowning-4/). - How Bart and I ran an improvised keynote at ATD: [Keynote eXtreme](https://www.workroom-productions.com/keynote-extreme/) - Investigating [Fuzzy Search](https://www.workroom-productions.com/fuzzy-search/), here on this site. For more, here's this site's [collection of articles](https://www.workroom-productions.com/tag/articles/). I'll bring over the older articles soon, but for now [here's my blog](http://workroomprds.blogspot.com). *I'll update this later with nicer features and a proper list.* ## Videos Here's a few – one to give you a feel for my preferred 2-minute style of teaching short, one online hands-on exploration as part of Ministry of Test's Exploration Week, and an on-stage talk, with stuff to play with, from Agile On The Beach 2018. [Workroom Productions - Private Site Access![](https://www.workroom-productions.com/favicon.ico)Private Site Access](https://www.workroom-productions.com/wicked-problems/) [Experience Report Live: Exploratory Testing a Product with James LyndsayWatch as James Lyndsay takes on Challenge 3 of the exploratory week live, the Exploratory Testing a Product challenge. He ops to test the Python Interpreter. ...![](https://www.ministryoftesting.com/assets/favicon-4b6ba0a4118ce928dfcf3d3230cdbb09347a5a9b8d98f244c79d0d56384758bb.ico)MoT![](http://www.ministryoftesting.com/assets/dojo-open-graph-01-16c70360eeba3459c3ffdfbefa6fdce584fec1a642891b65575446b35a8ec22d.png)](https://www.ministryoftesting.com/dojo/lessons/experience-report-live-exploratory-testing-a-product-with-james-lyndsay) *I'll build this out with a dedicated page and featured videos. I'm making many more videos to support my online workshops.* ## Papers It's been a while since a wrote a paper. Here, though, are a few which caught people's attention. - [The Importance of Data in Functional Testing](https://www.workroom-productions.com/papers/Importance%20of%20Data%20in%20Fn%20Te.pdf) - [Further Adventures in Session-Based Testing](https://www.workroom-productions.com/papers/AiSBTv1.2.pdf) - [Testing in an Agile Environment](https://www.workroom-productions.com/papers/Testing%20in%20an%20agile%20environment.pdf). You can still find my [collection of papers](https://workroom-productions.com/papers.html) on the old site. I'm working to keep the links to these pdfs the same, as they've been cited in others' work. *These are long papers, and .pdfs, so not particularly engaging to casual readers. I'll build this out with a dedicated page and featured papers. I may revist these papers in short chunks, writing about how my ideas have changed since writing them.* ## Outlines and Abstracts Every talk starts with an idea. To get an audience for that idea, you need something to take to a conference organiser. I've done big talks and tiny talks, long-form hands-on tutorials and events. I review papers for EuroSTAR, ATD, MoT (and so can you). I even run workshops on how to do talks and write outlines and abstracts. I'll post the stuff I send to conferences here; some old, some new. You'll see how I work, and I'll get – perhaps – a sense of what works for who. For now, here's a short [page of talk outlines](https://www.workroom-productions.com/tag/outlines/). --- *Below this point, you'll see my experiments with ghost templates and lists.* --- ### Stuff to Play With URL: https://www.workroom-productions.com/toys/ Last updated: 2025-11-04T11:33:11.000Z Play is a great way to learn: I build toys. My toys reward exploration, so are open to any approach. They tend to react to what you do, and their behaviour will tell you something about what's going on underneath. I don't particularly build games which have a winner and a loser, things where you have to do just the right thing within a timebox, toys which ask you to judge correct / incorrect, games of chance or guess-the-right-word games. Some toys are pretty abstract (the *Black Box Puzzles* tend to explore a code paradigm), some are to support exercises (*Converter* and *A Thousand Tiny Tests*), some are to support talks (*Simple Systems*, *Bring Me a Letter* and *Will it Shuffle*). *Questions for Testers* is a group game which involves exploring / predicting each other rather than a system I've built. I need to shift some of the deeper teaching things, such as *Are We Done Yet?* away from Flash, so that may take some time... You can play with them in crowds, or on your own. You're welcome to use them in whatever way you like – I hope you'll tell me how you've used them, and what they have helped you to understand. Each of these will get its own page on this site, as I build them out. ## BlackBox Puzzles 20+ little things, each of which is doing something simple-enough to describe in a tweet. Each takes 5-20 minutes to explore / wonder about / experiment with before you get a succinct description.You may find it takes you a lot less, or a lot longer. I encourage you to precisely describe what it does in terms of behaviour and underlying systems – so go deeper than a list of observations, or a "most of the time it..." . The UI is an important interface, of course, and know that the underlying system is generally built to be decoupled from the UI, and perhaps to respond to other interfaces. Big list, pages and pictures to come. Latest puzzle below. For the rest, go to ![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/2022/02/puzzle36-2-2.jpg) #### You'll make it crash by mashing buttons. Experiment to understand how. [Explore Puzzle 36](../puzzle-036/) ## Raster Reveal [Raster Reveal](../raster-reveal/) is a toy for exploring. If you're wondering about my process, such as it is, here is the [story of how I made Raster Reveal](../making-the-raster-reveal-exercise/). ## Questions for Testers A game to play in a group, which helps you to make connections and have conversations with colleagues and strangers. For now, go to: - to play - for background - [page](../questions-for-testers) on this site, still growing ## Bring Me a Letter An exercise exploring feedback. Bring me a Letter [part I](http://fastfb.workroomprds.com/bmarp1.html), [part II - 1](http://fastfb.workroomprds.com/bmarp2-1.html), [part II - 2](http://fastfb.workroomprds.com/bmarp2-2.html), [part II - 3](http://fastfb.workroomprds.com/bmarp2-3.html), [part III](http://fastfb.workroomprds.com/bmarp3.html) Based on ## Converter A trivial converter, with a host of surprises from a variery of sources. Supports exercises in working with different explorartion surfaces. [Converter v2](https://exercises.workroomprds.com/converter%5Fv2/) Convert meters into ... other things ## A Thousand Tiny Tests Supports exercises in bulk testing: building many linked exeriments to tell you sometime. [A Thousand Tiny Tests](http://exercises.workroomprds.com/thousandtinytests.html) Bulk testing with the generation done for you ## Will it Shuffle? Shows an interesting emergent behaviour, which differs across browsers while the code remains the same. Original by Mike Bostock. For now, ## What shall we Test Next? A idea generator. I'm not sure that testers need that. This one might bump you out of a rut, however. Coming when I get the site back up. ## Are we Done Yet? Using trends to decide whether to keep spending on testing. A simulator used to support workshops on test strategy and coverage, and to illustrate the power of diversity in test approach. Coming when I re-write in HTML/CSS/JS ## Performance Simulator Simulating response times. Design tests to stimulate varous systems to reveal (and to bypass) resource constraints, cacheing and more. Coming when I re-write in HTML/CSS/JS ## Simple Systems To illustrate how simple systems can behave in non-intuitive ways, and to help build intuition into stocks, flows, feedback and resonance. Toy 0 – [a very simple system](https://simplesystems.workroomprds.com/onebathtub.html) Toy 1 – [one tank, one flow](https://simplesystems.workroomprds.com/resonator.html) Toy 2 – [bathtubs part 1](https://simplesystems.workroomprds.com/bathtubs.html) | [bathtubs part 2](https://simplesystems.workroomprds.com/bathtubs2.html) ## Other peoples' stuff I've used these excellent resources; explore them, then stay on Dirk or Nicky's sites, or buy games by Bart and Ryan to see how to *really* do interactives. - Deterministic chaos – [Double Trouble](http://www.complexity-explorables.org/slides/double-trouble/) by Dirk Brockman - Emergent order – [Fireflies](http://ncase.me/fireflies/) by Nicky Case - [Bart Bonte](https://www.bontegames.com)'s delightful "colour" puzzles - Pink, Green, Black, Blue, Red, Yellow and more. Excellent stuff, and groovy music. Swift-ish to explore, with positive feedback on solving; solve it first, then try to describe *how* you solved it. - Grow Pixel / Ryan McLeod's awesome [Black Box app](https://apps.apple.com/us/app/blackbox/id962969578) on iOS. Try to work out what you can do to your device to trigger the 'solution' – and don't touch the screen... ### Stuff to Do Together URL: https://www.workroom-productions.com/exercises/ Last updated: 2024-12-10T10:35:13.000Z It's good to learn from an activity; a provocation of some sort, a process to engage with, a moment to reflect, and a chance to share. I've built dozens of conference events and tutorials – here are frameworks for exercises that I've used in those. They tend to work well with groups; I hope you find them valuable when you try them with your team. Most of these exercises arealready on my GitHub. Here, they're perhaps a bit tidied up, with some thought on how they've worked for me. Join the community to add your comments, and to learn from others. I run several of these as short online workshops, ideal for lunch-and-learn sessions. You'll find links to some exercises below, organised by workshop. *This page needs plenty of work; I need all the exercises linked here individually, and I need to organise thematically rather than offer a handful of bare links under an enigmatic workshop title. Each exercise will be a post on this site, and I'll run them from time to time with members.* ## Exercises My old exercise page contains links to Questions For Testers (not an exercise), Bulk Testing, Robot Sim (needs Flash), several Puzzles, and the paper and exercises (also need Flash) for Four Exercises for Teaching Exploratory Testing. It's out of date and incomplete, and that is one of the main reasons for moving to this site. [Workroom Productions: Exercises](http://exercises.workroomprds.com) ## Hands-on Lab / Work and Play Drop-in Lab at ATD, workshop at EuroSTAR 2018. [Work And Play](https://workandplay.workroomprds.com) Work and Play [Hands-On Lab – EuroSTAR 2018](https://handsonlab.workroomprds.com) Hands-on Lab ## Stay Sharp Workshop delivered at EuroSTAR2016\. Primarily social exercises, all relevant to testers, built to enthuse and re-engage testers with their work. [StaySharp: Keeping Testers Interested![](https://static.ghost.org/v5.0.0/images/link-icon.svg)](https://staysharp.workroomprds.com) [GitHub - workroomprds/StaySharp: Exercises for workshop on keeping testers interestedExercises for workshop on keeping testers interested - GitHub - workroomprds/StaySharp: Exercises for workshop on keeping testers interested![](https://github.com/fluidicon.png)GitHubworkroomprds![](https://opengraph.githubassets.com/838c900a14bee6dc706619102f41edeb82e14d25f8e135fdd43bfcd204c2c32c/workroomprds/StaySharp)](https://github.com/workroomprds/StaySharp) ## Speaker Prep Regular long drop-in workshop delivered at ATD and at other conferences. [Speaker Prep Day](http://speakerprepday.workroomprds.com) Speaker Prep Day – page full of exercises [GitHub - workroomprds/SpeakerPrepDay: Supporting material for Speaker Prep DaySupporting material for Speaker Prep Day. Contribute to workroomprds/SpeakerPrepDay development by creating an account on GitHub.![](https://github.com/fluidicon.png)GitHubworkroomprds![](https://opengraph.githubassets.com/d83979312e718fe4497ed82f329c4c736c784e33e85f46162c344f53a2a1c0af/workroomprds/SpeakerPrepDay)](https://github.com/workroomprds/SpeakerPrepDay) ## The TestLab Bart Knaack and I developed the basic idea of a space at conferences for testers to test within moments of meeting. We ran it, together and separately, at many events – and at EuroSTAR, ATD, Let's Test and the STAREast / STARWest conferences, it became a regular part of the program. We often had help, and those people have gone on to run TestLabs of their own. What an amazing, exploratory and collaborative thing to be part of! We've built so many things to do, over so many years. All ephemeral, most lost, some worth retrieving. Details to be added. [The TestLab · The TestLab](http://testlab.site) testlab.site ## TestLab Variant: BuildAnything Lab A 2-3 day open workshop at ATD2015 – ATD bravely bought in powertools and junk and IoT devices and robots and more. Bart Knaack and I took over the lobby, inviting people to drop in and build anything. Our only stipulation was that it should work with something else. And, looking back, that it should work repeatedly and reliably, that it should be safe, that building it should be safe and achievable and fun. Inevitably, the resulting machine was pointless, extraordinary and ran just once – but the build was collaborative, asynchronous, improvised, Agile, test-focussed, illuminating and memorable. We recently found this video... The Rube Goldberg machine at Agile Testing Days 2015 ## TestLab Variant: CollabLab A 2 -3 day Collaboration Laboratory at ATD2017\. We had a bunch of things for people to do that weren't directly testing software. Exercises built for this cross-fertilised StaySharp, SpeakerPrep and more. Repository (currently) on my local machine. \[JL Note: it's on \`/Users/james/Documents/testing, coding/collablab-exercises\`, and never got to [GitHub](https://github.com/workroomprds/CollabLab)\] [Collaboration Lab](http://collab.workroomprds.com) Action page / form (why have this?) ## Impro For Testers Workshop at Nordic Testing Days, EuroSTAR and for corporate clients. Repository (currently) local only. Recordings stored on `bulk`. --- ## GitHub links [GitHub - workroomprds/CollabLab: For the CollabLab, ATD 2017For the CollabLab, ATD 2017\. Contribute to workroomprds/CollabLab development by creating an account on GitHub.![](https://github.com/fluidicon.png)GitHubworkroomprds![](https://opengraph.githubassets.com/ffca9a866399e4523a37a958630c39eeda40729611694cdf74591ab28d121269/workroomprds/CollabLab)](https://github.com/workroomprds/CollabLab) [GitHub - workroomprds/StaySharp: Exercises for workshop on keeping testers interestedExercises for workshop on keeping testers interested - GitHub - workroomprds/StaySharp: Exercises for workshop on keeping testers interested![](https://github.com/fluidicon.png)GitHubworkroomprds![](https://opengraph.githubassets.com/838c900a14bee6dc706619102f41edeb82e14d25f8e135fdd43bfcd204c2c32c/workroomprds/StaySharp)](https://github.com/workroomprds/StaySharp) [GitHub - workroomprds/SpeakerPrepDay: Supporting material for Speaker Prep DaySupporting material for Speaker Prep Day. Contribute to workroomprds/SpeakerPrepDay development by creating an account on GitHub.![](https://github.com/fluidicon.png)GitHubworkroomprds![](https://opengraph.githubassets.com/d83979312e718fe4497ed82f329c4c736c784e33e85f46162c344f53a2a1c0af/workroomprds/SpeakerPrepDay)](https://github.com/workroomprds/SpeakerPrepDay) ### My Papers URL: https://www.workroom-productions.com/papers-page/ Last updated: 2021-12-01T12:42:47.000Z Most of these papers were written as an integral part of a conference presentation, to help the audience get some value after the talk was over, and to give greater depth to my own thinking. They're generally around 8000 words. --- **The Irrational Tester** We all act irrationally. When we as testers make decisions or give feedback, we influence the overall direction of a project. The ripples of any irrationality can go far. Understanding the patterns that underlie our irrationality will help us be better testers. This paper puts a test perspectiveon bias; why we so often labour under the illusion of control, how we lock onto the behaviours we're looking for, and why two people can use the same evidence to support opposing positions. It covers why timeboxes work, why independence matters, and the subtle de-biasing nudges that can help encourage us to stay on track. There are plenty of real-life examples of tester irrationality, and lots of references to the primary research papers. Download [The Irrational Tester](http://www.workroom-productions.com/papers/The%20Irrational%20Tester%20v1-06.pdf). Keynote at STARWest2009 ([video](http://www.stickyminds.com/Media/Video/Detail.aspx?WebPage=166)), the Danish SIGIST, CTG Software Testing Seminar, TCL's 'Star Testing' day, track at various other events. --- **Testing in an Agile Environment** It is hard to find a practical approach that allows a professional tester to achieve their full potential in an agile environment. This paper draws on experience of real-life agile projects, and will help testers recognise where they are bringing friction to an agile environment, help agile team members recognise where they may be incurring a 'testing debt' and identifies ways that testers can facilitate learning and bring value to an agile project. Download [Testing in an Agile Environment](https://workroom-productions.com/papers/Testing%20in%20an%20agile%20environment.pdf). Keynote at EuroSTAR 2008, IIR Finland, Unicom's Agile Testing Day London, Avenir's Agile Testing Day Oslo, London SIGiST, Ordina's Testing Masterclass and as a track presentation at STAREast, Agile 2008 (basis of "Agile is Groovy, Testing is Square") and others. Find out [more](https://workroom-productions.com/talk%5Ftesting%5Fagile%5Fenviron.html), or [bring the presentation to your team](https://workroom-productions.com/talks.html). --- **Exploratory Testing Notes** A suite of short papers, detailing my thoughts on questions that turn up frequently in my Exploratory Testing classes, and when working with ET teams. - ET Notes - [Why Exploration has a place in any Strategy](https://workroom-productions.com/papers/Exploration%20and%20Strategy.pdf) - ET Notes - [Scripting and Exploring](https://workroom-productions.com/papers/ET%20Script%20or%20Explore.pdf) - ET Notes - [The Importance of being Judgemental](https://workroom-productions.com/papers/Judging.pdf) - ET Notes - [What to Record](https://workroom-productions.com/papers/Record.pdf) - ET Notes - [Software testing diagram 1 and variants](https://workroom-productions.com/papers/SWT%20diag%201.pdf) --- **Things Testers Miss** Designers and Coders make bugs. Testers make tests - but when testers get it wrong, bugs end up in production. This paper is all about ways that testers can miss bugs, and ways that they can catch more. Download [Things Testers Miss](https://workroom-productions.com/papers/Things%20Testers%20Miss.pdf). Paper presented at StarEast05, EuroStar05 - find out [more](https://workroom-productions.com/talk%5Fthings%5Ftesters%5Fmiss.html), or [bring the presentation to your team](https://workroom-productions.com/talks.html). --- **'A Positive View of Negative Testing'** A practitioner's overview of negative testing, dealing with tests designed to make the system fail, and tests that are designed to exercise functionality that deals with failure. Details the overall aims and management of negative testing, and describes a variety of techniques used to select, derive and execute negative tests. Download [A Positive View of Negative Testing](https://workroom-productions.com/papers/PVoNT%5Fpaper.pdf). Paper presented as a keynote at STAREast 2003 - find out [more](https://workroom-productions.com/talk%5Fnegative%5Ftesting.html), or [bring the presentation to your team](https://workroom-productions.com/talks.html). --- **'Adventures in Session-Based Testing'** **with Niel van Eeden** Session-based testing is a management technique often used to measurem and control Exploratory Testing. It can can also form a foundation for significant improvements in productivity and error detection, partcularly in immature teams. This paper describes how two UK companies controlled and improved ad-hoc testing, and used the knowledge gained as a basis for ongoing, product sustained improvement. Download [Adventures in Session-Based Testing](https://workroom-productions.com/papers/AiSBTv1.2.pdf). Presented at STAREast 2002, Quality Week 2002, EuroSTAR 2002 and ASIAStar2003\. A variant presentation based on the same paper ('*Further Adventures in Session-Based Testing*', concentrating on tools and coaching) was presented at STARWest 2002\. Find out [more](https://workroom-productions.com/talk%5Fsession-based%5Ftesting.html), or [bring the presentation to your team](https://workroom-productions.com/talks.html). Winner of "Best Paper" at STARWest 2002. Winner of "Best Paper" at EuroSTAR 2002. --- **'From a Sow's Ear to a Silk Purse: Making the most of what you've got'.** The test team has a greater influence on the success of testing than any single process, tool or technique. Yet, under-resourced and over-stretched, it can be a source of weakness. This paper outlines some effective techniques to help get the best out of your team. Download the paper: [From a Sow's Ear...](https://workroom-productions.com/papers/MTM%5Fpaper.pdf). Presented as the QBIT keynote, QBIT Tools Fair September 2002, T-mobile Special Interest Group on 5 June 2003, AsiaSTAR 2003, Sydney, and (as a double-length special session) at EuroSTAR 2003 and the London SIGiST. Find out [more](https://workroom-productions.com/talk%5Fmaking%5Fthe%5Fmost.html), or [bring the presentation to your team](https://workroom-productions.com/talks.html). --- **'The Importance of Data in Functional Testing'** Details ways to improve your test data to allow it to help, rather than hinder, your testing. Includes a pair-wise approach to generate flexible datasets, and approaches to soft partitioning and data labeling. Also contains an analysis of commonly-found data-related issues and suggests solutions. Download [The Importance of Data in Functional Testing](https://workroom-productions.com/papers/Importance%20of%20Data%20in%20Fn%20Te.pdf). Presented at QW2001, STARWest 2001, QWE2002 and the BCS SIGiST. Find out [more](https://workroom-productions.com/talk%5Fdata%5Fin%5Ftesting.html), or [bring the presentation to your team](https://workroom-productions.com/talks.html). --- **'Test Managers, Chameleons of the Project World'** **with Julie Gardiner** This paper presents two models of test management. The first describes the test manager's interactions with other roles on the project, based on the interests represented by the manager in the interaction. The second describes the activities, deliverables, decisions and expectations of a test manager over the life of a project, paying particular attention to test phase. The paper will be useful to new and experienced test managers. Download [Test Managers, Chameleons of the Project World](https://workroom-productions.com/papers/TM-CPWv1.pdf). Presented at EuroSTAR 2003 and AsiaSTAR 2004. ### Talks URL: https://www.workroom-productions.com/talks/ Last updated: 2021-11-25T22:26:27.000Z You can bring some of the following talks to your own team. Each has been delivered at a major testing event, and most are supported with substantial papers. Use these talks to motivate your team, to give them a special event, and to stimulate debate. [Testing in an Agile Environment](https://workroom-productions.com/talk%5Ftesting%5Fagile%5Fenviron.html) Delivered at Ordina's Agile Testing Masterclass, the SIGiST and TestNet. Basis of EuroSTAR08 keynote. Supported by 7500 word paper [Things Testers Miss](https://workroom-productions.com/talk%5Fthings%5Ftesters%5Fmiss.html) Delivered at StarEast and EuroStar. Supported by 7500 word paper [A Positive view of Negative Testing](https://workroom-productions.com/talk%5Fnegative%5Ftesting.html) StarEast keynote, also delivered at EuroStar. Supported by 8500 word paper [Adventures in Session-Based Testing](https://workroom-productions.com/talk%5Fsession-based%5Ftesting.html) Delivered at StarWest, EuroStar (session and tutorial), AsiaStar (as keynote) supported by 8500 word paper [From a Sow's Ear to a Silk Purse: Making the Most of what you've got](https://workroom-productions.com/talk%5Fmaking%5Fthe%5Fmost.html) Delivered at EuroSTAR, SIGiST, supported by 5500 word paper [The Importance of Data in Functional Testing](https://workroom-productions.com/talk%5Fdata%5Fin%5Ftesting.html) Delivered at Quality Week, StarWest, EuroStar (session and tutorial), supported by 8500 word paper I'm an experienced speaker, regularly giving talks at major US and European conferences. I try to find an angle on a subject that has not been addressed before, and to deliver something new, with practical value, in each talk. Audiences find my talks both entertaining and informative. See [endorsements](https://workroom-productions.com/endorsements.html) for their reactions. --- **Other Talks** Some of the following talks are available as audio or video downloads. If you'd like to bring any of these to your team, please contact me. --- **Achievable Futures \- a Google TechTalk!** Over the last decade, we've seen huge changes in the commoditisation and ubiquity of computers, and have vastly more power at our fingertips. I believe that testing has failed to keep up with the times. This talk highlights important trends and asks if we can rise to their challenge. I think we can, and here I describe how we might use our existing skills, approaches and tools to achieve a bright, and rather different future. Watch the video of [Achievable Futures](https://www.youtube.com/watch?v=5BHS2T9tdiM) This talk was originally commissioned for the 10th anniversary of the [Swedish Association for Software Testing](http://www.sast.se/). The following recording was made at the event. I've split it into handy chunks. | Introduction | 4:00 | 1.6 Mb | [listen](https://workroom-productions.com/papers/AF%201.%20Introduction.mp3) | | ------------ | ----- | ------ | ---------------------------------------------------------------------------- | | Trends | 8:52 | 3.5 Mb | [listen](https://workroom-productions.com/papers/AF%202.%20Trends.mp3) | | Potential | 13:34 | 5.4 Mb | [listen](https://workroom-productions.com/papers/AF%203.%20Potential.mp3) | | Tools | 14:23 | 5.8 Mb | [listen](https://workroom-productions.com/papers/AF%204.%20Tools.mp3) | | Conclusion | 6:05 | 2.4 Mb | [listen](https://workroom-productions.com/papers/AF%205.%20Conclusion.mp3) | Download the slides as a .pdf: [Achievable Futures](https://workroom-productions.com/papers/Achievable%20Futures%20SAST.pdf) Also presented at Tampere University, [ScotTest](http://www.scottest.org.uk/) and the [London SIGiST](http://www.sigist.org.uk/). --- **'The Test Strategist's Toolbox'** Presented at AsiaSTAR 2004 in Canberra and STAREast 2005 in Orlando. Here's a [recording](https://workroom-productions.com/papers/TSTbox.mp3) \- 21Mb, 50 minutes long. --- **Automated Tricks for Manual Testers - a lightning Google TechTalk!** A very short talk from Google's 2006 London Test Automation Conference. Please watch the other speakers - but if you want to get to my part, it's around 33:45. Watch the 5-minute lightning talk [Automated Tricks for Manual Testers](http://www.youtube.com/watch?v=mR2VmDVgX6E). --- **'Tips for Testers: Bugs are all around us'** Ten top tips for testers: 10-minute presentation at the London SIGiST December 2003. [Listen to a recording of the presentation](https://workroom-productions.com/papers/Tips%5Ffor%5FTesters-bugs.mp3) (10 minutes, mp3, 2.5Mb). --- **Exploration and Introspection** Presented at SoftTest, Dublin, 29 September 2004\. Presented with "*Further Adventures in Session-based Testing.*" --- **'The Real Deal: How testing can tell you what you really want to know':** Slides are [here](https://workroom-productions.com/papers/Real%20deal%5Fs.pdf) (.pdf) | Planning | 3:44 | 0.5 Mb | [listen](https://workroom-productions.com/papers/TRD%201.%20Planning.mp3) | | -------------------------- | ---- | ------ | --------------------------------------------------------------------------------------- | | Information | 5:13 | 0.7 Mb | [listen](https://workroom-productions.com/papers/TRD%202.%20Information.mp3) | | What do you want to know? | 6:24 | 0.8 Mb | [listen](https://workroom-productions.com/papers/TRD%203.%20What%20do%20you%20want.mp3) | | How to get the information | 4:00 | 0.5 Mb | [listen](https://workroom-productions.com/papers/TRD%204.%20How%20to%20find%20out.mp3) | | Supporting the Process | 1:54 | 0.2 Mb | [listen](https://workroom-productions.com/papers/TRD%205.%20Support.mp3) | | Conclusion | 2:24 | 0.3 Mb | [listen](https://workroom-productions.com/papers/TRD%206.%20Conclusion.mp3) | Presented as the QBIT keynote, QBIT Tools Fair February 2003, London. ### Welcome, Subscriber URL: https://www.workroom-productions.com/welcome-free-subscriber/ Last updated: 2024-09-22T17:34:59.000Z Thank you for subscribing to the site! You'll now be able to see deeper content on many posts. _This page is for subscribers only._ ### Welcome! URL: https://www.workroom-productions.com/welcome-paid-subscriber/ Last updated: 2025-11-04T11:52:57.000Z Thank you for subscribing to the site and to the interactive sessions! You'll now be able to see deeper content, and you'll receive weekly invites to my workshops. _This page is for paying subscribers only._ ### TLP1 URL: https://www.workroom-productions.com/tlp1/ Last updated: 2021-12-03T16:34:59.000Z TPL1 content Here's a [Link](https://www.workroom-productions.com/home/) ## Posts ### Puzzle 41 URL: https://www.workroom-productions.com/puzzle-41/ Last updated: 2026-08-23T12:21:27.000Z The buttons and lamp obey a fairly simple principle. **What is it?** ### Puzzles supported by these lovely people [***Julie Gardiner***](https://talent-unleashed.com), [***Joep Schuurkes***](https://smallsheds.garden), [**Huib Schoots**](https://huibschoots.nl/about-me/), [*Ide Koops*](https://www.linkedin.com/in/koopside/?lipi=urn%3Ali%3Apage%3Ad%5Fflagship3%5Fsearch%5Fsrp%5Fpeople%3BXWQgAmL6QD%2BrTEFLDXdF4g%3D%3D), [*Kris Corbus*](https://www.linkedin.com/in/kriscorbus/), [Alan Richardson](https://www.eviltester.com), Rob van Steenbergen, Ioana Chiorean, Peter Houghton, Adun Urke, Christine Yen, Pascal Dufour all support me on [Patreon](https://www.patreon.com/workroomprds). Not for this puzzle though. Help me make more and I'll put your name on this list. There are other perks, too, eventually. Enjoy this? [Support another!](https://www.patreon.com/workroomprds) Built by James Lyndsay - [@workroomprds](http://twitter.com/workroomprds) © Workroom Productions 2026 ### Workroom PlayTime 063: Spot the Difference URL: https://www.workroom-productions.com/workroom-playtime-063-spot-the-difference/ Last updated: 2026-08-12T19:39:59.000Z This week's [Workroom PlayTime](https://www.workroom-productions.com/workroom-playtime/) ***is*** going to happen, but it is not going to be on Thursday (tomorrow / today depending on where you are). We'll try for [Friday 14 August at 7:30pm UK time](https://this-ti.me/?uts=1786732200&tz=Europe%2FLondon&name=Workroom+PlayTime+063). If that doesn't suit you, and you want to find another, ping me and if we can find a mutual one, I'll send another invite to subscribers. We'll run [Spot the Difference (Puzzle 40)](https://www.workroom-productions.com/spot-the-difference-puzzle-40/). We'll gather on Zoom. These exercises are for everyone, for free. [All subscribers](https://www.workroom-productions.com/#/portal/signup) get to play together most weeks, paid subscribers get a guaranteed slot or rerun if we're full. If you get this in an email or can see the **joining info** section below, then you're a subscriber and have access to the Zoom etc. for the workshop. If you'd like to bring a friend, you can do that. I'll say yes until I've got too many. Use this Savvycal link to [book your spot](https://savvycal.com/workroomprds/workroom-playtime-063) and get reminders. *Note to regulars: You should find it easier to use the Zoom link – it's in several places. Tell me if you like it or if it's a problem.* ↓ *signed-in subscribers will see joining info and links below* ↓ _This post is for subscribers only._ ### Spot the Difference (Puzzle 40) URL: https://www.workroom-productions.com/spot-the-difference-puzzle-40/ Last updated: 2026-08-12T19:30:45.000Z Open [Puzzle 40](https://www.workroom-productions.com/puzzle-040/) and [Puzzle 40b](https://www.workroom-productions.com/puzzle-040b/) in tabs. There is *one* substantial difference, and it's not obvious to the casual explorer. The exercise is to find it. You will find it by exploring, modelling and comparing. You will find it by comparing sources. You will find it by asking me for hints. You will find it by asking me. You might aim your exploration by thinking what I could easily change in one place that leads to a non-obvious change in behaviour. You can probably think of more. First person (to be sure of what it might be) gains kudos – and becomes a helper and hinter to the others, *if they ask*. Or maybe we'll all explore together, and so we'll find it together. At the start of the exercise, we'll agree collective ground rules – how long we'll spend, how collaborative we'll be. We might also exchange individual handicaps – maybe you won't look at the source, or you'll only use specific tools. ### Puzzle 40b URL: https://www.workroom-productions.com/puzzle-040b/ Last updated: 2026-08-08T21:36:57.000Z Describe how the buttons affect the lamps. This version has a single meaningful change from [Puzzle 40](https://www.workroom-productions.com/puzzle-040/) – you could use [diff](https://www.workroom-productions.com/diff-for-testers/) to find it, you could look at the behaviours, you could think of somewhere I might easily have dropped in a meaningful change that doesn't quite get picked up by typical validation tests. You could ask, of course. I don't typically put bugs in. This change is intentional, so arguably it isn't a bug. But the behaviour it introduces does, in some sense, break the symmetry of what you might have found in 40\. Off we go then. ### Currently supported by these lovely people [***Julie Gardiner***](https://talent-unleashed.com), [***Joep Schuurkes***](https://smallsheds.garden), [**Huib Schoots**](https://huibschoots.nl/about-me/), [*Ide Koops*](https://www.linkedin.com/in/koopside/?lipi=urn%3Ali%3Apage%3Ad%5Fflagship3%5Fsearch%5Fsrp%5Fpeople%3BXWQgAmL6QD%2BrTEFLDXdF4g%3D%3D), [*Kris Corbus*](https://www.linkedin.com/in/kriscorbus/), [Alan Richardson](https://www.eviltester.com), Rob van Steenbergen, Ioana Chiorean, Peter Houghton, Adun Urke, Christine Yen, Pascal Dufour all support me on [Patreon](https://www.patreon.com/workroomprds). Help me make more and I'll put your name on this list. There are other perks, too, eventually. Enjoy this? [Support another!](https://www.patreon.com/workroomprds) Built by James Lyndsay - [@workroomprds](http://twitter.com/workroomprds) © Workroom Productions 2026 ### Puzzle 40 URL: https://www.workroom-productions.com/puzzle-040/ Last updated: 2026-08-08T14:33:46.000Z Describe how the buttons affect the lamps. ### Currently supported by these lovely people [***Julie Gardiner***](https://talent-unleashed.com), [***Joep Schuurkes***](https://smallsheds.garden), [**Huib Schoots**](https://huibschoots.nl/about-me/), [*Ide Koops*](https://www.linkedin.com/in/koopside/?lipi=urn%3Ali%3Apage%3Ad%5Fflagship3%5Fsearch%5Fsrp%5Fpeople%3BXWQgAmL6QD%2BrTEFLDXdF4g%3D%3D), [*Kris Corbus*](https://www.linkedin.com/in/kriscorbus/), [Alan Richardson](https://www.eviltester.com), Rob van Steenbergen, Ioana Chiorean, Peter Houghton, Adun Urke, Christine Yen, Pascal Dufour all support me on [Patreon](https://www.patreon.com/workroomprds). Help me make more and I'll put your name on this list. There are other perks, too, eventually. Enjoy this? [Support another!](https://www.patreon.com/workroomprds) Built by James Lyndsay - [@workroomprds](http://twitter.com/workroomprds) © Workroom Productions 2026 ### Workroom PlayTime 062: Explore Puzzle 039 URL: https://www.workroom-productions.com/workroom-playtime-062-explore-puzzle-039/ Last updated: 2026-07-30T14:17:18.000Z This week's [Workroom PlayTime](https://www.workroom-productions.com/workroom-playtime/) is going to happen twice: [today (Thursday) at 21:30 London time](https://this-ti.me/?uts=1785441600&tz=Europe%2FLondon&name=Workroom+PlayTime+062), and again at [09:00 London time on Saturday](https://this-ti.me/?uts=1785571200&tz=Europe%2FLondon&name=Workroom+PlayTime+062). Do let me know if a particular time works better for you. Thank you for bearing with me as I took my time We'll play with [Puzzle 039](https://www.workroom-productions.com/puzzle-039/) together. This is the first in a series of puzzles over the next few weeks. We'll gather on Zoom. The BlackBox Puzzles puzzles are for everyone, for free (but you can become a patron if you like). Playing together is for [all subscribers](https://www.workroom-productions.com/#/portal/signup) and their guests. Paying subscribers get a guaranteed slot or rerun if we're full. If you can see the **joining info** section below, then you're a subscriber and have access to the Zoom etc. for the workshop. Use this Savvycal link to [book your spot](https://savvycal.com/workroomprds/workroom-playtime-062) for today or for Saturday. Savvycal will send you reminders. *Note to regulars: You should find it easier to use the Zoom link – it's in several places. Tell me if you like it or if it's a problem.* ↓ *signed-in subscribers will see joining info and links below* ↓ _This post is for subscribers only._ ### Puzzle 39 URL: https://www.workroom-productions.com/puzzle-039/ Last updated: 2026-08-01T11:11:46.000Z The buttons and lamps obey a fairly simple principle. **What is it?** ### Currently supported by these lovely people [***Julie Gardiner***](https://talent-unleashed.com), [***Joep Schuurkes***](https://smallsheds.garden), [**Huib Schoots**](https://huibschoots.nl/about-me/), [*Ide Koops*](https://www.linkedin.com/in/koopside/?lipi=urn%3Ali%3Apage%3Ad%5Fflagship3%5Fsearch%5Fsrp%5Fpeople%3BXWQgAmL6QD%2BrTEFLDXdF4g%3D%3D), [*Kris Corbus*](https://www.linkedin.com/in/kriscorbus/), [Alan Richardson](https://www.eviltester.com), Rob van Steenbergen, Ioana Chiorean, Peter Houghton, Adun Urke, Christine Yen, Pascal Dufour all support me on [Patreon](https://www.patreon.com/workroomprds). Help me make more and I'll put your name on this list. There are other perks, too, eventually. Enjoy this? [Support another!](https://www.patreon.com/workroomprds) Built by James Lyndsay - [@workroomprds](http://twitter.com/workroomprds) © Workroom Productions 2026 ### Workroom PlayTime 061: DeepWiki as a Tester URL: https://www.workroom-productions.com/workroom-playtime-061-deepwiki-as-a-tester/ Last updated: 2026-07-08T22:59:07.000Z This week's [Workroom PlayTime](https://www.workroom-productions.com/workroom-playtime/) is not going to be on Thursday – there's an expected conjunction of several commitments, and I've not found a time. So I'm shifting, for this week, to [Friday 10 July at 7:30pm UK time](https://this-ti.me/?uts=1783708200&tz=Europe%2FLondon&name=Workroom+PlayTime+061). Please **note the change of time**. And do let me know if this time works better for you. We'll play with tool DeepWiki, which reads code in public github repos, and writes documents. I recently [wrote about my experiences with DeepWiki](https://www.workroom-productions.com/using-deepwiki-as-a-tester/), as a tester – it's neither deep, nor a wiki, but it is interesting. If generated insights are in your wheelhouse, you might find my [adventures with Graphify](https://www.workroom-productions.com/local-llm-tooling-code-queries-with-graphify/) interesting, too, and that's all on a local LLM. Here are the [exercises](https://www.workroom-productions.com/exercise-deepwiki-as-a-tester/) we'll use. I hope you've got a github repo that you already know, but you'll see a couple of suggestions from me if not. We'll gather on Zoom. These exercises are for everyone, for free. [All subscribers](https://www.workroom-productions.com/#/portal/signup) get to play together most weeks, paid subscribers get a guaranteed slot or rerun if we're full. If you get this in an email or can see the **joining info** section below, then you're a subscriber and have access to the Zoom etc. for the workshop. If you'd like to bring a friend, you can do that. I'll say yes until I've got too many. Use this Savvycal link to [book your spot](https://savvycal.com/workroomprds/workroom-playtime-061) and get reminders. *Note to regulars: You should find it easier to use the Zoom link – it's in several places. Tell me if you like it or if it's a problem.* ↓ *signed-in subscribers will see joining info and links below* ↓ _This post is for subscribers only._ ### Exercise: DeepWiki – as a tester URL: https://www.workroom-productions.com/exercise-deepwiki-as-a-tester/ Last updated: 2026-07-10T18:28:53.000Z Bring a link to a public github repo – preferably, a repo you know. If you can't, then use - curl (glorious, huge, explored by me [here](https://www.workroom-productions.com/using-deepwiki-as-a-tester/)) or - a tiny thing of mine (my one is the core logic for a simple blackbox puzzle. You can run the tests at ) Paste it into DeepWiki (or just change the `github` to `deepwiki` in the URL): [DeepWiki | AI documentation you can talk to, for every repoDeepWiki provides up-to-date documentation you can talk to, for every repo in the world. Think Deep Research for GitHub - powered by Devin.![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/icon/icon-4a709292-f117-4ab2-b641-d9eb86172014.png)DeepWiki![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/thumbnail/opengraph-image-701a14d2-66cd-4e3c-9552-3853c9d47ca0.png)](https://deepwiki.com) ## Exercise ### Primer *5 mins* Primer: If you're coming into a new repo, where might you start? System diagrams? Tests? Entry points? Tech stack? User needs? Value? Risks? Open bugs? PRs? Decision records? Add your own... Deepwiki is based on the code and the config: it won't be appropriate for things not described in code. **From the list, what might those be?** ### **Explore** *10 mins* Dig in! #### My guidance (should you need it) - Find and read some sort of **Quick Start Guide*, pick holes in it, spot things you didn't already know, check them. My tiny repo is too pointless to have one... - Take a look at the table of contents (the list on the left) – what parts interest you as a tester? - Go look at a part you know well. Find accuracies, misconceptions, absences and over-complications. - Enumerate: Find several different ways that code is (or data are) represented. - Ask the thing a question – maybe about the tests, or the architecture. - Gauge whether the information is an up-to-date representation of the repo (hint: it's not). ### Share *5 mins* Share your thoughts ### Recent eye-catching testing articles URL: https://www.workroom-productions.com/recent-eye-catching-testing-articles/ Last updated: 2026-07-07T13:11:09.000Z _No content available._ ### Using DeepWiki as a Tester URL: https://www.workroom-productions.com/using-deepwiki-as-a-tester/ Last updated: 2026-07-10T17:39:46.000Z [DeepWiki](https://deepwiki.com/) reads code, and makes documentation to match. It's a Codebase Intelligence Tool. #### ****Codebase Intelligence tools**? You may be familiar with similar tools. I'm not about to survey the field, but I've already played with [Sourcegraph](https://www.workroom-productions.com/exchanges-with-sourcegraphs-cody-about-curl/) and [Graphify](https://www.workroom-productions.com/local-llm-tooling-code-queries-with-graphify/) here. Both those, and lots more, lean on [Tree Sitter](https://tree-sitter.github.io/tree-sitter/), which is [deterministic at its core](https://deepwiki.com/search/is-this-deterministic-software%5Fac1fedac-7eaa-43b6-885d-8b128c396823?mode=fast). And that crisp analysis is a DeepWiki link, of course... DeepWiki uses LLMs on top of algorithmic analysis, so while it does contextualise terse output into something a human might take in, it's not *definitive*. Unlike some of its rivals, it only really reads code/docs/config/imports/repo structure and not issues/commits/PRs/ADRs/meeting notes/minds, so it has no context about history nor intention. Which is probably to the good – inference about human action is best done by some*one*, not some*thing*. It is, of course, a sales funnel as well as a handy tool. It's easily and freely accessible, but only for open-source projects on GitHub. It's proprietary, so can be made not-free any time (or degraded, taken away, made dangerous or even improved). It won't use local models, so [DeepWiki](https://deepwiki.com) / [Devin](https://devin.ai) / [Cognition Labs](https://cognition.com) are subsidising and encouraging your use of a farty / thirsty / nosy / just awful token monster. It looks as though it could be linked to the projects it has processed. It is not. It looks authoritative. It is not. It is tempting to think of it as a great way to document an undocumented project. It's not, or at least, [*not really*](https://codersera.com/blog/deepwiki-vs-traditional-documentation-developer-decision-framework/)*.* Its un-checked output is eminently indexable and so contributes to the vast wave of slop that is turning search to shite. It feels [*not good* to the makers of the code](https://blopker.com/writing/12-deepwiki/). And, dammit, it's *not at all a* [*wiki*](https://en.wikipedia.org/wiki/Wiki). However, as code explainers go, it is as astonishing as you might imagine from a deterministic tool gussied up by a well-prompted large language model. Especially when that deterministic tool can stand on the shoulders of well-established and comprehensive open-source precursors. As an LLM-based tool, it will give you both ready-made analysis, and let you ask questions, and offer directions directly to the code to support its less-reliable answers. It's great for a tester wanting to better-understand an unfamiliar codebase, from broad strokes to fine details. Just so long as you stay a tester, sceptical of unsupported answers. And if you're a tester exploring a deliverable, rather than verifying the expectations of the humans making or using the system, then it offers an interesting exploratory surface representing that deliverable. Here's a short tour, roughly following me as I pottered around [curl](https://curl.se), vaguely wondering about some details of that project's testing. - Take a github URL. Change `github` to `deepwiki` and you'll get the documentation. Bosh. - ? ! - Do check the `Last indexed: ` date, top-left... - The docs are typically immediate – DeepWiki are clearly cacheing, so there's a compromise between up-to-date and not-needing-to-do-it-again - You'll see a sensible and clickable table of contents, which seem to relate to the project, not just to the files. I popped into [Test Suite Architecture](https://deepwiki.com/curl/curl/6.1-test-suite-architecture); you might go somewhere else. - Documentation includes readable diagrams, clickable links to highlighted core code, and you can ask questions. - Try the different ways of exploring the codebase - See how many ways to explore you can find - Consider what you might do to check that what you're absorbing is *right.* - When you've identified something you want to share, you can share the url with your team or with future you, and the analysis comes along with it. Here are some I might want to share: - [What patterns of tests actually get executed by valgrind](https://deepwiki.com/search/tell-me-about-tests%5Fcac76093-5542-4844-bd24-32261921419e?mode=fast) - [Tell me more about the helper methods in ExecResult. Where else appears to include helper methods?](https://deepwiki.com/search/tell-me-more-about-the-helper%5F5145de31-d88e-4b66-b632-2e5880455e31?mode=fast) - Again, these seem to be cached. For me, the first query takes \~1 minute, re-entering is immediate. So the results may be out of date, or they might be up-to-date but different from what you expected to share. You don't get to pick. You do need to know. Is what DeepWiki says true? Nope. Not that *I* can identify a specific untruth, because I don't know the truth well enough (and I've not spent the time to look for inconsistencies). But as [any fule kno](https://en.wikipedia.org/wiki/Nigel%5FMolesworth#St%5FCustard's), not being able to spot a lie doesn't make something true. [Here](https://news.ycombinator.com/item?id=45002092) and [here](https://blopker.com/writing/12-deepwiki/), DeepWiki said that something was true, when it was (somehow) in the code but not in the product. Generated text is ***all* imagined, every last word**. And I'm using 'imagined' to mean something we don't properly have a word for, using it in the sense of some inhuman and uncanny truth-adjacent noise. To my eye, DeepWiki's `imagination` seems close enough to give you good clues, seems guarded from offering plausible illusions about intent and reasons, and gives enough direct access to code that you can at least think about it yourself. You can be just as sceptical of its output as you are of your own conclusions: The map may be weird, but it can still help you to form your own educated opinion. It's still a world away from dibbling about, as I did [publicly in 2023](https://www.workroom-productions.com/exchanges-with-sourcegraphs-cody-about-curl/), with the tools of that time. And a world away again from thrashing around with `grep` , a whiteboard and some departed dev's scribbled handover notes. Go on. Go play, ask better questions than I did, and share them here. ### Support your Testing with Local LLMs URL: https://www.workroom-productions.com/support-your-testing-with-local-llms/ Last updated: 2026-06-27T21:09:32.000Z A double-length hands-on workshop, going out at Agile Testing Days 2026, with Bart Knaack and me. [Agile Testing Days | Europe’s Greatest Agile Software Testing FestivalJoin Europe’s leading software testing festival! Experience world-class keynotes, workshops & the unique Unicorn Spirit at Agile Testing Days in Potsdam.![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/icon/agiletd24-icon_color-7200a49d-4175-4cbc-9601-c09d6bffa955.svg)Europe’s Greatest Agile Software Testing Festivaltrendig![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/thumbnail/agiletd24-icon_color-c7191de2-6e7d-433d-a3f1-ed406fb86877.svg)](https://agiletestingdays.com/2026/session/support-your-testing-with-local-llms/) Expect this page to be updated from time to time. Here's what we're bringing to support the workshop. [Local LLM Tooling: Code Queries with \`graphify\`Supporting my local LLM with a code-graphing tool![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/icon/wpl-60x60-logo-semitrans-84a0715a-d58d-4fe9-b6d3-e2b83083cd3e.png)Workroom ProductionsJames Lyndsay![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/thumbnail/image-1-597bde03-de0e-4002-bcd4-b7cd435265e7.png)](https://www.workroom-productions.com/local-llm-tooling-code-queries-with-graphify/) ### Local LLM Tooling: Code Queries with `graphify` URL: https://www.workroom-productions.com/local-llm-tooling-code-queries-with-graphify/ Last updated: 2026-06-27T13:33:04.000Z I’m trying out a local LLM (`Qwen3.5:9b`) and harness (`Pi`) on a real project (`FlockXR`), running the whole thing on a 32GB M2Max with the agent running in a container (`Docker`) and the LLM running under `Ollama`. Running locally mandates using small models. Compared to frontier cloud models (thirsty, farty, costly, someone else’s), models that run on my hardware are slow, have small context windows, and seem dim. They need all the help they can get, and with the fixed compute resources available, that means they need more patience and more hand-holding. Right now, I’m working on hand-holding by managing context – the crisper my request, the more likely the model's output is valuable. I’m using minimal `pi` as my harness so that I can have some of the advantages of a coding agent without too much of the weight. Here, I’m playing with context management by using tools – the idea is that up-front work by a focussed tool can give the LLM some handholds, using a ‘skill’ to allow the LLM and the agent to interpret the results of the up-front work. As a tester, I want to know about the code structures of the projects I’m working on. We know tools to help with that for large repositories: `DeepWiki` gives them human-readable shape, `SourceGraph` lets humans and agents search, manage and understand them. I want something open-source that can help agents, and I've chosen [graphify](https://graphify.net) . `Graphify`'s up-front work is a deterministic analysis of the code, producing a `graph.json` . That file, and the `graphify` skill to read it, give later LLM-based code queries more power for fewer tokens. Lazily, I asked `pi` to use the `graphify` skill to run `graphify` itself. I was swiftly into the weeds: - turns out `graphify` would very much like to use an LLM, determinism or not. - between them, `pi` and `graphify` insisted on accessing `Ollama` via the wrong URL (`localhost` when the LLM is outside the container), - `graphify` wanted a key, while knowing that `Ollama` needs no key - on the host, requests from `pi`and from `graphify` via ollama into the LLM seemed to clash and the LLM sometimes stopped responding Taking a less-hands-off approach, I ran it successfully from the (container’s) commandline with: `OLLAMA_BASE_URL=http://host.docker.internal:11434/v1 OLLAMA_API_KEY=dummy graphify extract /workspace/flock --global --wiki --as flock --backend=ollama --model qwen3.5:9b --max-concurrency 1 --code-only` Once it was running, it took *ages*. `graphify`’s deterministic piece took a few minutes, but then switched into ‘semantic analysis’, which went silent. Looking at the requests made into `Ollama`, I could see that some were timing out after 10 minutes, and occasionally (less often) `graphify`’s log would tell me that a ‘chunk’ of 100 somethings had timed out. Sometimes the run would just fail with a timeout, sometimes it would end early and not produce anything, and when it finally got to a `graph.json`and a `GRAPH_REPORT.md`, the output was horrible – telling me that the ‘god nodes’ (the most-reference elements) were things like `Hi` and `fd`. Frustratingly, the artefacts were so off that the `graphify` skill typically told `pi` that the analysis hadn’t been done, and the ~~bastards~~ tools started all over again. Have I mentioned that I was working on this on the hottest June day that has yet been measured in the UK, and that the metal machine under my fingertips is running a *local* LLM? I took off my bow tie and my top hat, rolled up my sleeves and finally shut down my 300-tabbed monster browser. 💡 ****Managing context?** I want to see what I'm sending in my context, and one might hope that Ollama would log, somewhere, what it is being asked to do, and what it's returning. Some models advised me I could set `OLLAMA_DEBUG=1`, or even `OLLAMA_DEBUG=2` (and run ollama from the commandline in its owwn terminal) so I could watch the requests– but those didn't offer me much more than timestamped notes that a request had been made or returned. Watching the network with `sudo tcpdump -i lo0 -A 'tcp port 11434'` or `tcpdump -i any -A port 11434` did show me what was going on – and while it isn't readable, lets me see roughly whats in the context. On enquiry, I could see that`graphify` was merrily processing binary as code. `.graphifyignore` let me tell the tool to ignore images and sounds. I could see that `Hi` and obscure friends came from obfuscated built code, and ignored that, too. I recklessly ignored documents and html for good measure. After that work, `graphify` needed 5 minutes working deterministically and an hour or so working with the LLM doing semantic analysis on just 7 files to produce its `graph.json` file. But not the `--wiki` that I'm sure you noted I had asked for, above. I could run `graphify cluster-only /workspace/flock` to make a pretty graph. I did so, and about 90 minutes of LLM-thrashing later, it did. What a very pretty graph – and how irritatingly labelled. Several hundred labels for usefully-*named* elements, most of them *labelled* in the pattern `community nnn`. Anyway… With the analysis data all prepared for the skill, I asked my agent (in several conversations) about tests (I wrote a bit of Flock’s test harness), keyboard navigation and about error handling. My pattern was: ask initial – ask for more depth (based on what had been produced) – write report – \[new convo\] check facts – consolidate report. I skimmed outputs before taking the next step. And the output using this crisp context and this dinky local LLM is… OK. Readable. Correct, on my uninformed view, and more complete and informative than I might have imagined. References to code check out. It feels like it would be useful to me, as a tester. Have a look for yourselves: [error-handling-summaryerror-handling-summary.md9 KBdownload-circle](https://www.workroom-productions.com/content/files/2026/06/error-handling-summary.md "Download") [KEYBOARD\_NAVIGATIONKEYBOARD\_NAVIGATION.md12 KBdownload-circle](https://www.workroom-productions.com/content/files/2026/06/KEYBOARD%5FNAVIGATION.md "Download") I have no immediate idea how much power I’ve used, but let’s say it’s an 80W laptop for 5 hours at full blast, so about 0.4kWh. I’ve not extracted any water from the watertable for cooling, but I’ve had to change my own shirt. ### Workroom PlayTime 060: Sitegeist as an Exploratory Interface URL: https://www.workroom-productions.com/workroom-playtime-060-sitegeist-as-an-exploratory-interface/ Last updated: 2026-06-27T21:22:21.000Z This week's [Workroom PlayTime](https://www.workroom-productions.com/workroom-playtime/) is [Thursday 2 July at 4:00pm London time](https://this-ti.me/?uts=1783004400&tz=Europe%2FLondon&name=Workroom+PlayTime+060). Please **note the change of time**. We'll play with (abandonware) tool [Sitegeist as an Exploratory Interface](https://www.workroom-productions.com/sitegeist-as-an-exploratory-interface/). It needs setup – and I need you to come with it already set up. If you're not already set it up, you won't be able to play. Setup is easy, and you can do it yourself, but I'm happy to help if that helps you come along. [Pop something in my diary](https://savvycal.com/workroomprds/chat) to come and say hello. To set it up, you'll need a machine that has the right permissions to let you install browser extensions, and you'll need a [chromium browser](https://en.wikipedia.org/wiki/Chromium%5F%28web%5Fbrowser%29) (Chrome, Brave, Edge, Opera, Vivaldi and others). I probably won't be able to tell you exactly how to install it on your machine, but I hope I'll be able to be helpful. We'll gather on Zoom. These exercises are for everyone, for free. [All subscribers](https://www.workroom-productions.com/#/portal/signup) get to play together most weeks, paid subscribers get a guaranteed slot or rerun if we're full. If you get this in an email or can see the **joining info** section below, then you're a subscriber and have access to the Zoom etc. for the workshop. If you'd like to bring a friend, you can do that. I'll say yes until I've got too many. Use this Savvycal link to [book your spot](https://savvycal.com/workroomprds/workroom-playtime-060) and get reminders. *Note to regulars: I hope that I've got it working so that not only is there a link to the zoom in this email, but in the "usual zoom" text in the reminder emails that savvycal sends.* ↓ *signed-in subscribers will see joining info and links below* ↓ _This post is for subscribers only._ ### Workroom PlayTime 059: Play with Tenfold URL: https://www.workroom-productions.com/workroom-playtime-059-play-with-tenfold/ Last updated: 2026-06-22T14:10:43.000Z This week's [Workroom PlayTime](https://www.workroom-productions.com/workroom-playtime/) is [Thursday 25 June at 2:30pm London time](https://this-ti.me/?uts=1782394200&tz=Europe%2FLondon&name=Workroom+PlayTime+059). Please **note the time**, which is back to usual. We'll [Play with Tenfold](https://www.workroom-productions.com/play-with-tenfold-by-ink-and-switch/) – where we'll explore a toy made by [Ink and Switch](https://www.inkandswitch.com/). This will be a play-focussed session to see what we can find: their toy is both wide and deep, but doesn't particularly lend itself to exploration by iteratikon nor to the judgement that is fundamental to testing work. Next time, for [Workroom PlayTime 060](https://www.workroom-productions.com/workroom-playtime/#WorkroomPlayTime060), we'll finally get to [Sitegeist as an Exploratory Interface](https://www.workroom-productions.com/sitegeist-as-an-exploratory-interface/). It needs setup – to help you get set up, come to a short session immediately after Thursday's thing, or [pop something in my diary](https://savvycal.com/workroomprds/chat) to come and say hello. You'll need a machine that has the right permissions to let you install browser extensions, and you'll need a [chromium browser](https://en.wikipedia.org/wiki/Chromium%5F%28web%5Fbrowser%29) (Chrome, Brave, Edge, Opera, Vivaldi and others). It's easy, and you can do it yourself, but I'm happy to help if that helps you come play. We'll gather on Zoom. These exercises are for everyone, for free. [All subscribers](https://www.workroom-productions.com/#/portal/signup) get to play together most weeks, paid subscribers get a guaranteed slot or rerun if we're full. If you get this in an email or can see the **joining info** section below, then you're a subscriber and have access to the Zoom etc. for the workshop. If you'd like to bring a friend, you can do that. I'll say yes until I've got too many. Use this Savvycal link to [book your spot](https://savvycal.com/workroomprds/workroom-playtime-059) and get reminders. *Note to regulars: I hope that I've got it working so that not only is there a link to the zoom in this email, but in the "usual zoom" text in the reminder emails that savvycal sends.* ↓ *signed-in subscribers will see joining info and links below* ↓ _This post is for subscribers only._ ### Play with 'Tenfold' by Ink and Switch URL: https://www.workroom-productions.com/play-with-tenfold-by-ink-and-switch/ Last updated: 2026-06-20T16:17:42.000Z We'll play with [Tenfold](https://www.inkandswitch.com/project/tenfold), the interactive software built by [Ink and Switch and collaborators](https://www.inkandswitch.com) to celebrate 10 years of business. This Workroom Playtime is all *play* – I don't have a purpose to keep in mind. We'll just mess with it for 20 minutes, and talk about it as we go. You'll bring your own thing, we'll come to our own thoughts. It's a joy, it's a toy, and it was made with love. If you want to do some homework: [The Tenfold PlaygroundThe Tenfold Playground is a Patchwork tool to help you learn how to love.![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/icon/180x180-07ae85d0-9b21-4dc6-a987-e34dba879f01.png)Ink & Switch![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/thumbnail/playground-5f861340-e6ab-4893-8b57-1df7120d76c6.png)](https://www.inkandswitch.com/project/tenfold/playground/) [TenfoldThe story of Tenfold, a communal art project to celebrate ten years of the lab.![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/icon/180x180-7cd158ef-3f04-40d7-950b-4813e60a642b.png)Ink & Switch![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/thumbnail/L1010569-Edit-16x9-ebb3dac7-fa08-44c2-8b28-2310e69bbbb0.webp)](https://www.inkandswitch.com/project/tenfold/story/) And, if you love it enough, buy the t-shirt [Ink & SwitchRarified collectables, straight from the lab.![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/icon/ouacFRycr4FA04Bt-e33b9f85-f4a2-4309-b0aa-4ca8b047e367.webp)Ink & Switch![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/thumbnail/9x78knR6OSwdEDkG-c6a5a923-a36a-4509-acfc-17f7a5c9f67e.png)](https://merch.inkandswitch.com/en-gbp) ### Workroom PlayTime 058 URL: https://www.workroom-productions.com/workroom-playtime-058/ Last updated: 2026-06-18T13:08:34.000Z This week's [Workroom PlayTime](https://www.workroom-productions.com/workroom-playtime/) is [2:30 London time on Thursday 18 June](https://www.workroom-productions.com/r/e99549aa?m=3bf852c6-77ea-4b7e-9876-5f59cf1f1136). We'll do [Ever-Rolling Stream](https://www.workroom-productions.com/exercises-about-testing-and-systems-analysis/#ever-rolling-stream) from [Exercises about Testing and Systems Analysis](https://www.workroom-productions.com/exercises-about-testing-and-systems-analysis/). We'll gather on Zoom. These exercises are for everyone, for free. [All subscribers](https://www.workroom-productions.com/#/portal/signup) get to play together most weeks, paid subscribers get a guaranteed slot or rerun if we're full. If you get this in an email or can see the **joining info** section below, then you're a subscriber and have access to the Zoom etc. for the workshop. If you'd like to bring a friend, you can do that. I'll say yes until I've got too many. ↓ *signed-in subscribers will see joining info and links below* ↓ _This post is for subscribers only._ ### Exercises about Testing and Systems Analysis URL: https://www.workroom-productions.com/exercises-about-testing-and-systems-analysis/ Last updated: 2026-06-18T12:43:20.000Z Ideas for exercises, in various states of completion. Crucially, they're based in ideas of systems thinking. This is an attempt to write exercises that explore the ideas of testers as systems analysts. They're primarily designed for online use, for a handful of playful testers, and should last \~20 mins. Several have the potential to last much longer. _This post is for subscribers only._ ### Workroom PlayTime 057: Disaster Story URL: https://www.workroom-productions.com/workroom-playtime-057/ Last updated: 2026-06-11T13:58:17.000Z This week's [Workroom PlayTime](https://www.workroom-productions.com/workroom-playtime/) is [Thursday 11 June at 3:30pm London time](https://this-ti.me/?uts=1781188200&tz=Europe%2FLondon&name=Workroom+PlayTime+057). Please **note the time**, which is an hour later than usual. We'll do [Disaster Story](https://www.workroom-productions.com/playful-exercises-about-testing/#disaster-story) from playful exercises about testing. You'll note that we're not doing last week's promised longer session with Sitegeist. Soon. We'll gather on Zoom. These exercises are for everyone, for free. [All subscribers](https://www.workroom-productions.com/#/portal/signup) get to play together most weeks, paid subscribers get a guaranteed slot or rerun if we're full. If you get this in an email or can see the **joining info** section below, then you're a subscriber and have access to the Zoom etc. for the workshop. If you'd like to bring a friend, you can do that. I'll say yes until I've got too many. Use this to [book your spot](https://savvycal.com/workroomprds/workroom-playtime-057) and get reminders. ↓ *signed-in subscribers will see joining info and links below* ↓ _This post is for subscribers only._ ### No Workroom PlayTime 28 May URL: https://www.workroom-productions.com/no-workroom-playtime-28-may/ Last updated: 2026-05-25T21:43:07.000Z Hi all – there's no Workroom PlayTime this Thursday. Unexpected family commitments mean that I no longer have the slot on the Thursday, there's no time to spare on the Friday, and I'm too late in the week to usefully pull it sooner. We'll do [Sitegeist as an Exploratory Interface](https://www.workroom-productions.com/sitegeist-as-an-exploratory-interface/), and will focus on using an LLM-connected browser extension to help us to explore. ### Workroom PlayTime 056: Switching for Explorers URL: https://www.workroom-productions.com/workroom-playtime-056-switching-for-explorers/ Last updated: 2026-05-19T10:16:21.000Z This week's [Workroom PlayTime](https://www.workroom-productions.com/workroom-playtime/) is [Thursday 21 May at 2:30pm London time](https://this-ti.me/?uts=1779370200&tz=Europe%2FLondon&name=Workroom+PlayTime+056). We'll do [Switching for Exploratory Testers](https://www.workroom-productions.com/switching-for-exploratory-testers/), noticing how we switch methods while exploring, and bringing that back to software testing. That will be the end of the series: I'll gather the materials and offer an extended set in June. [Next week's workshop](https://www.workroom-productions.com/sitegeist-as-an-exploratory-interface/) will be different: not an exercise, longer and needing some preparation your end. It'll be on the weird, new, and already-possibly-orphaned browser extension [Sitegeist](https://sitegeist.ai), which acts as an interface between your web browser and an LLM, so can act as an [exploratory interface](https://www.workroom-productions.com/exploratory-interfaces/) when we're testing anything web-page ish. And it can work with a Local LLM, which floats my boat. On the subject of Local LLMs, I'm delighted to to let you know that Bart Knaack and I will be running a hands-on workshop called [Support your Testing with Local LLMs](https://agiletestingdays.com/2026/session/support-your-testing-with-local-llms/) at [Agile Testing Days](https://agiletestingdays.com/) in November. I'm *so* pleased to be building that for ATD's excellent crowd – Bart and I are learning a lot. We'll gather on Zoom. These exercises are for everyone, for free. [All subscribers](https://www.workroom-productions.com/#/portal/signup) get to play together most weeks, paid subscribers get a guaranteed slot or rerun if we're full. If you get this in an email or can see the **joining info** section below, then you're a subscriber and have access to the Zoom etc. for the workshop. If you'd like to bring a friend, you can do that. I'll say yes until I've got too many. Use this to [book your spot](https://savvycal.com/workroomprds/workroom-playtime-056-switching-for-explorers) and get reminders. ↓ *signed-in subscribers will see joining info and links below* ↓ _This post is for subscribers only._ ### Sitegeist as an Exploratory Interface URL: https://www.workroom-productions.com/sitegeist-as-an-exploratory-interface/ Last updated: 2026-07-02T18:30:11.000Z In this session, we'll play with the browser extension [Sitegeist](https://sitegeist.ai). I think it's interesting as an exploratory interface to anything in a web page. It's not unique, but it's on a sweet spot between accessible and powerful. So let's play with it and see what we find. This will be a longer session – I imagine an hour. You'll need to install Sitegeist as a browser extension – and for that you'll need a Chromium browser. You've probably already got Chrome. I've put it into Brave. You'll also either need a key for an LLM, or a local LLM. I'll set up a limited key for you, if you ask. So Sitegeist is a way to give the website you're looking at, to an LLM, along with your thoughts. It can do stuff through the UI, it can directly affect the DOM, it can read and run the JavaScript in the page directly / through user scripts / through the debugger. It let you look at what you've been doing, automates experiments if you ask, has memories and sessions behind the scenes, can make dashboards, can iterate and analyse, lets you set up skills for individual sites. For giggles, I opened [Puzzle 38](https://www.workroom-productions.com/puzzle-38/) and asked a Claude Haiku via the Sitegeist extension to hypothesise what was going on, to test that hypothesis, and to stop when it had a testable hypothesis that it could summarise in 300 characters. Took about 5 minutes of blinkenlighten, needed a nudge, then came back with a plausible description of the underlying principle. It clearly has handy tricks for exploratory testers, and I don't yet know what they are. ## Exercise Let's start with [Puzzle 38](https://www.workroom-productions.com/puzzle-38/) . - ask Sitegeist what events the UI responds to, and how it shows responses - ask it how it knows, if you're not satisfied with what it tells you - ask it to summarise all that in a markdown. - Download the `.md`. and start a new conversation. - Ask Sitegeist to build a tool to let you specify a sequence of L / R buttons, and to automate clicking that sequence, using the info in the markdown to help. *This is an exploratory interface. Maybe ask it to capture the lamps after the clicking is done, too. Maybe ask it for a speed parameter.* - try sequences to your own design - ask the thing to build a 'complete' set of sequences, iterate over them capturing the outputs, analyse the outputs with the aim of creating hypotheses about the way the buttons affect the lamps, then test those hypotheses by seeking disconfirmation as well as confirmation. - when it settles, ask it to check the code. Let's go to EvilTester's [Basic Shopping Cart](https://testpages.eviltester.com/apps/basiccart/?page=1&limit=10) app - ask it to observe as you explore. *This is the exploratory interface* - explore. Tell it what you note. *This is another way in* - ask it to summarise what it saw, and what you mentioned. - identify areas that are of interest. Ask it to explore those areas, and to report back with repeatable evidence. --- ## Sitegeist It's a browser extension for Chomium browsers. It was released in October 2025, and went open-source and (basically) unmaintained in March 2026. [Mario Zechner (@mariozechner.at)In Oct. 2025, I created sitegeist.ai, a browser agent extension for Chrome. Here it is beating the crap out of OpenAI Atlas by cheating. It is now OSS. Go forth, and fork/remix/make it your own. https://github.com/badlogic/sitegeist![](https://static.ghost.org/v5.0.0/images/link-icon.svg)Bluesky SocialMario Zechner![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/thumbnail/thumbnail-342aa37ce1ef8c202f6c6c3feca5590cf0fc549f751f571ebd732268a95f1e28.jpg)](https://bsky.app/profile/mariozechner.at/post/3mh4zle3jt22t) It was made by [Mario Zechner](https://mariozechner.at), who made [pi.dev](https://pi.dev), (the LLM coding agent lying under OpenClaw). Here are his [notes](https://mariozechner.at/posts/2025-12-22-year-in-review-2025/#toc%5F9) on Sitegeist – and also information about why he (at that time planned to take / subsequently took) it open-source. Mario left Sitegeist behind in March, and [took Pi with him](https://mariozechner.at/posts/2026-04-08-ive-sold-out/) into [Earendil](https://earendil.com) in April. --- So: what are we going to do? I'm not sure yet – but I'll scribble here as I try stuff out, you'll think of and try other things, and together we'll see what areas have potential, and what is just smoke and jelly. Try first: - a set of experiments across a range - summarising results of experiments and setting up more based on summary - a skill for a specific page - a skill for a specific class of oddness ### Workroom PlayTime 055: Iterating for Exploratory Testers URL: https://www.workroom-productions.com/workroom-playtime-055-iterating-for-exploratory-testers/ Last updated: 2026-05-04T13:27:40.000Z Happy [May Day](https://en.wikipedia.org/wiki/May%5FDay) / Beltane / Walpurgis Night / International Workers' Day! This week's [Workroom PlayTime](https://www.workroom-productions.com/workroom-playtime/) is [Thursday 7 May at 2:30pm London time](https://this-ti.me/?uts=1778160600&tz=Europe%2FLondon&name=Workroom+PlayTime+055). We'll do one of the exercises in [Iterating for Exploratory Testers](https://www.workroom-productions.com/iteration-exercises-for-exploratory-testers/), thinking about how we step through sets while testing. Next week, on 14th May, we'll run [Switching for Exploratory Testers](https://www.workroom-productions.com/switching-for-exploratory-testers/), which may be the end of this series for now. Later in the summer, I'm planning a series of Black Box Puzzles: Here's [Puzzle 38](https://www.workroom-productions.com/puzzle-38/), in case you've not come across the puzzles before. We'll gather on Zoom. These exercises are for everyone, for free. [All subscribers](https://www.workroom-productions.com/#/portal/signup) get to play together most weeks, paid subscribers get a guaranteed slot or rerun if we're full. If you get this in an email or can see the **joining info** section below, then you're a subscriber and have access to the Zoom etc. for the workshop. If you'd like to bring a friend, you can do that. I'll say yes until I've got too many. Use this to [book your spot](https://savvycal.com/workroomprds/workroom-playtime-055-iterating-for-exploratory-testers) and get reminders. ↓ *signed-in subscribers will see joining info and links below* ↓ _This post is for subscribers only._ ### Switching for Exploratory Testers URL: https://www.workroom-productions.com/switching-for-exploratory-testers/ Last updated: 2026-05-19T12:50:37.000Z *Trying to see how we switch approach while discovering things* When we explore, we might be working with a particular method until it's worth switching to another method. We might name the method, and consciously switch, or we might futz about and see what happens. We might futz about until we have named methods and deliberate switching, or we might get fed up with our patterns of work and drift to futzing about. In this exercise, we'll play with revealing information. In [*Raster Reveal*](https://www.workroom-productions.com/raster-reveal/)*,* you play with a block of colour to show details, and watch your mind to see when it sparks into hypothesis about what picture is revealed. In this exercise, a variant on Raster Reveal, we'll consciously limit our approaches, and observe what triggers a switch from one to another. For each exercise, you'll wave your mouse / finger over the block. We'll all do the same block at the same time. We'll all be revealing the same picture. We won't talk until the end of each exercise. ## Approaches - Dragging across the image in straight lines (of direction / spacing / length) - Making circles across (some part) - Making a spiral into (some part) - Quartering an area - Defining a perimiter - Digging into a detail - Opening a big, untouched bit ## Exercise 01 As you work, consider how you're approaching your exploration. How and where / why did you switch approach? ## Exercise 02 As you work, try to limit your approaches to the ones above, and limit those approaches to particular parameters. How and where / why did you switch approach? ## Exercise 03 Try to only use each approach once. What do you start with, and why? #### Example Start with a big slow spiral from outside to centre, do short diagonal lines, then little circles over areas of interest How and where / why did you switch approach? ## Exercise 04 this is a spare ## Conclude How dod you decide to switch approach? What was your experience of revelation in this experiment? Drawing from your experience just now, what do you recognise about how your mind works, when you recall your work exploratory testing? ## ## **Frameworks** ### **Stopping Heuristics for Exploratory Testing** Let’s distinguish switching from stopping. You switch when you cold be doing something better. Stopping is when you’ve run out of resource. A great way to run out is when you’ve set yourself a small budget – a timebox, a stack of discoveries: running out means it’s time to move on and switch to a different location or approach. Here’s my non-exhaustive list of situations to stop, and here’s a better [list from Michael Bolton and James Bach](https://developsense.com/blog/2009/09/when-do-we-stop-test): - out of resource: time / money / licenses - found enough (by number) - found enough (a big-enough problem) - found nothing of note after some time (before out of resource?) - answered all the outstanding questions (you did *have* questions?) ### Workroom PlayTime 054: Getting Stuck for Explorers URL: https://www.workroom-productions.com/workroom-playtime-054-getting-stuck-for-explorers/ Last updated: 2026-04-27T14:06:18.000Z This week's [Workroom PlayTime](https://www.workroom-productions.com/workroom-playtime/) is [Thursday 30 April at 2:30pm London time](https://this-ti.me/?uts=1777555800&tz=Europe%2FLondon&name=Workroom+PlayTime+054). We'll do [Getting Stuck for Explorers](https://www.workroom-productions.com/getting-stuck-for-explorers/) to think about being stuck when exploring, and what we might do about it. Next week, on the 7th May we'll do (one exercise in) [Iterating for Exploratory Testers](https://www.workroom-productions.com/iteration-exercises-for-exploratory-testers/). There are currently ten there, and a stack of notes for subscribers, so that's up for some refinement. Take a look: ping me if an exercise appeals to you. Or just laugh at my scattergun process. We'll gather on Zoom. These exercises are for everyone, for free. [All subscribers](https://www.workroom-productions.com/#/portal/signup) get to play together most weeks, paid subscribers get a guaranteed slot or rerun if we're full. If you get this in an email or can see the **joining info** section below, then you're a subscriber and have access to the Zoom etc. for the workshop. If you'd like to bring a friend, you can do that. I'll say yes until I've got too many. Use this to [book your spot](https://savvycal.com/workroomprds/workroom-playtime-054-getting-stuck-for-explorers) and get reminders. ↓ *signed-in subscribers will see joining info and links below* ↓ _This post is for subscribers only._ ### Iterating for Exploratory Testers URL: https://www.workroom-productions.com/iteration-exercises-for-exploratory-testers/ Last updated: 2026-05-07T13:30:10.000Z *still being built – trust nothing* Exploration needs iteration: if you stay in one place you’ll run out of new things to explore. Here are a collection of exercises that might help us think about how we iterate while doing exploratory testing. This is a growing collection of rough exercises. I'll refine and add over time. For Workroom PlayTime, we'll pick one, and I expect that we'll return to this over time. For [Workroom PlayTime 055](https://www.workroom-productions.com/workroom-playtime-055-iterating-for-exploratory-testers/), we'll do... ## Exercise 3 – What Iterables? An `iterable`, in this exercise, is a set of similar things that you can loop over, doing much the same thing to each, judging against similar obstacles. #### Examples - Where do all the links on this page go? - Can I open and play an example of each of these audio file formats? - If I take all available routes, what does my map of connections look like? *5 minutes* Think of how you might iterate over several of the following while exploring. - menu choices in app - buttons available on screen - regional variations in regulation - last week's fixes - error messages Share with a public sticky note when done. You might need to set out your context to make sense of your answer – you might have a different context for each, too. *5 minutes* Privately write down at least six more iterables that you might iterate over when exploring. On public sticky notes, share just three of those. *10 minutes* Share ways you might iterate over everyone els's new iterables. While doing that, note clusters in what y0u're acting on, the action you're taking, the observations you might make. --- ## Exercise 1 – Telling Stories about Testing *5 minutes* Write down several stories about exploratory testing. Stories should be swift and short: you need three or more. *5 minutes* For those stories, consider where you start, and how you move, how you steer, where you stop, what you cover. *10 minutes* Share some differences between your own stories. Explore some differences between each other's stories. ## Exercise 2 – Iteration Spotting *This is written as an introspective solo exercise. I think it would be better as a pair exercise; one person exploring, the other thinking of iterables that the first is using.* Explore {something}. Watch to see where you have a sense of iterating – you find a set, or a range, or something where you'll try a collection. As you explore, consider what characteristics the iterable has, and the method of iteration. Talk, conclude, share. ## ## Exercise 4 – Edges *10 mins* Pick one of the questions below. Write an answer to help you think - Iterating consumes something: starts with a pile of unused stuff, and transfers items over to a pile of used stuff. Why might you go back to some things that you have already used, when iterating? - What happens if your iteration approach (i.e. how you're stepping through / picking next) changes your results? Think of an example: - Do you 'steer' in an iteration? How do you steer? - How can you manage iteration across more than two variables? - What is the relationship between sampling and iterating in your own exploratory testing? 5 mins Read everyone else's Q&A. 5 mins Dot vote two of the Q&As – those with the most dots can volunteer to take / answer questions on their treatment. ## Exercise 5 – Internal / External Split into two groups. The groups do not have to be the same size. Discuss what makes an internal iterable, and what makes an external iterable. One group writes down as many internal iterables as possible. The other writes down as many external iterables. Switch roles. Swap lists, if you like. Do the other one. Share results. ## Exercise 6 – Common Iterables for Exploratory Testers See [tours](#tours) (below). Better: try {an artefact} and imagine what you'd do on {tours} ## Exercise 7 – Completion and Sampling Consider something real that you have explored – preferably software that you have tested. What is a ‘set’ that you can try every piece of? What is a ‘set’-ish that you can only try some of? How do you choose what to try? ## Exercise 8 – tools in a loop What tools can you use to iterate? What can you iterate over? Build an iterator (for real if you can, if you can't then do it as a thought experiment) - how do you read the results? - in what way is that iteration exploration? ## Exercise 9 – Steering Using {artefact} and at least one thing to iterate over, start exploring – and identify where you need to change the step size / direction. Resist – or give in. Why? Share your insights into the differences between iterating on a path, and steered iteration. ## Exercise 10 – Hold still Using {artefact}, iterate over {range} with {tool}. Now use {tool} to hold a value steady while everything else changes. Share how the different approaches to exploration felt. ## Exercise 11 – Lost toys Imagine you're looking for something – you know what it is, you had it in your hand a moment ago. It's got to be here somewhere. Now you can't see it. - How do *you* go about looking for it? - Write a short set of instructions for a child, to help them look for a misplaced toy in a similar circumstance. Doesn't have to b the same, if your approach works. - Pivot to software testing: in what situation might (one or both) of those approaches work? What are you iterating over? - Share that situation (and the approach) to the group, with a specific example if you have one. ## ## Sources and thoughts ### An Iterable has... My vocabulary: An *iterable* – something one can *iterate* over – will have individual *things*. Those things will be related in some way to make a *set*. A *range* is a one-dimensional set. Iterating on a range typically means going from one end to the other. More sophisticating iterations may go exponentially, or hunt by binary ?splitting. Examples of *discontiguous ranges* might be integers from 0 to 9, month names from January to December. Those have serially-related individuals. A *continuous range* might be numbers from 0 to 9 which has a delimited but infinite set of values (note that this is theoretical: when represented in a machine, that representation imposes quantisation). A *range* may have behaviour changes: imagine an integer age range for car insurance. Some *set*s may be multi-dimensional. See ESH variables in Explore It! for deeper. ?Set of sets? Some *set*s may be populations – made of individuals with loose relationships. Iterating involves picking an unused individual. Examples are functions, or navigation. We can iterate over `input`s and `output`s, of course – and with forethought we can iterate over `state`s, `record`s and other more internal representations. We can iterate over things that aren't behaviours, but do belong to the system – lines of code / decisions / data representations. We might also iterate, as exploratory testers, over externals – requirements, examples, tests. Every iteration has a sense of done-ness or coverage – and the many measures of coverage give a sense of the different ways we can iterate. I think of exploratory surfaces – a way that the context invites us to iterate. Behaviour is one surface. User inteface another. An API or CLI another. Code another. Bugs another. Change history another. Logs. Data. On we go. Step back. What *exercises*? If a range - Somewhere to start - A sense of 'step' – maybe regular, maybe - Somewhere to stop – a number of steps, a maximum or minimum size. If a population - a picker of some sort - a way to tell which individuals in the population have been tried If multi-dimensional, then what variables – or perhaps think in a different way ### Tours Just search for it. You'll find a stack; they won't cover the gamut of what you can think of, but may help you think of gaps. Make a list whichever ones appeal to you, and a couple that raise your eyebrows. Keep the list somewhere you can see all of them at once when you want to. Here are links to the ideas in [Ch 4 of Exploratory Software Testing](https://learning.oreilly.com/library/view/exploratory-software-testing/9780321647863/ch04.html) (Whittaker, 2009), [Ch 3 of Taking Testing Seriously](https://learning.oreilly.com/library/view/taking-testing-seriously/9781394253197/c03.xhtml) (Bach, Bolton 2025). They're in my materials from 2005 or so, so mush have been in common use among the gang earlier – it's not exactly a big metaphorical leap from "exploratory" to "tourism"... ### Stopping Heuristics for Exploratory Testing Let’s distinguish being stuck from stopping. Stuck is when you can’t move despite (something). Stopping is when you’ve run out of resource. A great way to run out is when you’ve set yourself a small budget – a timebox, a stack of discoveries: running out means it’s time to move on and switch to a different location or approach. Here’s my non-exhaustive list of situations to stop, and here’s a better [list from Michael Bolton and James Bach](https://developsense.com/blog/2009/09/when-do-we-stop-test): - out of resource: time / money / licenses - found enough (by number) - found enough (a big-enough problem) - found nothing of note after some time (equivalent to out of resource – but time's the one you can't buy) - answered all the outstanding questions (you did *have* questions...) ### Software Testing Live 07 URL: https://www.workroom-productions.com/mot-testing-live-07/ Last updated: 2026-04-23T12:53:23.000Z We'll test software that makes crossword / wordsearch grids: Give it words, and it will try to put those words onto a grid. ## For the workshop Target: [GridBuilder](https://exercises.workroomprds.com/tt%5Fmot%5Ftarget/index.html) Artefacts: [Tests](https://exercises.workroomprds.com/tt%5Fmot%5Ftarget/testRunner.html). [Coverage](https://exercises.workroomprds.com/tt%5Fmot%5Ftarget/coverage/index.html). [agents.md](https://exercises.workroomprds.com/tt%5Fmot%5Ftarget/AGENTS.md). [User instructions](https://exercises.workroomprds.com/tt%5Fmot%5Ftarget/userInfo.md). [Build log](https://ampcode.com/threads/T-019dba33-eed3-77ee-94c5-0e71717f3e06) via Amp. Here's a [shared Miro board](https://miro.com/app/board/uXjVHdtKsgY=/?share%5Flink%5Fid=138048162081) – we can use it to share notes etc. ## More MoT's page [Software Testing Live: Episode 07 - Testing transparentlyExplore the boundaries of AI-generated software by live-testing a word search tool and using collaborative techniques to identify gaps in logic despite passing all initial automated tests.![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/icon/favicon-2d674442d2b5d949f41d3acf8dcde2e6ac434709591a8b71f7fd7acc4e89f803.ico)Ministry of Testing![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/thumbnail/523uyhqvfm4p9med4xr6g2l93g5f)](https://www.ministryoftesting.com/events/software-testing-live-episode-07-testing-transparently) --- ## Backup version (working before workshop) Target: [GridBuilder](https://exercises.workroomprds.com/tt%5Fmot%5Fbackup/index.html) [Tests](https://exercises.workroomprds.com/tt%5Fmot%5Fbackup/testRunner.html). [Coverage](https://exercises.workroomprds.com/tt%5Fmot%5Fbackup/coverage/index.html). [agents.md](https://exercises.workroomprds.com/tt%5Fmot%5Fbackup/AGENTS.md) [Build log](https://ampcode.com/threads/T-019daf67-92d1-744d-ba2e-2ec616b178b3) via Amp. A [populated grid](https://exercises.workroomprds.com/tt%5Fmot%5Fbackup/?words=PEA%2COVERDONE%2CRAPIDLY%2CCOPENHAGEN%2CFORTINBRAS%2CREPTILE%2CNOTIFY%2CARRANGEMENT&width=10&height=10&all%5Fcross=true&allow%5Fdiagonals=false&allow%5Freverse=false): ## ### Workroom PlayTime 053: Play Styles and Exploratory Work URL: https://www.workroom-productions.com/workroom-playtime-053-play-styles-and-exploratory-work/ Last updated: 2026-04-23T13:07:20.000Z This week's [Workroom PlayTime](https://www.workroom-productions.com/workroom-playtime/) is [Thursday 23 April at 2:30pm London time](https://this-ti.me/?uts=1776951000&tz=Europe%2FLondon&name=Workroom+PlayTime+053). We'll do [Play Styles and Exploratory Work](https://www.workroom-productions.com/play-styles-and-exploratory-work/), to consider what exploratory testing work might suit particular styles of play. You might also want to know that I'm doing a 90-minute '[Exploratory Testing live](https://www.ministryoftesting.com/events/software-testing-live-episode-07-testing-transparently?s%5Fid=19039713)' session on MoT, also on Thursday, at 6:30pm London time. I'll generate some software, it'll pass its automated tests, then we'll all test it. And on the 30th (I know! *next* week – planning!) we'll do [Getting Stuck for Explorers](https://www.workroom-productions.com/getting-stuck-for-explorers/). We'll gather on Zoom. These exercises are for everyone, for free. [Free subscribers](https://www.workroom-productions.com/#/portal/signup) get to play together most weeks, paid subscribers get a guaranteed slot or rerun if we're full. If you get this in an email or can see the **joining info** section below, then you're a subscriber and have access to the Zoom etc. for the workshop. If you'd like to bring a friend, you can do that. I'll say yes until I've got too many. Go here to [book your spot](https://savvycal.com/workroomprds/workroom-playtime-053-play-styles-and-exploratory-work) and get reminders. ↓ *signed-in subscribers will see joining info and links below* ↓ _This post is for subscribers only._ ### Play Styles and Exploratory Work URL: https://www.workroom-productions.com/play-styles-and-exploratory-work/ Last updated: 2026-04-23T13:14:07.000Z In [Exploring Play Styles](https://www.workroom-productions.com/exploring-play-styles/), we considered how we like to play, and how that tendency might influence our exploration. In this exercise, we try a scenario to see what exploratory work might suit particular styles of playing. *For Workroom PlayTime 053: a* [*Miro Board*](https://miro.com/app/board/uXjVHduQuuA=/?share%5Flink%5Fid=229698717600) ## Exercise *10 mins, solo or together* You're able to recruit two people to your team\*, to focus on exploratory testing. None of them know your product, all know the business, and you've got a toolsmith to support them technically. You've got several candidates to choose from. No need to read them all (and do write your own if you prefer). Pick two, and write down some specific exploratory tasks you'd ask them to do (note: I've *not* made a list for you to pick from – your project and your tasks and your imagination are yours) **Why did you align these characters with those tasks?** *\`\* make this real if possible: a current team, or a team you can recall.* ## Debrief *10 minutes* Share, consider differences in the group, identify things you want to take away. OR Duplicate and fill in cards on the board --- ## Characters ### **Geoffrey** *Plays to win.* Geoffrey organises weekly poker nights (for money), and while he loves to win for real cash, donates those winnings to his sister's education. At work, he's proud of his custom project dashboard and loves to see his team's metrics improve – not only over time, but until they are indisputably better than his rivals. ### **Yuki** *Loves adapting and managing the unanticipated.* Yuki's a regular at a role-playing game cafe, and likes weekend city breaks with a localised escape room. At work she's know for staying late to dig deep into problems that others avoid. ### **Devon** *Plays with friends.* Devon organises quiz nights, and loves to set up questions for her friends' specialised knowledge (the more obscure the better). Devon has contacts throughout the organisation, and once restructured a workflow simply to bring two teams closer. ### **Priya** *Introspective, seeks growth through challenges.* Priya is a regular at an open jazz night that has a reputation for attracting passing musicians. At work, she's known for her enthusiasm for moving to different roles and departments, especially if it's not something she's done before. ### **Josh** *Playfulness is the only thing Josh takes seriously.* Josh is famous for his party games – never the same twice, and reliably hilarious. At work, he's got a history of long-form, elaborate pranks that stay (just) within professional limits. Teams compete to include Josh, although his work is not exceptional. ### **Aisha** *Experiments to learn.* Aisha's side hustle is a series of self-published fantasy novels, each in a different style but whose central character always has the same name. At work, she is happiest building prototypes and trying tools, and makes a habit of showing new people the ropes. ### Oliver *Plays to progress.* Oliver spends his weekends in junk shops and online, collecting sets and occasionally trading for more rare collectibles. At work, he is known for making great relationships with clients and suppliers – he works to understand their needs and problems and they trust him to tell their story. His sales increase every quarter. ### Sarita *Loves to experience shifts in her world view.* Sarita rides rollercoasters and games in VR when not campaigning for any number of political causes. She wants to have her 'mind blown' by new ideas and concepts, and is a regular and valued bringer of new ideas and 'what if's. --- ## Sources The characters have characteristics inspired by the two models of play styles #### Brian Sutton-Smith's Seven Rhetorics of Play - ****Progress** – growth is fun. Numbers going up. - ****Fate** – unpredictability (and one's reaction to it) is fun. Chance. - ****Power** – being better is fun. Winning. - ****Community identity** – feeling closer to others is fun. Group-making. - ****Imaginary** – experiments are fun. Safe learning. - ****Self** – discovering oneself is fun. Self knowledge. - ****Frivolity** – silliness is fun. Absurdism. More from [House of Nerdery](https://houseofnerdery.com/2025/06/04/exploring-play-sutton-smiths-seven-rhetorics/), [summary](https://sk.sagepub.com/ency/edvol/play/chpt/rhetorics-play-suttonsmith) from **Encyclopaedia of Play in Today's Society*, bio in [Wikipedia](https://en.wikipedia.org/wiki/Brian%5FSutton-Smith). #### Roger Caillois's Categories and Polarities ****Categories** - Competition (`Agon`) - Chance (`Alea`) - Mimicry (`Mimesis`) - Perception shift (`Ilinx`) ****Polarities** - Uncontrolled, improvised: rules change during game (`Paidia`) - Rewards skill, effort, patience, strategy: rules set before (`Ludus`) See [Wikipedia](https://en.wikipedia.org/wiki/Roger%5FCaillois) #### Possible play style by character - ****Geoffrey**: **Power, competition and crisp rules* - ****Yuki**: **Fate, chance and emergent rules* - ****Devon**: **Community identity and mimicry matter more than rules* - ****Priya**: **Self-discovery, mimicry and emergent rules* - ****Josh**: **Frivolity, perception shift and testing rules* - ****Aisha**: **Imagination and (setting / teaching) rules* - ****Oliver**: *Progress, mimicry / empathy and collaborative rules* - ****Sarita**: **Perception shift and emergent rules* ### Getting Stuck for Explorers URL: https://www.workroom-productions.com/getting-stuck-for-explorers/ Last updated: 2026-04-30T13:16:09.000Z How do you get *stuck* when exploring? More to the point, what does that tell use about getting *unstuck*? Exploration needs iteration: if you stay in one place you’ll run out of new things to explore. If stuckness is important to explorers, perhaps it’s important to exploratory testers. For Workroom PlayTime 054 – a [Miro board](https://miro.com/app/board/uXjVHaUn7Ac=/?share%5Flink%5Fid=925395031451) ## **Exercise** *5 minutes – prime your mind* Write down as many synonyms for being stuck as you can. See what metaphors they live in. *5 minutes – refine and aim* Think of ways that software exploration *doesn’t* match those metaphors. How does stuck feel there? Think of more synonyms and phrases that better suit. *10 minutes – get specific, share and find useful stuff* Share and classify metaphors around what’s happening when you explore. Share ways that you have used to get un-stuck. ## **Frameworks** #### ****example metaphors** (if you're stuck...) - stuck in the mud / in a rut, mired, gridlocked, dead end, hit a wall - lost, wandering, down a rabbit hole, sidetracked, can’t see the forest for the trees - in too deep, stuck in quicksand, drowning in details - tangled up, boxed in, trapped - brain fog, overthinking, analysis paralysis, fixated - gummed up, seized, rusty, - stunted, ossified, stagnant, ossified - burnt out, out of fuel, out of steam, sluggish ### **Inverting Motivators** In [Drive](https://en.wikipedia.org/wiki/Drive:%5FThe%5FSurprising%5FTruth%5FAbout%5FWhat%5FMotivates%5FUs), Dan Pink identifies the families of motivators: mastery, autonomy and purpose. This is taken from [Deci and Ryan’s Self-determination Theory](https://en.wikipedia.org/wiki/Self-determination%5Ftheory), which had autonomy, competence and relatedness. Matching demotivators might be incompetence / confusion / overwhem / futility, dependency / prescribed methods / powerlessness, meaninglessness / disconnection / absurdity. ### **Setbacks** [Amabile / Kamer’s Progress Principle](https://progressprinciple.com/progress-principle/) supposes that small setbacks can have a large effect when under pressure – and conversely, tiny steady achievements make an outsized difference when stuck. ### **Cognitive load** Sweller’s [theory of cognitive load](https://en.wikipedia.org/wiki/Cognitive%5Fload) tells us that we have a limited capacity, and get stuck when we have too much on our minds. A given situation has a certain *intrinsic* cognitive load, we bring an individual *germane* load to processing it, and there is an *extraneous* load in how it is presented. If we’re stuck but can adjust how we’re interacting with the artefact, we might get unstuck. ### **My Patterns of Stuckness** I get stuck when: - I’ve not found *enough* weidnesses – and when I’ve found *too much*. Typically I can’t decide what to do next, maybe becuase in the first instance I don’t have enough signal, and in the second I have too much noise. - it’s too long between weirdnesses. Typically I question my process and try to think of something different. - the weirdnesses show no hint of determinism, or when it’s not clear what the weirdnesses might mean. Seeing noise, not pattern, is a demotivator. - I can’t decide which weirdnesses is weirdest. - I can’t get my tools to work, and realise I’ve spent more time exploring my buggy tools than I have exploring the subject. ### Getting unstuck For me: - notes are a key to unlock stuckness – I can put information into notes for review and as a memory, use them to trigger new approaches, switch from a linear narrative to a growing picture and more. - I think of process goals (taking regular technical action / putting one foot in front of the other) rather than progress goals (getting to an endpoint / winning a race). - Accepting uncertainty is a reliable way to innovation. - Recognising my triggers (finding too few problems, feeling unqualified to judge) helps me move past. ### **Stopping Heuristics for Exploratory Testing** Let’s distinguish being stuck from stopping. Stuck is when you can’t move despite (something). Stopping is when you’ve run out of resource. A great way to run out is when you’ve set yourself a small budget – a timebox, a stack of discoveries: running out means it’s time to move on and switch to a different location or approach. Here’s my non-exhaustive list of situations to stop, and here’s a better [list from Michael Bolton and James Bach](https://developsense.com/blog/2009/09/when-do-we-stop-test): - out of resource: time / money / licenses - found enough (by number) - found enough (a big-enough problem) - found nothing of note after some time (before out of resource?) - answered all the outstanding questions (you did *have* questions?) ### How do LLM Makers Assess their Models? URL: https://www.workroom-productions.com/how-do-llm-makers-assess-their-models/ Last updated: 2026-06-23T16:18:55.000Z Anthropic, OpenAI, Google and others release System Cards / Model Cards for their LLMs. These documents describe a model's capabilities, limitations, and set out the evidence that the makers use for those claims. The [system card for Anthropic's Claude 3.5](https://www-cdn.anthropic.com/fed9cc193a14b84131812372d8d5857f8f304c52/Model%5FCard%5FClaude%5F3%5FAddendum.pdf) (June 24) has 6 pages of assessment and benchmarks from reasoning to safety. The [system card for Claude Mythos Preview](https://www-cdn.anthropic.com/08ab9158070959f88f296514c21b7facce6f52bc.pdf) (April 2026) has 200+ pages, on far more – including risks of deployment, honesty and evasion, responses to an irritating repetitive prompt, an analysis by a psychologist. Those cards describe the model – and as some aspects of a model are emergent, they include results of testing and exploring. The systems cards give us hints about how the makers have approached that work. They're illuminating for testers – not only in describing the models, but in describing how the makers seek to sense their models. We might learn from them, understanding the breadth and aims of their exploration of whatever it is that they've made. ## Exercise *10 mins – solo* There's no realistic hope of being able to understand a model card in a few minutes. So let's just dive in, and find something that we can share with other testers. Pick a model card for a model you've used (or have heard of). Skim it. Pick out something, of interest to you as a tester, that you'd like to share. *10 mins – collective* Let's exchange what we've found and talk. What are the surprises, and what are the signals. ## Sources [Model system cardsAnthropic is an AI safety and research company that’s working to build reliable, interpretable, and steerable AI systems.![](https://static.ghost.org/v5.0.0/images/link-icon.svg)![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/thumbnail/6d4a0d28992ade92d6fa63646fd9c9d318245c6c-2400x1260.jpg)](https://www.anthropic.com/system-cards) [OpenAI Deployment Safety Hub: System cards & other updatesReview OpenAI system cards and other safety updates for deployed AI systems. See how systems are evaluated, monitored, and improved over time.![](https://static.ghost.org/v5.0.0/images/link-icon.svg)OpenAI Deployment Safety Hub![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/thumbnail/deploymentsafety-social.png)](https://deploymentsafety.openai.com) [Model cardsOverviews of how an advanced AI model was designed and evaluated.![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/icon/google_deepmind_48dp.svg)Google DeepMind![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/thumbnail/JhHmWLAQUgqyIHM9Mt62AIUR3O3wzh-LLnRwV7dteeV-HzvF6IH6Fcv87YLa5icRUARDA_fOeuuUfPUhlssQJCfGh455T5VKQu1_1j6xhFp16qsM-w1200-h630-n-nu-rw)](https://deepmind.google/models/model-cards/) Model cards may have originated with this paper [Model Cards for Model ReportingTrained machine learning models are increasingly used to perform high-impact tasks in areas such as law enforcement, medicine, education, and employment. In order to clarify the intended use cases of machine learning models and minimize their usage in contexts for which they are not well suited, we recommend that released models be accompanied by documentation detailing their performance characteristics. In this paper, we propose a framework that we call model cards, to encourage such transparent model reporting. Model cards are short documents accompanying trained machine learning models that provide benchmarked evaluation in a variety of conditions, such as across different cultural, demographic, or phenotypic groups (e.g., race, geographic location, sex, Fitzpatrick skin type) and intersectional groups (e.g., age and race, or sex and Fitzpatrick skin type) that are relevant to the intended application domains. Model cards also disclose the context in which models are intended to be used, details of the performance evaluation procedures, and other relevant information. While we focus primarily on human-centered machine learning models in the application fields of computer vision and natural language processing, this framework can be used to document any trained machine learning model. To solidify the concept, we provide cards for two supervised models: One trained to detect smiling faces in images, and one trained to detect toxic comments in text. We propose model cards as a step towards the responsible democratization of machine learning and related AI technology, increasing transparency into how well AI technology works. We hope this work encourages those releasing trained machine learning models to accompany model releases with similar detailed evaluation numbers and other relevant documentation.![](https://static.ghost.org/v5.0.0/images/link-icon.svg)arXiv.orgMargaret Mitchell![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/thumbnail/arxiv-logo-fb.png)](https://arxiv.org/abs/1810.03993) Futurist Rob Hoeijmakers sets out his thoughts on the cards and the information (and signals) they hold: [Model Cards, System Cards and What They’re Quietly BecomingWhat are AI model cards, and why are they becoming the documents regulators will turn to first? I read a few and it taught me more than I expected.![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/icon/Circle-logo-2.png)Rob HoeijmakersRob Hoeijmakers![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/thumbnail/IMG_4983.jpeg)](https://hoeijmakers.net/model-cards-system-cards/) and on benchmarks and more deterministic tests [From Benchmarks to Evals: How We Measure AI and Why It MattersBenchmarks score models. Evals test them in real workflows. This is your guide to understanding how we measure and trust AI performance today.![](https://static.ghost.org/v5.0.0/images/link-icon.svg)Rob HoeijmakersRob Hoeijmakers![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/thumbnail/IMG_5761-1.jpeg)](https://hoeijmakers.net/ai-benchmarks-and-evals/) --- Swift note – The Economist points to these two. Maybe a short directoy of benchmarks? Or do we get just as much from testing ourselves with one? [WeirdML (v2)A benchmark of nonstandard ML engineering tasks in a variety of domains.![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/icon/favicon-0d9791a6-8be5-45ad-8918-508afaec611c.svg)Epoch AI![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/thumbnail/benchmarking-thumbnail-132778e9-dadf-4545-8be8-196fa91551e5.png)](https://epoch.ai/benchmarks/weirdml?view=graph&tab=release-date&metric=Accuracy) [Try Yourself - SimpleBench![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/icon/favicon-72beeefc-e862-46cc-9da8-7d8dbf73dd79.ico)SimpleBenchSimpleBench Challenge![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/thumbnail/hf-logo-pirate-a835a290-023e-44e7-b88a-0db2a43ddf9f.svg)](https://simple-bench.com/try-yourself) ### Workroom PlayTime 052: Exploring Play Styles URL: https://www.workroom-productions.com/workroom-playtime-052-exploring-play-styles/ Last updated: 2026-04-08T16:57:03.000Z This week's [Workroom PlayTime](https://www.workroom-productions.com/workroom-playtime/) is [Thursday 9 April at 2:30pm London time](https://this-ti.me/?uts=1775741400&tz=Europe%2FLondon&name=Workroom+PlayTime+052). That's **earlier** than last week. Sorry to keep moving it about. We'll do [Exploring Play Styles](https://www.workroom-productions.com/exploring-play-styles/), to consider how our favoured approach to play might influence out exploratory testing. We'll gather on Zoom. These exercises are for everyone, for free. [Free subscribers](https://www.workroom-productions.com/#/portal/signup) get to play together most weeks, paid subscribers get a guaranteed slot or rerun if we're full. If you get this in an email or can see the **joining info** section below, then you're a subscriber and have access to the Zoom etc. for the workshop. If you'd like to bring a friend, you can do that. I'll say yes until I've got too many. Go here to [book your spot](https://savvycal.com/workroomprds/workroom-playtime-052-exploring-play-styles) and get reminders. ↓ *signed-in subscribers will see joining info and links below* ↓ _This post is for subscribers only._ ### Exploring Play Styles URL: https://www.workroom-productions.com/exploring-play-styles/ Last updated: 2026-04-09T13:24:32.000Z *The ways we play influence the ways we test. In this exercise, we explore what that might mean.* ## Axioms for here and now... - **Exploration** is playing with intent (see *What is Exploratory Testing?* in [Alan Richardson's *Dear Evil Tester*](https://www.eviltester.com/page/deareviltester/)). - **Exploratory testing** is exploration with judgement. ## Exercise *10 mins, solo or together* How do you like to play? If you need a framework, have a look at one of the sources below. How has that tendency influenced your exploration, and by extension your exploratory testing? ## Debrief *10 minutes* Share, consider differences, identify things you want to take away. ## Sources #### Brian Sutton-Smith, Seven Rhetorics of Play - ****Progress** – growth is fun. Numbers going up. - ****Fate** – unpredictability (and one's reaction to it) is fun. Chance. - ****Power** – being better is fun. Winning. - ****Community identity** – feeling closer to others is fun. Group-making. - ****Imaginary** – experiments are fun. Safe learning. - ****Self** – discovering oneself is fun. Self knowledge. - ****Frivolity** – silliness is fun. Absurdism. More from [House of Nerdery](https://houseofnerdery.com/2025/06/04/exploring-play-sutton-smiths-seven-rhetorics/), [summary](https://sk.sagepub.com/ency/edvol/play/chpt/rhetorics-play-suttonsmith) from **Encyclopaedia of Play in Today's Society*, bio in [Wikipedia](https://en.wikipedia.org/wiki/Brian%5FSutton-Smith). #### Stuart Brown's Eight Personalities of Play - ****The Explorer** – novelty and learning is fun. - ****The Collector** – similarity, completeness and rarity is fun. - ****The Competitor** – winning, and being the winner, is fun. - ****The Creator** – making and mending is fun - ****The Director** – planning, thinking ahead, manipulating is fun - ****The Joker** – subversion is fun. - ****The Storyteller** – narrative is fun. - ****The Kinesthete** – moving is fun. See [Personalities of Play](https://nifplay.org/what-is-play/play-personalities/) #### Mildred Parten's stages of childhood play - ****Unoccupied** – seems scattered and aimless - ****Solitary** – directed, no interaction with others - ****Onlooker** – active watching, not joining in - ****Parallel** – same game, same place, not interacting - ****Associative** – more about the other players than the toys or activity - ****Cooperative** – sharing goals, assigning roles, negotiating rules See [wikipedia](https://en.wikipedia.org/wiki/Parten%27s%5Fstages%5Fof%5Fplay) #### Roger Caillois's Categories and Polarities ****Categories** - Competition (`Agon`) - Chance (`Alea`) - Mimicry (`Mimesis`) - Perception shift (`Ilinx`) ****Polarities** - Uncontrolled, improvised: rules change during game (`Paidia`) - Rewards skill, effort, patience, strategy: rules set before (`Ludus`) See [Wikipedia](https://en.wikipedia.org/wiki/Roger%5FCaillois) ### Workroom PlayTime 051: Exploring Bloom Filters URL: https://www.workroom-productions.com/workroom-playtime-051-exploring-bloom-filters/ Last updated: 2026-03-31T15:32:43.000Z This week's [Workroom PlayTime](https://www.workroom-productions.com/workroom-playtime/) is Thursday 2 April at 5pm London time. That's **later** than usual. We'll do [Exploring Bloom Filters](https://www.workroom-productions.com/exploring-bloom-filters/), to play with an algorithm, and see what that tells us about exploring, experimentation or about the algorithm. We'll gather on Zoom. These exercises are for everyone, for free. [Free subscribers](https://www.workroom-productions.com/#/portal/signup) get to play together most weeks, paid subscribers get a guaranteed slot or rerun if we're full. If you get this in an email or can see the **joining info** section below, then you're a subscriber and have access to the Zoom etc. for the workshop. If you'd like to bring a friend, you can do that. I'll say yes until I've got too many. Go here to [book your spot](https://savvycal.com/workroomprds/wp051) and get reminders. ↓ *signed-in subscribers will see joining info and links below* ↓ _This post is for subscribers only._ ### Exploring Bloom Filters URL: https://www.workroom-productions.com/exploring-bloom-filters/ Last updated: 2026-04-02T11:34:31.000Z *A data structure that tells you whether a desired item is definitely *not* in the set you're searching. Used as a first-point-of-contact in retrieving information because the dataset is small and fast compared with a direct search.* ## Exercise Go to this lovely visualisation . This particular bloom filter has 50 bits and uses three functions – every time you add a key, the filter will (deterministically) set 3 of the 50 bits. The check on the right lets you see whether a particular key is NOT in the set, by seeing whether the bits for that key are set, or not. Play with it – add keys, see the graph, check the response of the algorithm with the keys you've added and with terms you've not added. Track how your experiments change with what you see – keep an eye on what you're *doing*, and what you're *looking for*, and any *model* (hypothesis, rational, plausible guess) that you might be holding in mind. ## Debrief What did you observe? How did that change your experiment? How did you lean from your experiments? Let's reflect on this from the point of view of testing and exploring. OR What do you know now about bloom filters – how might you test them? ## Sources [Bloom filter - Wikipedia![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/icon/wikipedia-1.png)Wikimedia Foundation, Inc.Contributors to Wikimedia projects![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/thumbnail/500px-Bloom_filter.svg.png)](https://en.wikipedia.org/wiki/Bloom%5Ffilter) ### Workroom PlayTime 050: Power of Variety URL: https://www.workroom-productions.com/workroom-playtime-050-power-of-variety/ Last updated: 2026-03-03T18:03:38.000Z This week's [Workroom PlayTime](https://www.workroom-productions.com/workroom-playtime/) is [Thursday 5 March at 2:30pm London time](https://this-ti.me/?uts=1772721000&tz=Europe%2FLondon&name=Workroom+PlayTIme+050). That's 2:**30** – half past the hour. *If you've seen the email, you'll know that I mentioned 3 March in the subject. Ooops.* Throughout March, we'll run exercises to work through some ideas around how much we might spend testing, when we might stop, and how we might set up our test team and their skills. It's all about discovery, test management and handling a gamble, and it uses a toy simulation to get the ideas moving. This first one is [The Power of Variety](https://www.workroom-productions.com/power-of-variety/) **An experiment**: use this Savvycal link to [put a note in your calendar](https://savvycal.com/workroomprds/workroom-playtime-050-the-power-of-variety) and get an email beforehand. You don't need to do this to come, but I'd love to find out how it works for you. Be aware – I *think* it will show your name and email to other attendees. If that's a bother, don't use it – I'm trying to see if I can turn off the feature, but haven't yet found a way. We'll gather on Zoom. These exercises are for everyone, for free. [Subscribers](https://www.workroom-productions.com/#/portal/signup) get to play together every week. If you can see a **joining info** section below, then that gives you access to the Zoom etc. for the workshop. If you'd like to bring a friend, you can do that. I'll say yes until I've got too many. I'm still gathering audience for the [*Interesting Testing*](https://www.workroom-productions.com/interesting-testing/) Longplayer, I don't have enough people to run it at the end of this week. So I'm re-scheduling to later in March – [pick a day / time that works for you](https://savvycal.com/p/workroomprds/interesting-testing-longplayer-march-2026). ↓ *signed-in subscribers will see joining info and links below* ↓ _This post is for subscribers only._ ### Workroom PlayTime 049: What Makes a Great Talk? URL: https://www.workroom-productions.com/workroom-playtime-049-what-makes-a-great-talk/ Last updated: 2026-02-24T15:49:18.000Z This week's [Workroom PlayTime](https://www.workroom-productions.com/workroom-playtime/) is [Thursday 26 February at 2:30pm London time](https://this-ti.me/?uts=1772116200&tz=Europe%2FLondon&name=Workroom+PlayTime+049). That's 2:***30*** – half past the hour. We'll do [What Makes a Great Talk?](https://www.workroom-productions.com/what-makes-a-great-talk/), to reflect on talks we've seen from the audience which have positive lessons for us as speakers. That's the third (and last) part of February's [SpeakerPrep](https://www.workroom-productions.com/tag/speakerprep/) collection ([video](https://www.workroom-productions.com/speaker-prep-sessions-for-workroom-playtime/)). We'll gather on Zoom. These exercises are for everyone, for free. [Subscribers](https://www.workroom-productions.com/#/portal/signup) get to play together every week. If you can see a **joining info** section below, then that gives you access to the Zoom etc. for the workshop. If you'd like to bring a friend, you can do that. I'll say yes until I've got too many. I'm still gathering audience for the [*Interesting Testing*](https://www.workroom-productions.com/interesting-testing/) Longplayer, I don't have enough people to run it at the end of this week. So I'm re-scheduling to later in March – [pick a day / time that works for you](https://savvycal.com/p/workroomprds/interesting-testing-longplayer-march-2026). ↓ *signed-in subscribers will see joining info and links below* ↓ _This post is for subscribers only._ ### What Makes a Great Talk? URL: https://www.workroom-productions.com/what-makes-a-great-talk/ Last updated: 2026-02-24T15:17:04.000Z *focussing on the positives* ## Exercise Identify a session\* you loved. Share it. Take 5 minutes to recall the session – try to recall the talk itself, not just its message or location or your response. Write down how it made its impact, and why that is important to you. Share your thoughts (10 minutes + ) ## Debrief Any principles? \`\* `session` – no fixed format. Best if it's testing related. Avoid sessions by people in the room, please. --- ## From SpeakerPrep *An in-person session, where we have materials to share. I'll update with materials when I find them...* Use this exercise to share your principles, to share your inspirations, or to analyse some fine talks. ## Inspirations Whose session has inspired you? How did it inspire you? Share your example, and we’ll add it to our list. ## Principles What, for you, makes a great talk? Consider your principles by finding specific examples of actions *you* take. Consider your principles by remembering specific examples of things someone did while speaking that inspired you. Put down the principles, and add them to our gallery. ## Analyse a talk Watch a great talk (we have a list). What are you watching for? How does it manifest? How can you help yourself, and others, to be brilliant in a similar way? Add your principles to the gallery. ### Workroom PlayTime this week URL: https://www.workroom-productions.com/workroom-playtime-20260219/ Last updated: 2026-02-16T10:55:22.000Z There won't be a Workroom PlayTime this week – it's half term and I'll be playing with the family in [Kew](https://www.kew.org/kew-gardens) on Thursday. Next week we'll do "What makes a Great Talk" as the last in the [SpeakerPrep](https://www.workroom-productions.com/speaker-prep-sessions-for-workroom-playtime/) sessions. We will re-run January's [*Interesting Testing*](https://www.workroom-productions.com/interesting-testing/) series (tiny [video](https://vimeo.com/workroomprds/wplp01-interesting-testing), [more info](https://www.workroom-productions.com/workroom-playtime-longplayer-interesting-testing/)) in a longer (90-minute) session towards the end of this month. I'll run if I have 6 people, run more than once if I have too many, and if I have several who need a particular time or day, I'll run it for you and friends. I have a selection of times: to see them and indicate which might work for you, use [my savvycal link](https://savvycal.com/p/workroomprds/interesting-testing-longplayer-feb-26) (or reply to this if you prefer). ### WP 048: Cutting to the Core URL: https://www.workroom-productions.com/workroom-playtime-048-cutting-to-the-core/ Last updated: 2026-02-09T13:06:49.000Z This week's [Workroom PlayTime](https://www.workroom-productions.com/workroom-playtime/) is [Thursday 12 February at 2:30pm London time](https://this-ti.me/?uts=1770906600&tz=Europe%2FLondon&name=Workroom+PlayTime+048). That's 2:***30*** – half past the hour. In the second part of February's [SpeakerPrep](https://www.workroom-productions.com/tag/speakerprep/) collection ([video](https://www.workroom-productions.com/speaker-prep-sessions-for-workroom-playtime/)), we'll try [Cutting to the Core](https://www.workroom-productions.com/cutting-to-the-core/) to get our talks down to a single tiny idea. We'll gather on Zoom. These exercises are for everyone, for free. [Subscribers](https://www.workroom-productions.com/#/portal/signup) get to play together every week. If you can see a **joining info** section below, then that gives you access to the Zoom etc. for the workshop. If you'd like to bring a friend, you can do that. I'll say yes until I've got too many. I'm still gathering audience for the [*Interesting Testing*](https://www.workroom-productions.com/interesting-testing/) Longplayer. I'll run with 6, and I'll run a a time and day that suits us. In the Longplayer, we'll have more context, more conversation, and we'll rerun the exercises. Here is a tiny [video](https://vimeo.com/workroomprds/wplp01-interesting-testing), [more info](https://www.workroom-productions.com/workroom-playtime-longplayer-interesting-testing/). **Email me to give me a sign that you're interested**, and we'll find the right time for you. ↓ *signed-in subscribers will see joining info and links below* ↓ _This post is for subscribers only._ ### Cutting to the Core URL: https://www.workroom-productions.com/cutting-to-the-core/ Last updated: 2026-02-12T14:27:04.000Z *20-minute Workroom PlayTime* What's your fundamental idea for your talk? The one thing you want an audience to hear, the post-it version, the single short sentence you use in a bar chat? In this exercise, you'll work towards it – and it may be surprisingly different from where you started, and different again from your title. You'll need to start with a worked-on talk, or at least an abstract. This is a reduction, not an expansion. ## Exercise 1 – squish *5 mins* Squish your talk, {from opening to conclusion | the big points | the problem and its solution}, into 30 words or so. A few short sentences. You'll need to cut stuff out – and that's the point. Don't bother to cut conjunctions and pronouns: cut whole sections. Cut anything inferred or supporting. Amalgamate. Kill the jokes and teases and lists. If possible, keep **nothing** of your original – *rewrite*. The time for this is intentionally short. Don't polish. Don't share. It's fine if you hate what you've written. ## Exercise 2 – sloganise *5 minutes* You just squished 300 words into 30\. Now write 3\. Take your time to write at least one slogan summarising your whole talk. You'll probably need more than one word. You want to keep it under six. Go for three. Or ignore this, and write a rhyme. Or draw a picture. You're aiming for instant comprehension – that's all that matters. ## Exercise 3 – announce and polish *5-10 minutes* Use your slogan in a sentence with the group. You'll probably immediately think of a change. Share that, too. Listen to everyone's. What excites you? Share feedback if it is asked for. Offer suggestions, if that seems right. If there's time, share your thoughts on getting to this point. --- ## Ideas - Will you primarily inform, inspire, entertain, persuade, give an experience – or something else? - Consider all the ways you're putting that message across. Consider when in your talks you're setting out that message, what support for each occasion. - Consider your audience – who will use your message, and what emotional reaction do you hope they might have. - Consider different media – statement, picture, story, slogan, joke, example, counter-example, logical argument, sincere appeal, warning, vision. - Polish with rhyme, rhythm, alliteration, metaphor ### Workroom Playtime 047: Getting to a Great Abstract URL: https://www.workroom-productions.com/workroom-playtime-047-getting-to-a-great-abstract/ Last updated: 2026-02-02T17:38:07.000Z This week's [Workroom PlayTime](https://www.workroom-productions.com/workroom-playtime/) is [Thursday 5 February at 2:30pm London time](https://this-ti.me/?uts=1770301800&tz=Europe%2FLondon&name=Workroom+PlayTime+047). That's 2:***30*** – half past the hour. Starting February's [SpeakerPrep](https://www.workroom-productions.com/tag/speakerprep/) collection, we'll share and work on our conference proposals in [Getting to a Great Abstract](https://www.workroom-productions.com/getting-to-a-great-abstract-2/). We'll gather on Zoom. These exercises are for everyone, for free. [Subscribers](https://www.workroom-productions.com/#/portal/signup) get to play together every week. If you can see a **joining info** section below, then that gives you access to the Zoom etc. for the workshop. If you'd like to bring a friend, you can do that. I'll say yes until I've got too many. January's [*Interesting Testing*](https://www.workroom-productions.com/interesting-testing/) series was fun. I'll rerun the exercises (tiny [video](https://vimeo.com/workroomprds/wplp01-interesting-testing), [more info](https://www.workroom-productions.com/workroom-playtime-longplayer-interesting-testing/)) and will put them in context in a longer (90-minute) session sometime in February – probably a Weds / Thursday evening in the last couple of weeks of the month. I'll run if I have 6 people, run more than once if I have too many, and if I have several who need a particular time or day, I'll run it for you and friends. Reply to this to sign up. This is the first in a [SpeakerPrep](https://www.workroom-productions.com/tag/speakerprep/) collection, running in Workroom PlayTime sessions throughout February – here's a [page with a video](https://www.workroom-productions.com/speaker-prep-sessions-for-workroom-playtime/), and I'll add more information as we go. ↓ *signed-in subscribers will see joining info and links below* ↓ _This post is for subscribers only._ ### Speaker Prep Sessions for Workroom Playtime URL: https://www.workroom-productions.com/speaker-prep-sessions-for-workroom-playtime/ Last updated: 2026-02-02T17:16:37.000Z Bart Knaack and I developed a series of Speaker Prep activities to help people get their ideas into conferences and to communicate with audiences. In February 2026, I'll run some of those exercises online for Workroom PlayTime. ### Workroom PlayTime Longplayer: Interesting Testing URL: https://www.workroom-productions.com/workroom-playtime-longplayer-interesting-testing/ Last updated: 2026-02-16T10:36:00.000Z January's [*Interesting Testing*](https://www.workroom-productions.com/interesting-testing/) series was fun. I'll rerun the exercises and will put them in context in a longer (90-minute) session sometime in February – probably a Weds / Thursday evening in the last couple of weeks of the month. I'm calling it a ***Longplayer***. I'll run if I have 6 people, run more than once if I have too many, and if there are several of you who need a particular time or day, I'll run it for you and friends. Want to come? Email me – and if you're not already subscribed, subscribe. ## What we're doing I'll take the last three Workroom PlayTimes and run them all together. They've all, as it happens, got a machine reasoning / AI / LLM component to them – and they all present challenges or provocations to us, as testers. You'll get hands-on with these different parts: - We'll do two things with code, building that code from examples. We'll build `whatwords`, from [A Library with No Code](https://www.workroom-productions.com/a-library-with-no-code/), as a one-shot, just from examples. - Next, we'll use that library in a [Ralph Loop](https://www.workroom-productions.com/exercise-the-ralph-loop/): we'll set the loop up so that we're building tiny checkable deliverables using a fresh LLM each time. Each tiny task will be built test (check) first, each with a fresh context, with us simply checking and steering, not designing or coding. - We'll also be looking at the [ARC-AGI website and its tests](https://www.workroom-productions.com/playing-with-arc-agi-tests/): tests that are built to be easy for humans, hard for machines, and designed to assess and maybe guide the growth of machine reasoning. _This post is for subscribers only._ ### Getting to a Great Abstract 2 URL: https://www.workroom-productions.com/getting-to-a-great-abstract-2/ Last updated: 2026-02-05T11:58:48.000Z *20-minute Workroom PlayTime version. Original* [*here*](https://www.workroom-productions.com/getting-to-a-great-abstract/)*.* Your '*abstract*' summarises your content for someone else to judge. Here, we're dealing with what you write when you send your proposal to the program committee. #### who we're ignoring... Your abstract, no matter how polished, will get rewritten for different audiences. Conference advertising may need differnt text from the on-the-day guide. And you'll have your own brief verbal description, too for that moment where the small talk turns to "You're a speaker? What's your topic?" Who do you summarise for? - the people who decide what goes on the program - people who decide to engage with your content. - People who might decide to go to the event, based on what you and others have written. - Optimisation algorithms for search and summary and discovery - People who you bump into and who need a short verbal summary. - People who want to know what you've presented before. In this exercise, we'll start with real things we're working on, consider what works, make changes, get and give feedback. *Got something to share?* Put it on the Miro board. It does not have to be a polished abstract – and can be just an idea for a topic or a format if that's what you're working on. *Seeking feedback?* Put a 👍 if you want to know what works. Put a 👎 if you want to hear what doesn't. Put both if you fancy. ## Exercise: Quietly Consider Others *5 mins* Read another abstract. Decide what works. Decide what doesn't. Three of each, if you need a goal. Be very specific and write an alternative. If you like, formulate a principle (if you want to build on something, glance at Great / Poor abstract). You won't share these unless you want to. ## Exercise: Publicly Change Yours *5 mins* Apply (some of) that same judgement to *yours* – if you liked their catchy title, try working on your title. If you found it hard to understand a sentence, identify something of yours that needs work, and clarify. Make at least one change. Put your new version on the board. ## Debrief: Reflect *10 mins* We'll go round, presenting the changes we made. If you've asked for feedback, and if someone has feedback, we'll share. Celebrate any changes. Notice where your opinions match, and differ, on what makes a 'great' abstract. This is entirely good – and it enhances your unique angle, and positions you for the right audience. --- ## Example abstracts Pick one! *A Thousand Tiny Servers*, *Training a Tool to Test* and *How to be an Awesome Tester* are examples of potential, but un-developed talks, and date from 2018\. *Abstract Abstract* is a comedy example with plenty of things wrong, and is from 2019\. *Building and Handling Exploratory Interfaces* was written as a serous, but speculative, proposal in 2025\. I'll share its (rejection) feedback on request. [Anne-Marie Charrett](https://www.annemariecharrett.com) wrote *How to be an Awesome Tester* for the collection. [Fantasy Abstract – A Thousand Tiny ServersFantasy Abstract – A Thousand Tiny Servers.pdf16 KBdownload-circle](https://www.workroom-productions.com/content/files/2026/02/Fantasy-Abstract-----A-Thousand-Tiny-Servers.pdf "Download") [Fantasy Abstract – How to be an awesome testerFantasy Abstract – How to be an awesome tester.pdf17 KBdownload-circle](https://www.workroom-productions.com/content/files/2026/02/Fantasy-Abstract-----How-to-be-an-awesome-tester.pdf "Download") [Fantasy Abstract – Training a Tool to TestFantasy Abstract – Training a Tool to Test.pdf15 KBdownload-circle](https://www.workroom-productions.com/content/files/2026/02/Fantasy-Abstract-----Training-a-Tool-to-Test.pdf "Download") [Fantasy Abstract – Abstract AbstractFantasy Abstract – Abstract Abstract.pdf27 KBdownload-circle](https://www.workroom-productions.com/content/files/2026/02/Fantasy-Abstract-----Abstract-Abstract.pdf "Download") [Abstract – Building and Handling Exploratory InterfacesAbstract – Building and Handling Exploratory Interfaces.pdf50 KBdownload-circle](https://www.workroom-productions.com/content/files/2026/02/Abstract-----Building-and-Handling-Exploratory-Interfaces.pdf "Download") --- ## Support Captured at various events – thank you to participants! ### Great Abstract - catchy title - idea that has potential - indicates that thought and work has gone into communicating the message - /brief/ problem statement - real examples - demonstrates experience - clear takeaways - relevant to event - easy to read, draws you onwards - non-trivial - has substance – more than a tease, less than a paper - more than just the bright side – pitfalls, failures, antipatterns, pathologies - within word limit (and more than a couple of sentences) - offers interaction or demonstration --- ### Poor Abstract - claims expertise / authority - problem excludes content - poorly written, poorly proofread - unstructured or incoherent - buzzwords - single tool - happy paths only - One true method / “I’m right” - Emotionless - Tired topic ### Workroom Playtime 046: The Ralph Loop URL: https://www.workroom-productions.com/workroom-playtime-046-the-ralph-loop/ Last updated: 2026-01-23T12:33:17.000Z This week's [Workroom PlayTime](https://www.workroom-productions.com/workroom-playtime/) is [Thursday 29 January at 2:30pm London time](https://this-ti.me/?uts=1769697000&tz=Europe%2FLondon&name=Workroom+PlayTime+046). That's 2:***30*** – half past the hour. Continuing January's [*Interesting Testing*](https://www.workroom-productions.com/interesting-testing/) theme, we'll go play with a Ralph Loop – a stupid but persistent way to iterate with an LLM towards working, tested code. We'll use [Exercise: The Ralph Loop](https://www.workroom-productions.com/exercise-the-ralph-loop/). That exercise is currently very under-formed... We'll gather on Zoom. These exercises are for everyone, for free. [Subscribers](https://www.workroom-productions.com/#/portal/signup) get to play together every week. If you can see a **joining info** section below, then that gives you access to the Zoom etc. for the workshop. If you'd like to bring a friend, you can do that. I'll say yes until I've got too many. `046` will bring this short series on [*Interesting Testing*](https://www.workroom-productions.com/interesting-testing/) to а close. I'll rerun the exercises and to put them in context in a longer (90-minute) session sometime in February. `047` will be the first in a series from [SpeakerPrep](https://www.workroom-productions.com/tag/speakerprep/), focussing on ideas and abstracts for conference sessions in the regular Workroom PlayTime sessions. ↓ *signed-in subscribers will see joining info and links below* ↓ _This post is for subscribers only._ ### Exercise: The Ralph Loop URL: https://www.workroom-productions.com/exercise-the-ralph-loop/ Last updated: 2026-01-28T23:58:32.000Z *Workroom PlayTime – you'll probably need to run on your own machine to get hands-on.* What's a Ralph Loop? A Ralph Loop is an approach to generating code. #### more... By putting a simple loop around code generation, a 'fresh' LLM is always used to generate code. A 'fresh' LLM has all-relevant context, un-polluted by the LLM's answers or by irrelevant earlier tasks. This **forces* design and spec work out of the generation context, **forces* small and checkable tasks, **forces* documentation that allows a new LLM to pick up where the old one died. By **forces*, I mean that the approach needs that pre- and post- work to happen in order to be useful. How is the relevant to testers – or at least to a testing mindset? With design and spec work under human control and out of generation steps, testers can influence as valuably as ever. With checkable tasks, testers describe the checks (and audit the ways those checks are made). With the weird things LLMs do, testers can develop their whiskers for the weird and get right into the loop, stopping to reframe and re-aim the work. In this exercise, we'll use a one-liner to iteratively generate a pure function. We'll use [my own version of ](https://github.com/workroomprds/whatwords%5Fralphed%5Fwp046): [A Library with no Code](https://www.workroom-productions.com/a-library-with-no-code/). My version has been made into a python library. I've asked Claude to spec out some approaches to adding daylight saving time, and run a Ralph Loop through several iterations – each time working through one of the small checkable tasks. I've got checkpoints, so I can 'rewind' my environment, and re-run the loop. The repo has commits, so you should be able to do that, too. ## Dragging a repo into a Ralph Loop To get to a working exercise, I started with something that I knew could generate working code – Drew Breunig's examples and instructions to generate a relative-time library. I asked Claude Code to generate tests from the examples, and code to pass those tests. Then I asked Claude to come up with a a spec to add daylight saving time (which implies knowing what and where those times are), to make a plan of action to make useful examples into tests and to write code to pass those tests, and to update the docs. Once I had fair docs, I set a Ralph Loop on the problem of making the code. I've taken this repo through a the following [(here is the commit history)](https://github.com/workroomprds/whatwords%5Fralphed%5Fwp046/commits/main/). Note: I failed to fork it properly, so it's no longer linked to the original. - a conversation with an LLM that took the code-free library to a[ python implementation](https://github.com/workroomprds/whatwords%5Fralphed%5Fwp046/commit/24453e4aca970e2dba7cca0e7aec5a3badc9ed0a) in `/bin` with passing tests derived from existing examples (sadly: *derived from*. I've got a js one that uses the examples directly as tests) - a conversation with an LLM that led to a [set of docs](https://github.com/workroomprds/whatwords%5Fralphed%5Fwp046/commit/b9d206f6b30ef7b501939a7ce73c16055eb49cfb) describing a change - several further conversations, each an individual turn around a Ralph Loop, each typically producing two commits – one for tests / test-passing code, and one for the documentation, updating `decisions.md`, [implementation\_plan.md](https://github.com/workroomprds/whatwords%5Fralphed%5Fwp046/blob/main/implementation%5Fplan.md) and [progress.md](https://github.com/workroomprds/whatwords%5Fralphed%5Fwp046/blob/main/progress.md) . - a final commit where I noticed that the [repo was missing the vital prompt.md ](https://github.com/workroomprds/whatwords%5Fralphed%5Fwp046/commit/18dbb50cd27c16b07b733ad562fd58e75e43a27d), and included `INSTALL.md` and `handover.md` which I had binned after a couple of runs round the loop. To get properly hands-on, you'll need a safely-sandboxed environment with a tool-using agent (and its tools) installed, a key with enough credit, and the repo. To get to this point, I have used [sprites.dev](https://sprites.dev) for the sandboxed environment, I've installed Claude Code and Pytest on that environment, and I've connected the environment to my github account and to my Anthropic account. I can't duplicate that environment and I can't easily share commandline access to the one I have. **So for Workroom PlayTime 046**, I don't think I can guarantee to set that up for you in a worthwhile way. You can do it yourself, or I'll demo it – we'll watch and talk. If you want to do it yourself, then you'll need to set up something to the environment described above (sprites.dev is very handy for this and has Claude installed and talks with various services and tools), clone the repo, and use the commandline below to pass the prompt into your agent (`claude` in the prompt – meaning [Claude Code](https://claude.com/product/claude-code)) over and over again. You'll also need to rewind the commits a bit, but to keep `prompt.md` and probably bin `handover.md` and `INSTALL.md`. ### Costs The Ralph Loop seems to be a thought experiment that happens to work as an software generation technique, and is *not* an efficient use of tokens (as a proxy for real money, energy or resources). Money: In my experiment, each iteration (i.e. each small checkable task) used $1-$2 and took 5-10 minutes – so it took the best part of an hour and $10 to add daylight savings to the algorithm. By comparison, my 'magic loop' was cheaper, useing deterministic tools to run tests, check syntax, and do commits, and \~80 people playing used under $20 in an afternoon. Claude Code can 'one-shot' the base library above (and its tests) for around $3, and Amp can build the [Testing Transparently](https://www.workroom-productions.com/testing-transparently/) web page app for around $4. On the other hand, the approach gives an LLM a good shot at greatness, by giving it great context. If you've got a competent local LLM or unlimited tokens, then maybe the cost is less of a problem. ## Exercise You've got the repo and you're all set up with an agent? Smashing. If not, I'll share my screen – and you've still got the files to look at. - Make a branch - Rewind a few commits - Run the one-line bash command below. - Watch your CLI agent do its thing while talking - Review while talking – see ***Debrief*** below ### Ralph Loop ```shell while true:; do cat prompt.md | claude --dangerously-skip-permissions; done ``` Yeah – it is *that* simple/stupid. And that's because the Ralph Loop is *just a loop* – the infrastructure (local and specific / vast and generic) is elsewhere. The key things this enforces / expects / needs are: - fresh new context every new task - lots of documentation (specs, examples, plans, glossaries) to give context to the prompt below - a test harness (and runnable tests) - working code that needs a small change - a tool-enabled LLM that can write and run checks, write code, notice / diagnose / fix problems, work with git, parse the docs, prioritise work and set up new work. - a human watching, judging and steering where necessary In terms of what's actually going on... - You set an initial list of small checkable tasks, each with some completion criteria. You might use an LLM for that. This setting of checkable tasks is absolutely necessary, and is where our skills as testers get used. - The Ralph Loop effectively queues up fresh LLMs. - It starts a new LLM-on-the-commandline instance, passing the prompt. - That *fresh* LLM fills its own context from the docs and picks its own task from a list in a doc. It goes to work, and you watch it, use a second terminal window to check its work, fire up another task, or go get a cup of tea. - If the LLM goes off the rails (again, tester skills), *you kill it*. Before you let the Ralph Loop restart, you reframe the tasks and the instructions (tester skills) – and that is what shapes the final product. - When the LLM completes a task, it tidies up for the next LLM, and waits for you to approve. - You can fiddle with the docs etc. here. If you need to fiddle with the prompt, you'll need to kill the loop. Core thought: If the task is done, *you kill it* and the Ralph Loop ushers in the next LLM. #### contents of my `prompt.md` ```md study handover.md study implementation_plan.md pick the one most important thing to do that has not yet been done. That is your target. Concentrate only on the target. IMPORTANT * tests come first – add examples to tests.yaml for your target and run the tests. Expect the tests for your new examples to fail – your target does not yet exist in the code * change existing code to pass the tests * ALWAYS run the tests after changing the code * when the tests pass, commit * after committing, append a short narrative to progress.md, adjust and add to decisions.md, and change implementation_plan.md to reflect your completed target, and any new or unfnished work. We're not going to use handover.md again. ``` Note – I had imagined, based on previous runs, that the tests would ingest the examples, and dumbly based this prompt on that assumption. But in this repo they don't. I'll ask for it explicitly next time. ## Debrief *pick one* - We'll compare and contrast our building experiences - We'll look at the artefacts it makes, and see what they tell us - We'll consider what we can give the loop to help it – and how much tests are part of that - We'll consider the usefulness of a system called into being without tests, and yet which makes up its tests as it goes. ## Sources Here's Geoffrey Huntley's [earlier post](https://ghuntley.com/ralph/), [later post](https://ghuntley.com/loop/), and recent ['first principles' demo/video](https://www.youtube.com/watch?v=4Nna09dG%5Fc0). Here's a more measured [video](https://www.youtube.com/watch?v=%5FIK18goX4X8) from Matt Pocock. Here's a couple of fair descriptions: [Autonomous Loops](https://paddo.dev/blog/ralph-wiggum-autonomous-loops/) (more to do with the Claude plugin) and [2026 – the Year of the Ralph Loop](https://dev.to/alexandergekov/2026-the-year-of-the-ralph-loop-agent-1gkj) (around the cursor plugin). Neither are great as they focus on specific implementations rather than the approach as a thought-provoking technique. --- That [Ralph plugin](https://github.com/anthropics/claude-plugins-official/tree/main/plugins/ralph-loop)? Haven't used it, and staying closer to the ~~metal~~ commandline gives me enought insights. [The originator says](https://ghuntley.com/ralph/) the plugin 'isn't it', and as the plugin seems to treat context fundamentally differently I'm inclined to agree. It won't be part of this workshop. ### Workroom Playtime 045: ARC Foundation tests URL: https://www.workroom-productions.com/workroom-playtime-045-arc-foundation-tests-2/ Last updated: 2026-01-20T12:24:08.000Z This week's [Workroom PlayTime](https://www.workroom-productions.com/workroom-playtime/) is [Thursday 22 January at 2:30pm London time](https://this-ti.me/?uts=1769092200&tz=Europe%2FLondon&name=Workroom+PlayTime+045). That's 2:***30*** – half past the hour. Continuing January's [*Interesting Testing*](https://www.workroom-productions.com/interesting-testing/) theme, we'll go [play with the tests](https://www.workroom-productions.com/playing-with-arc-agi-tests/) from the [ARC Prize Foundation](https://arcprize.org)'s collection which are designed to check human and AI reasoning. There are a bunch of insights that we might get from the collection, from its results and from our own responses. To get moving with that, we'll do some of the tests ourselves, together. We'll gather on Zoom. These exercises are for everyone, for free. [Subscribers](https://www.workroom-productions.com/#/portal/signup) get to play together every week. If you can see a **joining info** section below, then that gives you access to the Zoom etc. for the workshop. If you'd like to bring a friend, you can do that. I'll say yes until I've got too many. Next week, we'll take a look at the Ralph Loop, which will bring this short series on [*Interesting Testing*](https://www.workroom-productions.com/interesting-testing/) to а close. I'll rerun the exercises and to put them in context in a longer (90-minute) session sometime in February. We'll mostly run exercises from [SpeakerPrep](https://www.workroom-productions.com/tag/speakerprep/) in February, focussing on ideas and abstracts for conference sessions in the regular Workroom PlayTime sessions. ↓ *signed-in subscribers will see joining info and links below* ↓ _This post is for subscribers only._ ### Playing with ARC-AGI tests URL: https://www.workroom-productions.com/playing-with-arc-agi-tests/ Last updated: 2026-01-22T14:27:28.000Z *Workroom PlayTime045* ## Exercise *15 minutes playing + 5 mins conversation* ### Kickoff Demo – let’s pick an AGI-1 task from [Explore All Tasks](https://arcprize.org/tasks/) and do it together. ### Do one yourself Try (AGI-1 task) [ARC-AGI Task #03560426](https://arcprize.org/tasks/03560426/) alone and keep a note of your hypotheses. ### Work together Try (AGI-1 / 2 task) [ARC-AGI Task #12422b43](https://arcprize.org/tasks/12422b43/) together – listen for hypotheses Then let’s try to solve (AGI-1 / 2 task) [ARC-AGI Task #136b0064](https://arcprize.org/tasks/136b0064/) together. *If we've got time, we'll go back to solving solo / in pairs.* ### Debrief *5 mins – though we will be talking about most of this as we go. The following topics are suggestions if we feel quiet.* - How did the tests evaluate reasoning? - Compare several tests. What is different in how they evaluate? - Does this 'feel' like a challenge to reasoning, or simply a sweet spot for tests? - Looking at the [leaderboard](https://arcprize.org/leaderboard), AGI-2 is a greater challenge to current models than AGI-1, and was developed as AGI-1's challenges were surmounted. Can this reflect a change in reasoning capability? ## Extend Only humans managed [ARC-AGI Task #25094a63](https://arcprize.org/tasks/25094a63/) and [ARC-AGI Task #212895b5](https://arcprize.org/tasks/212895b5/). Why might that be? JL note: JL's machines have a further extension to 90/120 mins. ### Workroom PlayTime in 2026 URL: https://www.workroom-productions.com/workroom-playtime-in-2026/ Last updated: 2026-01-13T15:26:12.000Z *Mostly from the message sent announcing* [*Workroom PlayTime 044*](https://www.workroom-productions.com/workroom-playtime-044-2/) ## Plans for 2026 I'll continue to run short, (roughly) weekly, online interactive workshops throughout 2026\. I've learnt a lot doing it in 2025, had fun, met new people and enjoyed their company. Wahey! Here's an idea of what's planned... - **January** – [*Interesting Testing*](https://www.workroom-productions.com/interesting-testing/): other people's projects that challenge my perspective, and may challenge yours. - **February** – [*SpeakerPrep*](https://github.com/workroomprds/SpeakerPrepDay), to work on ideas and abstracts - **March** – [*Testing Decisions*](https://www.workroom-productions.com/tag/testing-decisions/), looking at costs and returns. Maybe *Puzzles...* - **April** – [*Unix* *tools for testers*](https://www.workroom-productions.com/unix-tools-for-testers/), returning to 2025's emergent theme - **May** – *Agents for Testers*, because that's what Bart and I are learning about... You'll see that I plan to run themed sets in 2026, roughly one set a month. A couple of weeks after the end of a set, I'll run a long† session to re-cover some of the ground – adding context and materials around the exercises. From now‡*,* all short Workroom PlayTime sessions will be open to all subscribers. I may restrict the longer sessions to paying subscribers, or set a fee for non-subscribers. I'll try to run most Thursdays, but times are likely to be *more* irregular. This may be good – outside your working hours may be better for some of you, others may be able to come without setting an alarm or staying up late. It's *bad* for marketing – and a pain in the diary for all. I will continue to experiment with time and format this year; running the same session a couple of times for different geographical groups, possibly offering an asynchronous path, dipping into video and audio. (†) By 'short', I mean 20 minutes or so. By 'long', I mean 90 minutes or so. (‡) I write 'from now', but it's been that way since... um... May? ### Workroom PlayTime 044 (and onwards into 2026) URL: https://www.workroom-productions.com/workroom-playtime-044-2/ Last updated: 2026-01-13T15:25:58.000Z Workroom PlayTime 44 is on [Thursday 15 January at 2pm London time](https://this-ti.me/?uts=1768485600&tz=Europe%2FLondon&name=Workroom+PlayTime+44). We'll meet on Zoom, and we'll look at a [library without code](https://www.workroom-productions.com/a-library-with-no-code/). ## Plans for 2026 I'll continue to run short, (roughly) weekly, online interactive workshops throughout 2026\. I've learnt a lot doing it in 2025, had fun, and enjoyed your company. Wahey! And **thank you**! Here's an idea of what's planned... - **January** – [*Interesting Testing*](https://www.workroom-productions.com/interesting-testing/): other people's projects that challenge my perspective, and may challenge yours. - **February** – [*SpeakerPrep*](https://github.com/workroomprds/SpeakerPrepDay), to work on ideas and abstracts - **March** – [*Testing Decisions*](https://www.workroom-productions.com/tag/testing-decisions/), looking at costs and returns. Maybe *Puzzles...* - **April** – [*Unix* *tools for testers*](https://www.workroom-productions.com/unix-tools-for-testers/), returning to 2025's emergent theme - **May** – *Agents for Testers*, because that's what Bart and I are learning about... You'll see that I plan to run themed sets in 2026, roughly one set a month. A couple of weeks after the end of a set, I'll run a long† session to re-cover some of the ground – adding context and materials around the exercises. From now‡*,* all short Workroom PlayTime sessions will be open to all subscribers. I may restrict the longer sessions to paying subscribers, or set a fee for non-subscribers. I'll try to run most Thursdays, but times are likely to be *more* irregular. This may be good – outside your working hours may be better for some of you, others may be able to come without setting an alarm or staying up late. It's *bad* for marketing – and a pain in the diary for all. As a weird coping mechanism, I may run the same session a couple of times for different geographical groups, or I might offer an asynchronous path: I welcome any suggestions and will experiment more with time and format this year. One thing that was clear in 2025 was that Workroom PlayTime is better with friends. **Tell contacts, bring colleagues**: your word-of-mouth will help build a gang that enjoys playing and learning and testing together. As a bonus, all these exercises are in the world now, for anyone, forever. Maybe you'll find something that is handy to play with in your own gang. I need *your* nudges to keep making them – I'm deeply grateful for your support. (†) By 'short', I mean 20 minutes or so. By 'long', I mean 90 minutes or so. (‡) I write 'from now', but it's been that way since... um... May? ↓ *signed-in subscribers will see joining info and links below* ↓ _This post is for subscribers only._ ### A Library with no Code URL: https://www.workroom-productions.com/a-library-with-no-code/ Last updated: 2026-01-15T14:01:31.000Z *In the* [*Interesting Testing*](https://www.workroom-productions.com/interesting-testing/) *series, and Workroom* [*PlayTime*](https://www.workroom-productions.com/workroom-playtime/) *044* [Drew Breunig](https://www.linkedin.com/in/drewbreunig) proposes a software library expressed as tests, and a prompt. ## Exercise We'll read the short article – read it now if you've come here before the workshop. Then, we'll talk about it, using the library and the examples to answer questions and think more concretely. ## Sources ### Article [A Software Library with No CodeDo we still need libraries of 3rd party code when AI agents are this good?![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/icon/apple-touch-icon-2.png)Drew BreunigDrew Breunig![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/thumbnail/ikea_instructions-1.jpg)](https://www.dbreunig.com/2026/01/08/a-software-library-with-no-code.html) ### Library [GitHub - dbreunig/whenwords: A relative time formatting library, with no code.A relative time formatting library, with no code. Contribute to dbreunig/whenwords development by creating an account on GitHub.![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/icon/pinned-octocat-093da3e6fa40-16.svg)GitHubdbreunig![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/thumbnail/whenwords)](https://github.com/dbreunig/whenwords) ### Examples [GitHub - dbreunig/whenwords-examplesContribute to dbreunig/whenwords-examples development by creating an account on GitHub.![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/icon/pinned-octocat-093da3e6fa40-17.svg)GitHubdbreunig![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/thumbnail/whenwords-examples)](https://github.com/dbreunig/whenwords-examples) ## Thoughts Similar to testing for [pure functions](https://en.wikipedia.org/wiki/Pure%5Ffunction) UTC doesn't have daylight savings – if we needed daylight savings, we'd need a location and a lookup. [I ported JustHTML from Python to JavaScript with Codex CLI and GPT-5.2 in 4.5 hoursI wrote about JustHTML yesterday—Emil Stenström’s project to build a new standards compliant HTML5 parser in pure Python code using coding agents running against the comprehensive html5lib-tests testing library. Last …![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/icon/favicon-10.ico)Simon Willison’s WeblogSimon Willison![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/thumbnail/justjshtml-better-card.jpg)](https://simonwillison.net/2025/Dec/15/porting-justhtml/) ### Interesting Testing URL: https://www.workroom-productions.com/interesting-testing/ Last updated: 2026-02-12T14:16:43.000Z I find these provocative and perhaps challenging – let's dig in. ## January 2026 – Workroom Playtime [A Software Library with No CodeDo we still need libraries of 3rd party code when AI agents are this good?![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/icon/apple-touch-icon-1.png)Drew BreunigDrew Breunig![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/thumbnail/ikea_instructions.jpg)](https://www.dbreunig.com/2026/01/08/a-software-library-with-no-code.html) [ARC PrizeARC Prize is a $1,000,000+ nonprofit, public competition to beat and open source a solution to the ARC-AGI benchmark.![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/icon/favicon-2.png)ARC Prize![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/thumbnail/og-image-default.jpg)](https://arcprize.org) [claude-plugins-official/plugins/ralph-loop at main · anthropics/claude-plugins-officialOfficial, Anthropic-managed directory of high quality Claude Code Plugins. - anthropics/claude-plugins-official![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/icon/pinned-octocat-093da3e6fa40-15.svg)GitHubanthropics![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/thumbnail/claude-plugins-official)](https://github.com/anthropics/claude-plugins-official/tree/main/plugins/ralph-loop) We'll run all those together in late Feb [Workroom PlayTime Longplayer: Interesting Testing90-minutes online interactive workshop on Interesting Testing.![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/icon/wpl-60x60-logo-semitrans-43.png)Workroom ProductionsJames Lyndsay![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/thumbnail/photo-1576950133428-53b513ba7847.jpeg)](https://www.workroom-productions.com/workroom-playtime-longplayer-interesting-testing/) ## In the queue [How the GNU coreutils are testedTools and techniques used to test coreutils![](https://static.ghost.org/v5.0.0/images/link-icon.svg)](https://www.pixelbeat.org/docs/coreutils-testing.html) [curl - Tests Overview![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/icon/curl-symbol.svg)Tests Overview![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/thumbnail/curl-white-symbol.svg)](https://curl.se/dev/tests-overview.html) [Baiting the botLLM chatbots can be engaged in endless “conversations” by considerably simpler text generation bots. This has some interesting implications.![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/icon/https-3A-2F-2Fbucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com-2Fpublic-2Fimages-2Fb3e0ef3d-e542-4fb9-9b01-9f69cf091103-2Fapple-touch-icon-180x180.png)Conspirador NorteñoConspirador Norteño![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/thumbnail/https-3A-2F-2Fsubstack-post-media.s3.amazonaws.com-2Fpublic-2Fimages-2F63471085-de45-4c5a-9121-7d34489ba7f1_1998x1286.png)](https://www.conspirator0.com/p/baiting-the-bot) Mutation testing [Exploratory Interfaces 4: TestsExplore the passing tests of a built product to explore the product.![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/icon/wpl-60x60-logo-semitrans-39.png)Workroom ProductionsJames Lyndsay![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/thumbnail/photo-1574645434327-cb7970fe2e13-1.jpeg)](https://www.workroom-productions.com/exploratory-interfaces-4-tests/) ### No Workroom PlayTime on 8 Jan URL: https://www.workroom-productions.com/no-workroom-playtime-on-8-jan/ Last updated: 2026-01-07T22:04:29.000Z Apologies. I can't get my ducks in a row. See you next week. James ### Happy 2026 URL: https://www.workroom-productions.com/happy-2026/ Last updated: 2026-01-06T14:56:27.000Z We're all in 2026 together. _This post is for subscribers only._ ### Workroom Playtime 043: Building Short Exercises URL: https://www.workroom-productions.com/workroom-playtime-043-building-short-exercises/ Last updated: 2025-12-16T19:27:25.000Z This week's [Workroom PlayTime](https://www.workroom-productions.com/workroom-playtime/) is [Thursday 18 December at **3:00pm** London time](https://this-ti.me/?uts=1766026800&tz=Europe%2FLondon&name=Workroom+PlayTime+043). I've spent a year on short exercises – converting some from my existing materials, building some, delivering regularly. It's time for [a short exercise *about* short exercises](https://www.workroom-productions.com/building-short-exercises/). Let's not spend time on why they're useful, or what I've learned, or on a handful of polished principles. Let's learn by scribbling down new exercises and sharing. Maybe you're an old hand. Maybe not. Come to this and build your hundredth exercise, or your first. We'll share and take away new perspectives. We'll gather on Zoom. These exercises are for everyone, for free. [Subscribers](https://www.workroom-productions.com/#/portal/signup) get to play together every week. If you can see a **joining info** section below, then that gives you access to the Zoom etc. for the workshop. If you'd like to bring a friend, you can do that. I'll say yes until I've got too many. Here's a rough plan of what's coming up: - No sessions on 25th December, or 1 January (you might have something planned, yourself). - In 2026, we'll do more puzzles, more stuff for speakers, more thinking... and maybe fewer tools. - I'll re-run some exercises, and experiment more with timings and lengths, regularity and uncertainty. By the time we restart, in early January, we can mark a year of Workroom PlayTime – **thank you for coming on this journey.** ↓ *signed-in subscribers will see joining info and links below* ↓ _This post is for subscribers only._ ### Building Short Exercises URL: https://www.workroom-productions.com/building-short-exercises/ Last updated: 2025-12-18T16:11:43.000Z *A meta-exercise for exercise builders* ## Setup *Gather, read this – we'll move on in less than 5 mins* In this, we'll seek insights around how we build short interactive learning experiences. We'll split our time and attention roughly equally between making something, and sharing what we've made. You can work on something you've made, something you've experienced, or spectate. While we're gathering, we'll exchange: > What's *short*? ## Work *under 10 minutes. Alone or in pairs. Aiming to put text on the Miro board describing the exercise's activity, goal and more.* There are four paths through the next piece of work, depending on whether you've got an exercise in mind, whether you built it, and whether you want a guide. Open one of the choices below... *if you have made an exercise and want a guide to working on it...* ### Structured Path In this, you'll follow a sequence that has (sometimes) worked for me. You'll work *fast* – breadth being more important than depth. Even fast, we may not have enough time to do all this – if you're going over time, finish *two.* You can always come back to this another time if interested. #### Note a fundamental choice – 1 min Why short? Rather than longer... Why an interactive exercise? Rather than a poem or a lecture or a graph or a cartoon or whatever? *Describe your goal in bringing people together. Write something on how your goal might be reached with a short exercise, how your goal might be reached with an interactive exercise, and how that might be done coherently.* #### Constraints and necessities – 1 min Do you need a computer? An IDE? Do you need to be outdoors? Do you need to be familiar with something? Does this need to be an interactive exercise? Does it need to be short? Does your choice of a short interactive exercise work with the topic, or is there some strain? *Write a sentence on how you'll need to prepare to let your exercise be short exercise, and another show you'll need to prepare to let your exercise be interactive.* #### Participant activity – 2 mins *Describe the actions your participants will take, and the time they'll spend.* #### Get stuff to stick – 1 min What proportion of your short timebox will you use? Do you have a common pattern of interaction that you use? *Describe the actions you and the participants will take to review and absorb the ideas that have turned up.* #### Review what you've written – 1 mins Re-read, looking for things that need to be changed: coherence, ordering, supporting, time given. *Make changes* --- *if you have made an exercise and do *not* want a guide...* ### Unstructured Path Take 5 minutes to write / re-read / edit / tidy up a description of a short interactive exercise – try to give yourself a good chance of feeling done in 5. You'll describe participant activity and your goal, and what you've done to make your exercise both short and interactive. *Write wherever you like, make sure it gets on to the Miro board* --- *if you have experienced a good exercise, but haven't got one to bring...* ### Descriptive Path Have you enjoyed a short interactive exercise? What do you remember about it? Write down the apparent goal, the activities that participants did, and more – especially around what enabled it to be short, and how interactivity worked. Take 5 minutes to write / re-read / edit / tidy up your description – try to give yourself a good chance of feeling done in 5. *Write wherever you like, make sure it gets on to the Miro board* --- *if you prefer to see what others are doing rather than writing something...* ### Observer's Path Have a think about what works about a short interactive exercise, and what doesn't. Do you have principles? Do you have examples? Why might someone want to build a short interactive exercise? Why might someone want to participate in a short interactive exercise? No need to write. If you're happy to share your opinions, let me know. ## Share *5 minutes collective, going round and hearing exercises* Give a swift spoken summary of your activity and goal, and what you've done to make the exercise both short and interactive. We'll cut the time to suit the exercises available – more content means less time, and if there are many exercises we'll follow the hover. Keep track of your reactions – confirmations, lightbulb moments, points of resistance ## Reflect *5 minutes to exchange* To talk about things that - confirmed what you already knew - restructured your thoughts - surprised you - challenged you --- ## Sources I recall Jerry Weinberg's workshop and use [his book series on interactive exercises](https://leanpub.com/b/ExperientialLearning). I use the Liberating Structures app to give me ideas for interactions and wrapups. [Liberating Structures - Introductionliberating structures, social invention.net, microstructures, disruptive innovation, behavior change, collaboration, social invention, diffusion of innovation, strategy, transformation, heuristics, complexity science, emergence![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/icon/favicon-1.png)IntroductionKeith McCandless, Henri Lipmanowicz![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/thumbnail/00_LS-Menu.png)](https://www.liberatingstructures.com) Lisa Crispin reminded me to include: [Training from the Back of the Room!: 65 Ways to Step As…From Sharon L. Bowman, the author of the best-selling T…![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/icon/favicon-9.ico)GoodreadsSharon L. Bowman![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/thumbnail/8141935.jpg)](https://www.goodreads.com/book/show/8141935-training-from-the-back-of-the-room) ## Extension This extends the *structured* path above. ### Why are you doing this? Be clear with yourself: are you leaning towards sending out information, giving people a learning experience, or something else? #### Illustrative Examples If I'm sending out information, perhaps I will be... - looking for an activity to fit the learning objectives I have in mind - putting information into people's brains - demonstrating / repeating / practicing a skill - checking their learning against something authoritative If I'm interested in giving people a learning experience, perhaps I will be... - doing an activity and looking for learning points to emerge - giving an experience - opening people's brains to change / reframing - seeing what they have discovered collectively If I'm interested in something else, perhaps I will be... - exploring ideas alongside the participants - trying out something I've made before - aiming primarily to entertain There are hybrids, of course. I very often find myself experiencing something, thinking that I've learned from it, recognising that I've learned something that is usefully teachable to my peers, then setting up and refining somehow-similar experiences. Here's a story about one such occasion. [Making the RasterReveal ExerciseBuilding an exercise involves much wrangling and working around unexpected behaviours.![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/icon/wpl-60x60-logo-semitrans-38.png)Workroom ProductionsJames Lyndsay![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/thumbnail/photo-1620672900676-216af0dee6f1)](https://www.workroom-productions.com/making-the-raster-reveal-exercise/) ### Workroom Playtime 042: Exploratory Interfaces 4 – Tests URL: https://www.workroom-productions.com/workroom-playtime-042-exploratory-interfaces-4-tests/ Last updated: 2025-12-10T23:15:27.000Z This week's [Workroom PlayTime](https://www.workroom-productions.com/workroom-playtime/) is [Thursday 11 December at **4:30pm** London time](https://this-ti.me/?uts=1765470600&tz=Europe%2FLondon&name=Workroom+PlayTime+042). I'm sorry for the late notice – finding a time has been a challenge. We'll be thinking of *tests* as an exploratory interface. We won't directly be judging the tests, but reframing with the idea that if a subject passes its tests, then those (passing) tests describe the subject directly enough to be explorable. [Exploratory Interfaces 4: TestsExplore the passing tests of a built product to explore the product.![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/icon/wpl-60x60-logo-semitrans-36.png)Workroom ProductionsJames Lyndsay![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/thumbnail/photo-1574645434327-cb7970fe2e13.jpeg)](https://www.workroom-productions.com/exploratory-interfaces-4-tests/) 2 We'll gather on Zoom. These exercises are for everyone, for free. [Subscribers](https://www.workroom-productions.com/#/portal/signup) get to play together every week. If you can see a **joining info** section below, then that gives you access to the Zoom etc. for the workshop. If you'd like to bring a friend, you can do that. I'll say yes until I've got too many. Here's a rough plan of what's coming up: - `042` 11 December: Exploratory interfaces 4: Tests - `043` 18 December: Going meta – let's build short exercises! - No sessions on 25th December, or 1 January (you might have something planned, yourself). So that's a full year of Workroom PlayTime – **thank you for coming on this journey.** In 2026, we'll do more puzzles, more stuff for speakers, more thinking... and maybe fewer tools. I'll re-run some exercises, and experiment more with timings and lengths. ↓ *signed-in subscribers will see joining info and links below* ↓ _This post is for subscribers only._ ### Exploratory Interfaces 4: Tests URL: https://www.workroom-productions.com/exploratory-interfaces-4-tests/ Last updated: 2025-12-04T15:15:34.000Z We can explore tests; reading through them to judge whether they are any good. Here is a different framing... - If tests are executable, then those tests describe the product in a way that can be programatically judged. - If the existing product *passes those tests*, then the tests **describe the existing product**. Not what we want from it, but the product itself. We can explore a product by exploring its passing tests. ## Exercise ### Explore existing tests *10 mins* Let's look at the tests for a [crossword maker](https://www.workroom-productions.com/testing-transparently/). Let's look specifically at those tests which describe valid and invalid crosswords. What can we discover about the crossword maker by collating those tests that describe what it considers to be valid and invalid crosswords? We'll take `subject_e`, and start by looking at - Note: you can see the assertions by clicking on each test. You'll need to look at the source of the tests to see the setup etc. ### More exploring Once you're in the (test) source, have a look around to see if we need to widen our search. ### Debrief *5 mins* What did we learn about the *product* from its tests? *Key question*: does the product care about valid and invalid crosswords? ### Extend... > We can explore a product by exploring its passing tests. *Can* we? ### Puzzle 38 URL: https://www.workroom-productions.com/puzzle-38/ Last updated: 2025-12-04T15:23:33.000Z The buttons and lamps obey a simple principle. **What is it?** *Not* currently supported by these lovely people I have a [Patreon](https://www.patreon.com/workroomprds): Patrons get their name on new Puzzle, get to see it early, and generally get to be involved. However, I'm not putting this out as a Patreon-supported puzzle, so my patrons don't get perks. Unless they want them. I will restart the Patreon next year, and there will be a short series of new puzzles. Use [Patreon](https://www.patreon.com/workroomprds) to help me make more, and I'll put your name on the list. Enjoy this? [Support another!](https://www.patreon.com/workroomprds) Built by James Lyndsay - [@workroomprds](http://twitter.com/workroomprds) © Workroom Productions 2025 ### Workroom Playtime 041: Ah, um... URL: https://www.workroom-productions.com/workroom-playtime-041-ah-um/ Last updated: 2025-12-04T15:24:47.000Z This week's [Workroom PlayTime](https://www.workroom-productions.com/workroom-playtime/) is [Thursday 4 December at **4pm** London time](https://this-ti.me/?uts=1764864000&tz=Europe%2FLondon&name=Workroom+PlayTime+041). That's *4*pm. ***4***. I pre-announced the session as something on exploratory interfaces... but when I went to work, I thought that I was delivering a new puzzle. As of now, we have *both* (and this is an updated page). [Puzzle 38Puzzle 38 – swift to explore. Probably.![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/icon/wpl-60x60-logo-semitrans-37.png)Workroom ProductionsJames Lyndsay![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/thumbnail/puzzle_038-2-1.png)](https://www.workroom-productions.com/puzzle-38/) [Exploratory Interfaces 4: TestsExplore the passing tests of a built product to explore the product.![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/icon/wpl-60x60-logo-semitrans-36.png)Workroom ProductionsJames Lyndsay![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/thumbnail/photo-1574645434327-cb7970fe2e13.jpeg)](https://www.workroom-productions.com/exploratory-interfaces-4-tests/) We'll gather on Zoom to ... play. These exercises are for everyone, for free. [Subscribers](https://www.workroom-productions.com/#/portal/signup) get to play together every week. If you can see a **joining info** section below, then that gives you access to the Zoom etc. for the workshop. If you'd like to bring a friend, you can do that. I'll say yes until I've got too many. Here's a rough plan of what I plan to cover in the rest of the year: - `041` 4 December: Building exploratory interfaces - `042` 11 December: a New Blackbox puzzle - `043` 18 December: Going meta – let's build short exercises! I don't think we'll do 25th December, or 1 January (you might have something planned, yourself). So that's a full year of Workroom PlayTime – **thank you for coming on this journey.** In 2026, we'll do more puzzles, more stuff for speakers, more thinking... and maybe fewer tools. I'll re-run some exercises, and experiment more with timings and lengths. ↓ *signed-in subscribers will see joining info and links below* ↓ _This post is for subscribers only._ ### Add Commas URL: https://www.workroom-productions.com/add-commas/ Last updated: 2025-11-27T20:52:45.000Z This [tiny tool](https://www.workroom-productions.com/tag/tiny-tool/) adds commas to every line on your list. It's a bookmarklet, and takes advantage of your browser's permissions to go to work on your clipboard. I made it in response to a need that turned up in our [Crafting Custom Tools](https://www.workroom-productions.com/crafting-custom-tools/) tute at Agile Testing Days 2025\. Did I write it? No, I asked Claude: `Give me a bookmarklet to change the clipboard by adding a comma on the end of every line`, and this is part of what it made. Then I noticed that, for *some* text, it adds a newline with a comma in between all lines but the last – and that's because it was looking for `LF` line endings and not coping as well with `CRLF` endings. Drag the link below to your bookmark bar. Copy a list, hit the bookmarklet, paste the list. [add commas](javascript:%28async%28%29=>{let t=await navigator.clipboard.readText%28%29;let n=t.split%28/\r?\n/%29.map%28e=>e.trim%28%29?e.trim%28%29+',':e%29.join%28'\n'%29;await navigator.clipboard.writeText%28n%29;alert%28'Commas added to each line!'%29}%29%28%29) Wrinkles: - It may ask you for permission the first time you use it with a particular site and browser. - On Chrome, I've just answered to "BBC News wants see text and images copies to your clipboard" – so take care what sites you give permission to, and revoke that in the site settings. - On Safari I get a miniature modal with just the word `paste` in it, every time. - It acts invisibly, and I've seen it *just not work* (amusingly, during a demonstration to a room-full of testers at ATD). So it has an alert that tells you it has done something. Edit the alert out if you want. - I've no idea what circumstances it needs to work: It works on text and an Excel snippet, via Safari, for me. But for rich text, via Firefox, on your machine? That's for you to find out. There's perhaps more work to do to make thus a transferrable or robust tool. But this is an *ephemeral* tool, and I'd maybe just make a few more, or be conscious as I used it, rather than seek to understand its every wrinkle. #### The need One of our participants copied long columns of IDs, and needed them to have commas on the end at the point when they pasted the information into the next thing. Bookmarklets can futz with your clipboard, so this seemed a simple option. ### Workroom PlayTime 040: Explore Generated Code II URL: https://www.workroom-productions.com/workroom-playtime-040-explore-generated-code-ii/ Last updated: 2025-11-26T12:11:26.000Z This week's [Workroom PlayTime](https://www.workroom-productions.com/workroom-playtime/) is [Thursday 27 November at **3:30 pm** Berlin time](https://this-ti.me/?uts=1764167400&tz=Europe%2FBerlin&name=Workroom+PlayTime+040). That's *half three*. ***15:30***. And BERLIN time – which means 14:30 London time. We'll be on Zoom, and live at AgileTD at the table behind reception near the big door. I'm sitting at that table as I write this, with some of you... We'll explore different versions of generated code – [four versions of a crossword generator](https://www.workroom-productions.com/testing-transparently/). These exercises are for everyone, for free. [Subscribers](https://www.workroom-productions.com/workroom-playtime-039-stuck-overflow/#/portal/signup) get to play together every week. If you can see a **joining info** section below, then that gives you access to the Zoom etc. for the workshop. If you'd like to bring a friend, you can do that. I'll say yes until I've got too many. - `041` 4 December: Building exploratory interfaces (longer session) - `042` 11 December: a New Blackbox puzzle - `043` 18 December: Going meta – let's build short exercises! I don't think we'll do 25th December, or 1 January (you might have something planned, yourself). So that's a full year of Workroom PlayTime – **thank you for coming on this journey.** In 2026, we'll do more puzzles, more stuff for speakers, more thinking... and maybe fewer tools. I'll re-run some exercises, and experiment more with timings and lengths. ↓ *signed-in subscribers will see joining info and links below* ↓ _This post is for subscribers only._ ### Testing Transparently URL: https://www.workroom-productions.com/testing-transparently/ Last updated: 2025-12-01T15:24:40.000Z *ATD2025 keynote, Elizabeth Zagroba and James Lyndsay* > Scenario: “the tests all pass – but this is *generated* code. **What is weird?**” *This has been built for a conference audience of testers, and for Elizabeth and James. But they’re imagining this as a web-based tool to allow teachers to swiftly build crosswords for kids. Commercially, it’s part of a package, rather than a paid-for service.* ### subject2\_e – used at AgileTD Join us exploring this [word grid generator](https://exercises.workroomprds.com/tt%5Fsubject2%5Fe) – version `subject2_e` Test Resources: [test runner](https://exercises.workroomprds.com/tt%5Fsubject2%5Fe/testRunner.html), [user info](https://exercises.workroomprds.com/tt%5Fsubject2%5Fe/userInfo.md), [decisions made while building](https://exercises.workroomprds.com/tt%5Fsubject2%5Fe/decisions.md), [build diary](https://exercises.workroomprds.com/tt%5Fsubject2%5Fe/diary.md) and [agent info](https://exercises.workroomprds.com/tt%5Fsubject2%5Fe/claude.md). Examples: [keynote words 10x10](https://exercises.workroomprds.com/tt%5Fsubject2%5Fe/?words=ELIZABETH%2CJAMES%2CKEYNOTE%2CNOVEMBER%2CPOTSDAM%2CTRANSPARENTLY%2CTESTING%2CDAYS%2CTESTING%2CAGILE&width=12&height=12&all%5Fcross=true&allow%5Fdiagonals=false&allow%5Freverse=false), [3x3](https://exercises.workroomprds.com/tt%5Fsubject2%5Fe/?words=MOB%2CJIM%2CBOB%2CJOB&width=3&height=3&all%5Fcross=true&allow%5Fdiagonals=false&allow%5Freverse=false), [5x10](https://exercises.workroomprds.com/tt%5Fsubject2%5Fe/?words=OTTER%2CFIFTY%2CTYPEWRITER&width=10&height=5&all%5Fcross=true&allow%5Fdiagonals=false&allow%5Freverse=false) --- ## Other versions ### final\_a Explore this [word grid generator](https://exercises.workroomprds.com/tt%5Ffinal%5Fa) – version `final_a` Test Resources: [test runner](https://exercises.workroomprds.com/tt%5Ffinal%5Fa/testRunner.html), [user info](https://exercises.workroomprds.com/tt%5Ffinal%5Fa/userInfo.md), [decisions made while building](https://exercises.workroomprds.com/tt%5Ffinal%5Fa/decisions.md), [build diary](https://exercises.workroomprds.com/tt%5Ffinal%5Fa/diary.md) and [agent info](https://exercises.workroomprds.com/tt%5Ffinal%5Fa/AGENTS.md). Examples: [keynote words 10x10](https://exercises.workroomprds.com/tt%5Ffinal%5Fa/?words=ELIZABETH%2CJAMES%2CKEYNOTE%2CNOVEMBER%2CPOTSDAM%2CTRANSPARENTLY%2CTESTING%2CDAYS%2CTESTING%2CAGILE&width=12&height=12&all%5Fcross=true&allow%5Fdiagonals=false&allow%5Freverse=false), [3x3](https://exercises.workroomprds.com/tt%5Ffinal%5Fa/?words=MOB%2CJIM%2CBOB%2CJOB&width=3&height=3&all%5Fcross=true&allow%5Fdiagonals=false&allow%5Freverse=false), [5x10](https://exercises.workroomprds.com/tt%5Ffinal%5Fa/?words=OTTER%2CFIFTY%2CTYPEWRITER&width=10&height=5&all%5Fcross=true&allow%5Fdiagonals=false&allow%5Freverse=false) ### final\_b Explore this [word grid generator](https://exercises.workroomprds.com/tt%5Ffinal%5Fb) – version `final_b` Test Resources: [test runner](https://exercises.workroomprds.com/tt%5Ffinal%5Fb/testRunner.html), [user info](https://exercises.workroomprds.com/tt%5Ffinal%5Fb/userInfo.md), [decisions made while building](https://exercises.workroomprds.com/tt%5Ffinal%5Fb/decisions.md), [build diary](https://exercises.workroomprds.com/tt%5Ffinal%5Fb/diary.md) and [agent info](https://exercises.workroomprds.com/tt%5Ffinal%5Fb/AGENTS.md). Examples: [keynote words 10x10](https://exercises.workroomprds.com/tt%5Ffinal%5Fb/?words=ELIZABETH%2CJAMES%2CKEYNOTE%2CNOVEMBER%2CPOTSDAM%2CTRANSPARENTLY%2CTESTING%2CDAYS%2CTESTING%2CAGILE&width=12&height=12&all%5Fcross=true&allow%5Fdiagonals=false&allow%5Freverse=false), [3x3](https://exercises.workroomprds.com/tt%5Ffinal%5Fb/?words=MOB%2CJIM%2CBOB%2CJOB&width=3&height=3&all%5Fcross=true&allow%5Fdiagonals=false&allow%5Freverse=false), [5x10](https://exercises.workroomprds.com/tt%5Ffinal%5Fb/?words=OTTER%2CFIFTY%2CTYPEWRITER&width=10&height=5&all%5Fcross=true&allow%5Fdiagonals=false&allow%5Freverse=false) ### subject2\_f Explore this [word grid generator](https://exercises.workroomprds.com/tt%5Fsubject2%5Ff) – version `subject2_e` Test Resources: [test runner](https://exercises.workroomprds.com/tt%5Fsubject2%5Ff/testRunner.html), [user info](https://exercises.workroomprds.com/tt%5Fsubject2%5Ff/userInfo.md), [decisions made while building](https://exercises.workroomprds.com/tt%5Fsubject2%5Ff/decisions.md), [build diary](https://exercises.workroomprds.com/tt%5Fsubject2%5Ff/diary.md) and [agent info](https://exercises.workroomprds.com/tt%5Fsubject2%5Ff/claude.md). Examples: [keynote words 10x10](https://exercises.workroomprds.com/tt%5Fsubject2%5Ff/?words=ELIZABETH%2CJAMES%2CKEYNOTE%2CNOVEMBER%2CPOTSDAM%2CTRANSPARENTLY%2CTESTING%2CDAYS%2CTESTING%2CAGILE&width=12&height=12&all%5Fcross=true&allow%5Fdiagonals=false&allow%5Freverse=false), [3x3](https://exercises.workroomprds.com/tt%5Fsubject2%5Ff/?words=MOB%2CJIM%2CBOB%2CJOB&width=3&height=3&all%5Fcross=true&allow%5Fdiagonals=false&allow%5Freverse=false), [5x10](https://exercises.workroomprds.com/tt%5Fsubject2%5Ff/?words=OTTER%2CFIFTY%2CTYPEWRITER&width=10&height=5&all%5Fcross=true&allow%5Fdiagonals=false&allow%5Freverse=false) --- ## Building These were built using agentic tool [amp.code](https://ampcode.com/), with the following prompt: > We've lost the css, index and all the javascript. We do have the documentation, and all the tests. All the tests were passing before. Please read everything in the project (but ignore the change history). Re-create the source from the docs and the tests, ensuring that all the tests pass and that the decisions are respected. > *When you're done, document your actions in a file @diary.md* You can build versions yourself, using the files above. The versions prefixed `subject` were built with amp running primarily on Claude Sonnet 4.5\. The versions prefixed `final` were built with slight differences to the documentation and tests, and by that time amp was using GeminiPro. While we would have preferred the slightly simpler needs and tighter tests of the `final` versions, the generated UIs had enough irritations that we felt participants would be distracted by those bugs, rather than play with the deeper code. After the keynote, we ran `subject2_e` through testers.ai's tools, which reported [finding 0 bugs](https://www.linkedin.com/posts/jameslyndsay%5Fagiletd-activity-7399918091951939584-2C0Q?utm%5Fsource=share&utm%5Fmedium=member%5Fdesktop&rcm=ACoAAAAyiC8BkEoqTadPzV9zHdb5crUj0ZK30y0). ![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/2025/12/image.png) ### Workroom Playtime 039: Stuck Overflow URL: https://www.workroom-productions.com/workroom-playtime-039-stuck-overflow/ Last updated: 2025-11-19T13:54:35.000Z This week's [Workroom PlayTime](https://www.workroom-productions.com/workroom-playtime/) is [Thursday 20 November at **3pm** London time](https://this-ti.me/?uts=1763650800&tz=Europe%2FLondon&name=Workroom+PlayTime+039). That's *3*pm. ***3***. We'll gather on Zoom for [Stuck Overflow](https://www.workroom-productions.com/playful-exercises-about-testing/#stuck-overflow) – a playful exercise about testing. These exercises are for everyone, for free. [Subscribers](https://www.workroom-productions.com/#/portal/signup) get to play together every week. If you can see a **joining info** section below, then that gives you access to the Zoom etc. for the workshop. If you'd like to bring a friend, you can do that. I'll say yes until I've got too many. - `040` 27 November: from ATD! I might run [The Consultant](https://www.workroom-productions.com/playful-exercises-about-testing/#the-consultant). - `041` 4 December: Building exploratory interfaces (longer session) - `042` 11 December: a New Blackbox puzzle - `043` 18 December: Going meta – let's build short exercises! I don't think we'll do 25th December, or 1 January (you might have something planned, yourself). So that's a full year of Workroom PlayTime – **thank you for coming on this journey.** In 2026, we'll do more puzzles, more stuff for speakers, more thinking... and maybe fewer tools. I'll re-run some exercises, and experiment more with timings and lengths. ↓ *signed-in subscribers will see joining info and links below* ↓ _This post is for subscribers only._ ### Quick Trick Automation URL: https://www.workroom-productions.com/quick-trick-automation/ Last updated: 2025-11-19T00:18:53.000Z A 25-minute session at Agile Testing Days #### Promise - How to use combinations of simple (OS) tools to make you life easier - Disposable tools can make a difference in efficiency and productivity - Automating tedious tasks can be so much more fun than the tasks themselves. Automate away your tediousnesses In testing we are bound to run into tedious tasks that we have to do as part of testing, but we would rather offload to be able to do the interesting parts. In this presentation James and Bart will teach you how quick automation tricks can get you out of having to do some of these tasks manually. Clever usage of combinations of simple tools and knowing how to use the powertools most of us have on their laptops any way, can drastically improve efficiency productivity and fun: - in executing complex testcases, - doing intricate comparisons or - setting up structured datatypes. See how James and Bart tackle these issues and get a sneak preview of the toolkit they have built during their tutorial earlier in the week. ## Tiny tool gallery [tiny tool - Workroom ProductionsShort, delicate, specific, swiftly-made tools![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/icon/wpl-60x60-logo-semitrans-34.png)Workroom Productionstiny tool![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/thumbnail/photo-1542359498-13ebad248020)](https://www.workroom-productions.com/tag/tiny-tool/) ## Workshop-built tools ### Crafting Custom Tools URL: https://www.workroom-productions.com/crafting-custom-tools/ Last updated: 2025-11-26T23:39:42.000Z *A* [*tutorial at Agile Testing Days*](https://agiletestingdays.com/2025/session/the-art-of-crafting-your-custom-tools/) Huib, Bart and James #### Promise A lot of the activities you need to do as part of test execution/validation/testautomation you can ‘automate/generate’ by building your own tools, either by combining existing tooling or GEN AI. Key learnings - Automate to simplify test processes and access new ways of working. - Defer dull and precise testing jobs to simple tools. - Use existing tools, libraries and GEN-AI to build your own tools. Build the Right Tool for Right Now **Efficient testing needs good tooling – and the right custom tools enable powerful, focussed test approaches. In this innovative workshop, you’ll combine common, accessible tools to build automation for your own testing tasks.* **We’ll focus on testing-relevant work like comparing datasets, generating bulk data, wrangling environments, enhancing unit tests with custom matchers and exploring code with generated UIs. We’ll use spreadsheets, shell prompts, and LLMs to access simple automation that perfectly suits your teams’ needs.* **After a grounding in the basics of building tools for tasks, our three experienced presenters will help you work together on custom tools. By the end of this interactive tutorial you will have the confidence and starting skills to work on your own toolset – and a handful of example tools that you’ve already built. Enhance your testing powers by making the right tool for your work.* #### Tools We can all buy or find tools – this is about making tools, generally by chaining and redeploying tools we can all access. A **tool* is something we control that does something better than we can unequipped. In the physical world, tools generally have a place where the human interacts, a part that changes force or size or precision, and a place that interacts with whatever is being worked on. Tools generally outlast the work, but disposable / ephemeral tools are important, if easy to overlook. A tool may not do the work, but make some part of the work possible; we can do some without a ruler or a dust sheet, but we might well not choose to start the work. In testing, we use tools all the time. To work as professionals and as craftspeople, James (+) reckons you need to be able to make the tools you need. ## Logistics We run in 08:30 – 10:00, 10:30 – 12:00, 13:00 – 14:30, 15:00 – 16:00 You'll need something internet-connected that you can type on. We have [a Miro board](https://miro.com/app/board/uXjVJmeW60k=/?share%5Flink%5Fid=262468796980), just in case. ## Background These materials may be useful during the workshop [Parts for ToolsTool making from what you find![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/icon/wpl-60x60-logo-semitrans-23.png)Workroom ProductionsJames Lyndsay![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/thumbnail/photo-1615746363486-92cd8c5e0a90-1.jpeg)](https://www.workroom-productions.com/parts-for-tools/) #### Tools to make tools These let you build custom tools – often around flowchart-like or graph-like IDEs: - - [PowerAutomate](https://make.powerautomate.com/) (Win) Bart, - [AutoHotKey](https://www.autohotkey.com) (Win), - [Automator](https://support.apple.com/guide/automator/welcome/mac) (Mac) James, - Shortcuts ([Mac](https://support.apple.com/en-gb/guide/shortcuts-mac/apdf22b0444c/mac) / [iOS](https://apps.apple.com/us/app/shortcuts/id915249334)) James, - [Keyboard Maestro](https://www.keyboardmaestro.com/main/) (Mac) James , - [IFTTT](https://ifttt.com) (web + remote services) James, - [Zapier](https://zapier.com) (web + remote services), - [n8n](https://n8n.io) (Docker-based UI for tools, from command-line to LLM integrations) - [Ansible](https://docs.ansible.com) (commandline tool to configure servers) #### Example of how tools are used and chained This example uses commandline tools. You can imagine the same with Excel, or try threading something similar in n8n. - It may be that you've got a log from a test system and you want to extract all lines that refer to a particular transaction that seemed problematic. You'd `grep`. - Maybe you've got several logs, and you need to look for every reference. You could merge them while paying attention to the timestamps with a `sort`, then `grep`, and you'd join those two with a pipe `|`. - Maybe those logs are remote, but you can get them with `cURL`. You `|` them into `sort`. Maybe you want to add a suffix to each line, so that you can see the source. You `curl` | `sed` | `sort` | `grep` - Maybe you can't read the pages of scrolling weirdness. You `|` the `grep` into `tail`, to see the last 10 lines. - At the end, you've got a single hard-to-read but repeatable command `curl` | `sed` | `sort` | `grep` | `tail` – you use it once and throw it away. Or keep it in a handy spot. ### Workbench We've set up a server which is running several [instances of VSCode](https://github.com/coder/code-server), accessible via the browser. Via VSCode, you'll access the command line, the file system, and whatever else is available. I hope that you feel relatively familiar with it. #### Access 1. Go to [cct01](https://cct01.workroomprds.com) or [cct02](https://cct02.workroomprds.com) or [cct03](https://cct02.workroomprds.com) and pick the env that seems most ... you. Tell the others in your group so you don't clash names. Your password is `password`. VSCode may want to set up a more-complete env, so skipor move on until you see something like this: ![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/2025/02/image-1.png) 1. Open up the files and the commandline, and choose the `terminal` tab if you need to. 2. The chunk of green text is your *prompt* – where you'll type commands. Try `pwd` to see where you are: probably `/home/«whatever you chose»` ![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/2025/02/image-2-1.png) 1. Use *Open Folder* to get a dropdown, probably pointing to the same place (with a `\` on the end. If it's not that, change it and hit enter to see the filesytem. If asked, you *trust the author*. You'll see a filesystem sidebar on the left, and you may need to open the terminal again to get to (roughly) here: ![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/2025/02/image-4.png) So now you've got a browse-able file system, and commandline access to a remote server, all in your browser. Which is what we need to start. *One irritation: copy / paste from your host may not work for you. Or it might. If nit doesn't try right-click. I'm working on the config. It works for me...* 1. If you're unfamiliar with VSCode, take a moment to look around. Open the three-bar menu top-left, resize the parts if you want to, see the tooltips. #### VSCode Parts - ****Mode Panel** (Narrow leftmost Sidebar): - - The menu icon at the top gives access to the VSCode menus - The icons below change what the panel to its side does – file explorer, search, git, debugger, extensions. - ****Explorer Panel** (Left Sidebar): - - Shows your files and folders / search results / git history etc. - ****Editor Area** (Center-right): - - Currently shows a "Welcome" tab - Has quick actions like "New File..." and "Open File..." - When you select a file, you'll see it here. - Can have tabs, and can be split vertically - ****Terminal** (Bottom): - - Command-line interface, with prompt - You can run commands directly here without leaving the editor - "OUTPUT" tab: Displays program output - "TERMINAL" tab: For command-line operations - ****Top Bar Elements**: - - Command bar for VSCode commands - Top right buttons show layouts - ****Bottom Status Bar**: - - Shows useful information like Current file type (bash), Layout settings, Additional status indicators ## Materials *Note: headings below "fold up". They're left open for search, and (until I sort out a css thing, don't have foldup / down arrows)* ### Setup 50 mins #### Exercise: Scope *10 mins warm up, ending with tools you want to take home* Talk to your table. - Who are you? - What do you expect to do/learn today? - What tool would YOU like to build for work? *At the end,* we'll share with the room – and in particular we'll try to capture the tools you want to build, and work towards them. *Hint: think small.* #### Who we are *5 mins* #### Exercise: How to Count (Using Tools) *10 mins: Using tools, sharing thoughts* Count the "Bart"s, "Huib"s and "James"s in ATD's page for this tutorial - Use whatever method you like. - Then use a different method or tool. [Agile Testing DaysAgile Testing Days - November 23 - 26, 2024 in Potsdam, Germany - Europe’s GreaTest Agile Testing Conference for Software Testers, Developers & Managers![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/icon/agiletd24-icon_color.svg)trendig![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/thumbnail/agiletd24-icon_color.svg)](https://agiletestingdays.com/2025/session/the-art-of-crafting-your-custom-tools/) *At the end,* compare what you've done #### Exercise: Parts for Tools / triggers for tools / tools for tools *10 mins: Hands-on with a framework to help think about tools* *10 mins: sharing what we use* ##### This exercise will get us thinking about building tools from parts. We'll consider any executable software tools, from browser bookmarklets to Ansible playbooks. We'll necessarily consider context: Where a tool works, what it works on, when it might be needed, and what you do to run it. [Exercise: Parts for ToolsPurposefully picking parts for tools.![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/icon/wpl-60x60-logo-semitrans-33.png)Workroom ProductionsJames Lyndsay![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/thumbnail/photo-1615746363486-92cd8c5e0a90-1-2.jpeg)](https://www.workroom-productions.com/exercise-parts-for-tools/) [Parts for ToolsTool making from what you find![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/icon/wpl-60x60-logo-semitrans-24.png)Workroom ProductionsJames Lyndsay![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/thumbnail/photo-1615746363486-92cd8c5e0a90-1-1.jpeg)](https://www.workroom-productions.com/parts-for-tools/) **Gallery:** Our tools - James' [tiny tools](https://www.workroom-productions.com/tag/tiny-tool/) – scripts, typically. Including bookmarkets, scripts and prompts. - Bart's *Tool Migration Spreadsheet* - Huib's handy sites #### Bart's spreadsheet – experience report I was asked to convert the export of one test management tool into importable content into another test management tool, where the output contained 2 files (one with the testcases one with the teststeps), that needed to be combined into single file, so I wrote a excel macro to get it done. Excel Macro (non-commercial) at ... (**Bart)* #### Huib's handy sites - [bit.ly/testersplayground](http://bit.ly/testersplayground) - [bit.ly/testingtoolsrp](http://bit.ly/testingtoolsrp) - [bit.ly/evaluatetools](http://bit.ly/evaluatetools) - [bit.ly/alternativetools](http://bit.ly/alternativetools) - - - - [RANDOM.ORG - True Random Number Service](https://www.random.org/) - - ##### ***At the end,*** we'll refine or revise what tools we'd like to take home ### Part 1: Tools already on your laptop *35-40 mins to Break 1* #### Group: Tools Exchange *10 mins* Round Robin: *What tools do you already have with you* #### Demo: Bart example in Google sheets and Excel *5 mins* Using a Macro *@Bart: send spreadsheet for JL to attach.* #### Exercise: Find the Right Data to Use *15 mins* We're going to build a data-picking tool – you can use macros and / or formulae in sheets (or whatever else, in whatever tool you like) Here are two data sets: [datadata.csv37 KBdownload-circle](https://www.workroom-productions.com/content/files/2025/11/data.csv "Download") [usedused.csv2 KBdownload-circle](https://www.workroom-productions.com/content/files/2025/11/used.csv "Download") The first dataset, for accounts, has several fields: - `UniqueID` - `amount` - `status` (unknown, bad, questionable, good) - `new_Flag` (True/False) - `suspended_Flag` (True/False) The second is a list of `UniqueIDs` already used in your testing. ##### Part 1: Data Picker Write a macro in a spreadsheet of your choice to find an account to fit the following specs: - positive amount of money in the account - status = good - suspended account - have not been used in an earlier test ##### Part 2: Data Aggregator Sum up the money in all suspended accounts ##### Stretch Alert the user when there are no available 'suspended' accounts with -ve money. *At the end,* let's summarise for each other **what tooling is viable for the tool we want to take home**. #### #### Optional Exercise – Custom Build **20 mins; build and show off a tool* Opportunity hunt – what can you imagine that has value to you or someone in your workplace. Pick one for your group. Make a prototype with the tools you have available right now. #### Marketplace *5-10 mins into the break.* What have we built? What did we learn? ### Part 2: Simple Tools for Complex Tasks *85 - 90 mins to lunch* *Let's review the tools we want to go home with.* #### *Command line tools* *5 mins on workbench and commandline tools* The command line was the primary way to interact with computers until the advent of GUIs in the mid 80s. Coincidentally, the IEEE's POSIX standard turned up in 1988, and codified many of the command-line tools used by UNIX. In this section, we'll get to grips with some of those tools. [UNIX philosophy](https://en.wikipedia.org/wiki/Unix%5Fphilosophy) proposed that those command line tools should - do one thing - work together - handle text streams as a universal interface #### Your Workbench If you've got a development environment that you'd like to use, use it. You'll want to clone the following for the exercises. If not, we've got VSCode in the browser for you. Go to [cct01](https://cct01.workroomprds.com) or [cct02](https://cct02.workroomprds.com) or [cct03](https://cct03.workroomprds.com) depending on which table you're at. Pick a name to be "you" for this session, log in (password is `password`), and pick `VSCode`. Once VSCode in the browser settles, you'll want to open the terminal. More details and pictures in the section **Background*: **Workbench* above. #### #### Exercises *40 mins* *We'll go to work on `cat`, `grep`, `sort`, `head` and the redirection operators – and modify based on the room.* - `grep ` – [info](https://www.workroom-productions.com/grep-for-testers/) – [exercises](https://www.workroom-productions.com/grep-tool-exercises/) - `sort` – [exercises](https://www.workroom-productions.com/playing-with-sort/) - `cat`, `head`, `tail` – [info](https://www.workroom-productions.com/cat-head-and-tail-for-testers/) – [exercises](https://www.workroom-productions.com/exercises-for-cat-head-and-tail/) - Redirection operators `|`, `>` `>>`, `<` (and `tee`) – [info](https://www.workroom-productions.com/redirection-operators-for-testers/), [exercise](https://www.workroom-productions.com/exercise-redirection-operators/) - `diff` – [info](https://www.workroom-productions.com/diff-for-testers/) – [exercises](https://www.workroom-productions.com/diff-tool-exercises/) - `curl` – [exercise](https://www.workroom-productions.com/curl-exercises-for-testers/) - `xargs` – [](https://www.workroom-productions.com/xargs-for-testers/)[exercise](https://www.workroom-productions.com/xargs-for-testers/) - `jq` – [exercise](https://www.workroom-productions.com/exercise-jq-for-testers/) #### Build *30 mins* Think about what you want to build. Put it into context – when is it used, where does it run, what does it work on. Have a look at the [framework on Patterns for Tools](https://www.workroom-productions.com/parts-for-tools/#framework), and consider the parts that fit into those patterns. - Map out some part of the thing that you want to take with you. - Get one part of the (tool you need) to work as a proof of concept. - Consider how you'll plug it into another part. #### Not building your own tool? You'll build something that extracts information from a file. That need means you'll need to feed that file into a filter, then manage the filtered information. We'll take logs as an example. ****Short exercise:** Using the log from the grep exercise, find the top 10 IP addresses that have the message `Directory index forbidden`. Here's the link ****Longer Exercise:** Let's talk. Here are a couple of sources for files that are explorable. - [GitHub - SoftManiaTech/sample\_log\_files: A large collection of system log datasets for log analysis researchA large collection of system log datasets for log analysis research - SoftManiaTech/sample\_log\_filesGitHubSoftManiaTech](https://github.com/SoftManiaTech/sample%5Flog%5Ffiles) - [GitHub - logpai/loghub: A large collection of system log datasets for AI-driven log analytics \[ISSRE′23\]A large collection of system log datasets for AI-driven log analytics \[ISSRE′23\] - logpai/loghubGitHublogpai](https://github.com/logpai/loghub) #### Marketplace *5-10 mins, into lunch* - What have we built? - What did we learn? - Are we closer to what we wanted to build for our work? ### Part 3: Generated Tools *90 mins to Break 2* LLMs can help you build custom tools – if you've got an example and a clear need, and you feel able to spot where the tool might be doing something wrong, you'll get there more easily. #### LLMs and tools LLMs can **act* as tools – but they're inconsistent, limited and expensive. A better use, if we don't need the inherent text-processing and randomness abilities of LLMs, is to ask the LLM to **make* a tool. #### Possible prompting patterns Start with an example of input → output if transforming, or of event → action if triggering. - Describe the situation (attach examples if you can), and ask for a re-statement and options to address it with automation. If the response could be refined, **re-ask rather than converse.** - Ask for executable code and how to set up the necessary infrastructure. - Inspect the code, and run it under your own permissions. **Test it.* - If the results aren't satisfying, describe the results in the same conversation that had the code. Ask for options to improve. Make changes to the **initial* prompt and go from the top. - If the results are satisfying, inspect again, ask for clarifications, ask an LLM to explain the code in a different conversation, test more, then check in the code **if you're happy to take responsibility for it*. If it's an ephemeral tool, keep it around for as long as you need it – then chuck it. Potentially keep track of the final tool prompt request. #### Using LLMs You probably want to use LLMs that you already have set up, or you ,may be familiar with using popular LLMs via their web pages or apps. If you'd like to use them on the commandline, we have temporary keys for OpenAI and for Anthropic, and you'll use them via the [commandline llm tool](https://llm.datasette.io/en/stable/). !! Activate the tool with `source ~/llm-env/bin/activate` Example: - `llm -m claude-haiku-4.5 "Who are you?"` #### Demo: Generating Bulk Data Bart – asking the LLM to make the data Huib – asking an LLM to make a tool to make the data #### Exercise: commandline tool *20 minutes* LLMs are sometimes good at commandline tools. They're better at *explaining* a commandline tool. For this exercise, we'll ask for the tool (in various LLMs) and then ask it to explain. [Tiny Tool: checkRefs.shA tiny tool to see which files include one or more of a set of strings.![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/icon/wpl-60x60-logo-semitrans-31.png)Workroom ProductionsJames Lyndsay![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/thumbnail/photo-1564901523975-b18a3eb1d11f)](https://www.workroom-productions.com/tiny-tool-checkrefs-sh/) #### Exercise: Custom asserts *20 minutes* BART - js matcher for google sheet [Exercise: Custom AssertsPlay with building custom asserts![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/icon/wpl-60x60-logo-semitrans-25.png)Workroom ProductionsJames Lyndsay![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/thumbnail/photo-1604881991405-b273c7a4386a-1.jpeg)](https://www.workroom-productions.com/custom-asserts-exercise/) [Custom Asserts – what and whyWhat they are, why they’re handy, how to start![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/icon/wpl-60x60-logo-semitrans-26.png)Workroom ProductionsJames Lyndsay![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/thumbnail/photo-1604881991405-b273c7a4386a-1-1.jpeg)](https://www.workroom-productions.com/custom-asserts/) #### Exercise: Exploratory interface *20 minutes* [Exploratory Interfaces alongside unit testsLooking inside the code to see what it offers, then playing with it.![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/icon/wpl-60x60-logo-semitrans-27.png)Workroom ProductionsJames Lyndsay![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/thumbnail/photo-1730578726210-82ce0f955496-1.jpeg)](https://www.workroom-productions.com/exploratory-interfaces-alongside-unit-tests/) [Exploratory InterfacesFinding good ways to explore a system![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/icon/wpl-60x60-logo-semitrans-28.png)Workroom ProductionsJames Lyndsay![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/thumbnail/bug-background-600x200.jpeg)](https://www.workroom-productions.com/exploratory-interfaces/) #### Wrangling Environments James uses Ansible to work with environments. There's such a broad set of possibilities to this that we won't be doing a hands-on exercise in this (though if you have something you'd like to automate, off we go!) Here's an example of one re-usable part: enabling https on a new environment. This example also includes an LLM's initial attempt, and the refinements made to it. [Environments: Cert workAnsible to get a certificate by leveraging CloudFlare with LetsEncrypt![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/icon/wpl-60x60-logo-semitrans-29.png)Workroom ProductionsJames Lyndsay![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/thumbnail/photo-1628529706799-b424a9fee803-1.jpeg)](https://www.workroom-productions.com/environments-cert-work/) #### #### Error Messages As a debugging trick, it can be handy to ask the LLM to read an error message alongside the code and data. #### Duff code Ask an LLM to review code (and critique) #### Purposeful code Ask an LLM to guess the intent of your latest tests – if it can't tell, you need to be more clear. Ask an LLM to guess the intent of someone else's old tests – can the LLM help you to understand the intent? #### Iterating with LLMs [Environments: Cert workAnsible to get a certificate by leveraging CloudFlare with LetsEncrypt![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/icon/wpl-60x60-logo-semitrans-29.png)Workroom ProductionsJames Lyndsay![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/thumbnail/photo-1628529706799-b424a9fee803-1.jpeg)](https://www.workroom-productions.com/environments-cert-work/) [Clear up your LinkedIn feedA bookmarklet to afford me some control over my LinkedIn feed.![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/icon/wpl-60x60-logo-semitrans-30.png)Workroom ProductionsJames Lyndsay![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/thumbnail/photo-1744646338661-c1e6530dfef4)](https://www.workroom-productions.com/clear-up-your-linkedin-feed/) [Tiny Tool: checkRefs.shA tiny tool to see which files include one or more of a set of strings.![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/icon/wpl-60x60-logo-semitrans-31.png)Workroom ProductionsJames Lyndsay![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/thumbnail/photo-1564901523975-b18a3eb1d11f)](https://www.workroom-productions.com/tiny-tool-checkrefs-sh/) #### Marketplace 3 *10-15 mins, into afternoon break* - What have we built? - What did we learn? - Are we closer to knowing what we might use to build what we want our work? ### Part 4: Build your own *60+ mins* We'll work with you to build (some of) the tools you wanted. We expect that we'll coalesce into small groups using command-line tools, your existing tools, and LLMs. At the end, we want several new tools that the group has built together, that can demonstrably do work that you find useful. ### Wrap Up *25-30 mins* #### Final Marketplace - *15 minutes* Let's show each other what we have built - *5 mins* What did we learn? - *5 mins* Our next steps, to take when we're back at work. ## After the workshop ## ### Parts for Tools URL: https://www.workroom-productions.com/parts-for-tools/ Last updated: 2025-11-24T05:42:31.000Z A page setting out a framework I use when thinking about making tools. Why a framework? Because it's handy. I don't know what tools I'll have to hand. A framework for how I can use (and misuse) the things I find around the place can help me to be more resourceful. A tool, for here, is something that acts for us. We use a tool because we work better with the tool. Tools can be pretty-much *anything*, by these lights: a checklist, a process, thirty years experience, a pocketknife. For this framework, we're mostly considering *executable software*. We'll need to think about where our tools typically run, what triggers them, and what they operate on. See the examples. #### Why make tools? As anyone who has seen a specialist's collection of hammers knows, the **right* tool for the job can differ from the **wrong* tool by what might seem to be small details. If you can't configure an existing tool just right, you need to **make* your tool. You'll need to be purposeful, to experiment, and to use the tools you have around you. You configure your tools, combine them, and use your new resulting tool for just the purpose you intended. And then you pick it up again for the **wrong* thing, and stick the sharp end through your thumb. Careful of what you build. #### Examples - I have a Bookmarklet tool that runs when I click a bookmark button. It acts on the front page in my browser, and gives me a form to send that to a bookmarking service. info -> action - I have an AppleScript tool that runs on my Mac, on request from my dock. It works with my browser and summarises its state in an alert. info->aggregate - I have a commandline tool that runs on my Mac when I trigger it, that sees what files with particular extensions in the current directory are referenced in files in another directory. info->filter->aggregate - I have an IFTTT tool that is triggered when someone takes a particular action on my website. It reacts to a webhook and sends me an email. - I have a HomeKit tool that runs in my home cloud, and turns some lights on at sunset. - I have an Ansible tool that runs on request on my Mac: it fires up and configures a remote virtual server, provisions it, loads up information, gets ssl certificates, makes a web page and starts several services. ## Context: where a tool runs, what it acts on, when it is used. A pocketknife is primarily a *hand* tool. It lives in your pocket and acts on things that you can pick up. You pick up the tool when you need to change something physical. Thirty years experience is primarily a *verbal* tool. It lives in your head, and primarily acts on things with which you can *converse*. The tool is activated (it's hardly your choice) when the context switches to something that matches experiences. Executable software tools have context, too – where they live, what they act on, what activates them. A big part of their context is how one links them together: triggers and flows. This depends on the technology you're using. My Mac has apps. I can link those apps together with Keyboard Maestro, Shortcuts, Automator, raycast, Msty – data gets in typically with an extraction from an app or a file. The end point is when one of those linking tools causes an app to take action, to put transformed information into a file, to show aggregated info in a dialog box. I'll reach out across the internet with IFTTT, Postman, curl, n8n, ansible. My servers have command line tools. Data typically gets in from files or is generated and sent to `stdout` by one of those tools. The endpoint is typically a file or an action. My services live on the internet. Data typically is sent from one to another with an https call to an API endpoint – I might use `ping` to check for a connection and `curl` (or Postman) to get from one machine to another. I've used `ssh` and `sftp` too, and webhooks to send events. I need to think carefully about credentials. When I'm bringing up new servers, I'll use \`ansible\`, CloudFlare and DigitalOcean droplets. When I want a swift environment, I'll use github. When I want a tiny scrit to run unattended, I'll use cloudflare workers, and if I want We're all using LLMs – and that's a whole other context. Replicate, OpenRouter... ## Testing and testability Tiny tools need testing, too. Handily, tiny tools often have a clear input and expected output. As short bits of code, they can be inspected, if not always easily. If their output is disposable, then perhaps it's easier to play. However, as they're often an amalgamation of existing tools, you may trigger behaviours that are well outside your aims – and of course they're running with your permissions, so they may have the rights to do bad stuff. I throw my inputs through as I'm building the tool, keep a few input / output pairs around, look out for weirdnesses if I'm using a tool after I've built it, try to build in robustness if the tool seems useful-enough to keep around. On the command line, I can use `tee` to see data at intermediary steps and some executable tools have an option to let me 'dry-run' before I run a command: `ansible` has `--check` and `git` has `--dry-run` . `xargs` has `-p` to ask my permission, and I can often chose to `echo` rather than execute a command. For local macros, and for actions which complete in an app, I might run with a very reduced and sacrificial set of data rather than taking a chance. I don't tend to automate tests for small tools – unless they're running as a permanent part of a test suite, like a [custom assert](https://www.workroom-productions.com/custom-asserts/). ## Framework I get to the right custom tool more easily if I choose to think. This framework helps me think – it's not a definitive set of patterns for tools. A whole class of tools transforms data. Let's start with those, because we recognise and can reuse their parts. - There's usually some part that **gathers or produces information**. That information has a **source**, typically a list, which may be in a file from some person, or be extracted from the file system, the running system, the network, or generated. - Then that data is **transformed** – and tools that work on **lines** are different from those that work on **columns**, because we've got a typical pattern where each line is a similar and countable item, and each column tells us how those items differ. Transformations **stack**: you might filter for the lines you want, cut out only those characteristics that matter, sort by size, and filter again to highlight some extreme. - The transformed data is **aggregated** or **compared**, taking you to some set of information that is comprehensibly small and directly relevant. And your tiny tool lets you repeat the action, so it can be tested, tuned and replayed when the initial data has changed. - These parts need to be **linked together**. That linkage may be very environment-specific and dependent on where you are in a [systems interconnection model](https://en.wikipedia.org/wiki/OSI%5Fmodel). Unix systems have the redirection operators, corporates have middleware, the internet has https and (much) more, Macs have whatever restrictive / enabling flow diagram tool the OS is currently favouring... Another class of tools takes action – again, on the file system, the system, something remote – and also perhaps on the tool and its environment. Perhaps that action is scheduled. Another class measures something – time, space (storage or memory), CPU taken, files locked, permission. If I start to recognise that something needs careful judgement or a particularly twitchy algorithm or process, then that's an indication that I probably need to split my end point: **I need *two* tools** and at least one of them might involve tools which get a person involved. ## My parts list – make your own This is rather unix-focussed. You'll make your own for the tools and technologies you have at your fingertips. #### About this list While the framework above has been kicking around for ages, I made this tools list with assistance from an LLM, with reference to the POSIX commands and GNU core utilities. Then I fiddled with it. It's a work in progress. My prompt to Claude 4 Sonnet was: \> We'll be making several table of posix commands (list attached). I want the following tables: commands that produce information (like `ls` and `ps` and `cat` and `curl`), commands that change information by filtering or rearranging rows (like `grep`, `sort`, `uniq`, `head` and `tail`), commands that change information by filtering columns (like `cut`), commands that change characters (like `tr`), commands that aggregate or combine (like `wc`, `paste`, `sum`), commands which take action (like `curl`, `mv`, `sh`), commands that set or change context (like `cd`), redirection commands (like `>`, `|` and – for me – `tee`), commands to start a human interface (like `less` and `top`), commands that persist (like `cron`) or which measure (like `time`, `du`). Commands can be in more than one table. If a command is not i any table, put it in a ‘miscellaneous’ table. \> Your tables will have columns. Each command in a table should have all columns filled in if possible (if empty, make a note below the table). Columns include: description in 10 words or less | whether the command is for files (like `ls` and `lsof`), processes (like `ps` and `lsof`), networks (like `netstat`), the system (like `ps`, `stat`, `which`) or something else (categorise if I’ve missed something), a column to indicate whether the command is (notably popular, or deprecated, or superseded by something), a column to indicate if the POSIX command is in the GNU coreutils (also attached) and whether (file, text or shell), a column to indicate the other tables a command can be found in. ### Info – gather / produce #### files and text | Command | Description | | ------- | ------------------------------ | | ls | List directory contents | | find | Return files to match criteria | | cat | Concatenate and print files | | df | Report free storage space | | du | Estimate file space usage | | file | Report type of files | | seq | Generate number sequences | | echo | Display text | #### system | Command | Description | | ------- | ---------------------------- | | ps | Report process status | | date | Report system date and time | | env | Report environment variables | #### newtwork / remote | Command | Description | | --------------------------- | ---------------------------------- | | curl | get info from an internet source | | wget | get a file from an internet source | | ping | contact a network address | | ifconfig netstat ss iproute | network information | #### Excel / sheets / notepad / VSCode | Command | Description | | ------------------------------------------------- | ---------------------------------- | | **excel**: paste csv / tsv | import rows, splitting into colums | | **copy** from somewhere, **paste** somewhere else | | ### Filter and Change #### filter / rearrange rows | Command | Description | | ----------------------------- | ------------------------------- | | grep | Search text for pattern | | sort | Sort lines of text files | | uniq | Report or filter repeated lines | | head | Copy first part of files | | tail | Copy last part of files | | join | Merge files on common field | | diff | Compare two files | | **excel**: use filters / sort | | #### filter columns | Command | Description | | -------------------------------------------- | -------------------------------------- | | cut | Cut selected fields from lines | | paste | Merge corresponding lines of files | | awk | Pattern scanning and processing | | basename dirname | Extract filename / directory from path | | **excel**: drag + drop / cut or copy + paste | | #### change characters | Command | Description | | -------------------------------- | ----------------------------- | | tr | Translate characters | | sed | Stream editor | | expand / unexpand | Convert tabs to / from spaces | | printf | Format and print text | | **excel**: global find / replace | | | **excel**: replace on import | | ### Aggregate and Combine | Command | Description | | ------------ | -------------------------- | | wc | Count lines, words, bytes | | paste | Merge lines of files | | join | Join files on common field | | sort | Sort and merge files | | sum cksum | Checksum | | test | Evaluate expressions | | excel: pivot | | ### Take action #### on files | Command | Description | | ----------------- | ------------------------------------------- | | \> \>> | write to / append to a file | | mv cp cp | Move or rename / copy / remove files | | mkdir rmdir | Make / remove directories | | chmod chown chgrp | Change file permissions / ownership / group | | ln | Link files | | touch | Change file timestamps | #### trigger and act on processes | Command | Description | | ------- | ------------------------- | | kill | Terminate processes | | sh | Shell command interpreter | | nohup | Run immune to hangups | | xargs | do a command on all items | #### remote action | Command | Description | | ------- | --------------------------------- | | curl | send info to an internet endpoint | | ping | check response | | ssh | log into a remote | ### Set or change context | Command | Description | | ------- | ---------------------------- | | cd | Change working directory | | env | Set environment for commands | ### Linking and redirecting | Command | Description | | --------- | ------------------------------ | | tee | Duplicate standard output | | \| (pipe) | pipe output to input | | \> | Output redirection - overwrite | | \> | Output redirection – append | | < | Input redirection | ### Measuring | Command | Description | | ------------- | -------------------------------------------------------- | | time | Measure command execution time | | du | Measure disk usage | | df | Measure filesystem space | | ps \+ options | various measures of processes i.e. memory / CPU / uptime | ### Scheduling / repeating | Command | Description | | ------- | --------------------------- | | cron | Schedule periodic tasks | | at | Execute commands later | | watch | execute a command regularly | ### Start a human interface | Command | Description | | ------- | -------------------------- | | less | Page through files | | more | Display files page by page | | top | Display running processes | | vi | Visual text editor | | ed | Line text editor | ### Environments: Cert work URL: https://www.workroom-productions.com/environments-cert-work/ Last updated: 2025-11-12T12:16:33.000Z If you're setting up an environment that has some sort of presence on the internet, you're going to need a certificate so that people can get to it with `https`. #### What's my use case? My play environments use VSCode in the browser. VSCode expects encrypted connections, so needs `https`. Your use case might be able to live with `http`, but I reckon that's probably only temporary. Here's a short ansible task that I built, along with LLM assistance. On the machine that has the site that needs a certificate: - It installs python libraries to do the work. - It makes a credentials file including the secret cloudflare credential, restricting its access. - It checks whether a cert exists (it doesn't want to request a new cert if it already has one) and whether it is about to expire (and gets a new cert if it expires within 30 days). - It gets the cert from LetsEncrypt, - using [certbot](https://eff-certbot.readthedocs.io/en/stable/using.html#certbot-commands) with the [dns-cloudflare plugin](https://certbot-dns-cloudflare.readthedocs.io/en/stable/) (installed earlier) - indicating the domain with `-d` - passing the cloudflare credentials as a way to identify to Certbot (and LetsEncrypt) that the entity requesting a cert is entitled to that cert. There are other ways - using `certonly` to only gets the cert. This pops it in the default `/etc/letsencrypt/live/` directory, and leaves setup to somewhere else. An `--ngnix` option would set up ssl here, but I'm leaving it until later. This may change. - agreeing to everything without pause - It gets rid of the credentials file (whether it got a cert or not) - It reloads the nginx server (so the cert gets used) - It displays a message (for log needs). #### Why use a credentials file? Q: After going to the trouble of keeping a credential secret, why write it to a file? A: Because the [certbot cloudflare plugin](https://certbot-dns-cloudflare.readthedocs.io/en/stable/) insists. #### Learn from my mistake In my initial script, I didn't delete the credentials file, leaving it visible to anyone with access to my play environments... and my environments give command-line access, so that's **anyone*. Oh dear. It **was* restricted to `root`, but with the access I give workshop participants, even that might be flaky. Thankfully, cloudflare's credentials can be limited. Still, yikes. #### How did I use LLMs and what did they get wrong My Ansible – indeed, my `yaml` – is poor, and I used Claude in various guises to get me to this point. Initially, over several LLM requests and fiddling, I got something 'working' – as in I got an Ansible-built repeatable environment that used my configuration and secrets and matched my description. It was a single awful file. I used my smarts, and more LLM requests, to actually **read* the file and split it into coherent parts. This was one part, and was aggregated from chunks from all over the monolith. When I looked at it a third time, I recognised that it was doing undesirable things. - it wrote credentials to a file, then left it in place for any fule to find. Not to **read*, though. - if a package manager on the server was still busily installing, this script flaked out when asking to install `certbot`. So it flaked out the first time, most times. It needed to retry, not barf. - it created an `ssl` directory and did other ssl setup tasks – those were duplicated later, and some indeed needed to be done later. So those tasks are deferred, but I imagine I'll come back to this. - it suggested a set of entirely spurious possible options to the `certbot` command when I wanted to pass credentials directly. Although it had used the 'right' plugins, it didn't have information to use them. I went to find the page, gave it to the LLM, and asked it to double-check my work. I also asked the LLMs to check my approaches and to explain things to me. It's particularly handy to give the right webpage / manual page as an attachment, because I can ask questions and get answers that make sense of the liberally-scattered special terms. I use [Msty](https://msty.ai) and a [jina.ai](https://jina.ai/) key to do that. ```yaml - name: Install certbot and DNS plugins apt: name: - certbot - python3-certbot-nginx - python3-certbot-dns-cloudflare state: present force_apt_get: yes environment: DEBIAN_FRONTEND: noninteractive async: 300 poll: 10 when: inventory_hostname in groups['new_droplets'] ## often needs a couple of retries as packages are locked. retries: 5 delay: 30 until: result is succeeded register: result - name: Create Cloudflare credentials file for certbot copy: content: | dns_cloudflare_api_token = {{ cloudflare_api_token }} dest: /etc/letsencrypt/cloudflare.ini mode: '0600' owner: root group: root - name: Check if Let's Encrypt certificate already exists stat: path: "/etc/letsencrypt/live/{{ group_name }}.{{ main_domain }}/fullchain.pem" register: cert_exists - name: Check certificate expiration if it exists command: openssl x509 -in "/etc/letsencrypt/live/{{ group_name }}.{{ main_domain }}/fullchain.pem" -noout -checkend 2592000 register: cert_valid failed_when: false when: cert_exists.stat.exists ## Needs to use a credentials file, as option to use token directly does not exist ## Set to always delete the credentials file, whether cert obtained or not ## puts cert into /etc/letsencrypt/live/{{ domain }}/ - block: - name: Get Let's Encrypt certificate using DNS challenge command: certbot certonly --dns-cloudflare --dns-cloudflare-credentials /etc/letsencrypt/cloudflare.ini --dns-cloudflare-propagation-seconds 60 -d {{ group_name }}.{{ main_domain }} --email {{ certbot_email }} --agree-tos --non-interactive when: not cert_exists.stat.exists or (cert_exists.stat.exists and cert_valid.rc != 0) notify: reload nginx always: - name: Remove Cloudflare credentials file file: path: /etc/letsencrypt/cloudflare.ini state: absent - name: Display certificate status debug: msg: "Certificate for {{ group_name }}.{{ main_domain }} {{ 'already exists and is valid' if (cert_exists.stat.exists and cert_valid.rc == 0) else 'was created/renewed' }}" ``` ### Workroom Playtime 038: Parts for Tools URL: https://www.workroom-productions.com/workroom-playtime-038-parts-for-tools/ Last updated: 2025-11-12T00:41:29.000Z This week's [Workroom PlayTime](https://www.workroom-productions.com/workroom-playtime/) is [Thursday 13 November at **5pm** London time](https://this-ti.me/?uts=1763053200&tz=Europe%2FLondon&name=Workroom+PlayTime+038). That's *5*pm. ***5***. I might just have Bart Knack and Huib Schoots with me... We'll gather on Zoom to work with a [short exercise about custom tools](https://www.workroom-productions.com/exercise-parts-for-tools/). In it, we'll think about and play with choosing what you'll make your custom tools from, and what you'll use those parts for. This will become part of a [tutorial at Agile Testing Days](https://agiletestingdays.com/2025/session/the-art-of-crafting-your-custom-tools/) in Potsdam in November. These exercises are for everyone, for free. [Subscribers](https://www.workroom-productions.com/#/portal/signup) get to play together every week. If you can see a **joining info** section below, then that gives you access to the Zoom etc. for the workshop. If you'd like to bring a friend, you can do that. I'll say yes until I've got too many. This is your last chance to see the **trial run of** [**Testing Transparently**](https://agiletestingdays.com/2025/session/testing-transparently/) – the keynote that [Elizabeth Zagroba](https://elizabethzagroba.com) and I are working on for ATD. We'll do it on Zoom at [9:30am (UK time) on Friday 14 November](https://this-ti.me/?uts=1763112600&tz=Europe%2FLondon&name=TestingTransparently+runthrough). That's *this* week. Ping me for an invite if you've not already heard from us. Bart Knaack and Huib Schoots and I planned to run a bigger chunk of the [Crafting Custom Tools tute](https://agiletestingdays.com/2025/session/the-art-of-crafting-your-custom-tools/). Can't guarantee it at this late stage, but **email me to find out if and when we do.** Here's a rough plan of what I plan to cover in the rest of the year: - `039` 20 November: «something fun» - `040` 27 November: from ATD! - `041` 4 December: Building exploratory interfaces (longer session) - `042` 11 December: a New Blackbox puzzle - `043` 18 December: Going meta – let's build short exercises! I don't think we'll do 25th December, or 1 January (you might have something planned, yourself). So that's a full year of Workroom PlayTime – **thank you for coming on this journey.** In 2026, we'll do more puzzles, more stuff for speakers, more thinking... and maybe fewer tools. I'll re-run some exercises, and experiment more with timings and lengths. ↓ *signed-in subscribers will see joining info and links below* ↓ _This post is for subscribers only._ ### Exercise: Parts for Tools URL: https://www.workroom-productions.com/exercise-parts-for-tools/ Last updated: 2025-11-18T23:25:46.000Z In this exercise, we'll play with tool patterns, not with tools, and see if that helps us think of opportunities. A tool, for here, is something that acts for us. A pattern is something that tells us about the parts of a tool, and how they fit together. #### Tools, and the **right* tool We use a tool because we do a better job with the tool. We may have to **make* a new tool to have the **right* tool– and making means you need to use the tools you have around you. You select tools based on what they can do, configure those tools to do what you need, combine tools so that they work together, and use the resulting 'right' tool with care for the purpose you intend. And then you pick it up because it's there, use it for the wrong thing, and stick the sharp end through a thumb. ## Exercise 0 – warm up *5 mins max* Let's talk about what we build tools from, and what we're comfortable with. Your tools may look like processes, or shortcuts, or templates – for the rest of this exercise, we'll tend towards tools that you can run on a computer. We'll talk about where our tools typically run, and what they typically operate on. See #### Examples (copy) - I have a JavaScript Bookmarklet tool that runs when I click a bookmark button. Running in my browser, it acts on the current page, and gives me a form to send the page's URL and title to a bookmarking service. info -> action - I have an AppleScript tool that runs on my Mac, on request from my dock. It works with my browser and summarises its state in an alert. info->aggregate - I have a commandline tool that runs on my Mac when I trigger it, that sees what files with particular extensions in the current directory are referenced in files in another directory. info->filter->aggregate - I have an IFTTT tool that is triggered when someone takes a particular action on my website. It reacts to a webhook and sends me an email. - I have a HomeKit tool that runs in my home cloud, and turns some lights on at sunset. - I have an Ansible tool that runs on request on my Mac: it fires up and configures a remote virtual server, provisions it, loads up information, gets ssl certificates, makes a web page and starts several services. ## Exercise 1 – play with patterns *5-10 mins* > A pattern is something that tells us about the parts of a tool, and how they fit together. Bring a pattern you know – or pick one from below. To get in touch with the pattern, think of one or more tools you've used or made, and how they fit the pattern. The examples may help (or they may not – use as you see fit). Look in the tool list for tools that fit the parts of the pattern. Using the pattern you're familiar with, think of another couple of tools using different parts. Your imagined tools do need to be purposeful, and they don't have to be test related. Then we'll exchange ideas. We'll get a feel for the ways that we can structure custom tools, and look out for ideas for tools we could use in work. ## Exercise 2 – propose new custom tools *5-10 mins* With luck, you've got some thoughts about something that a custom tool would help with. You might recognise a pattern, or you might usefully build a custom tool for yourself. Sketch out your custom tool on the Miro board. Show your custom tool to the group. Talk about what you'll use, how you'll grow it, how you know it's working. ## Tool pattern directory Here are a few rough patterns to help us think about tool types as we build. *Examples... coming.* ### tool + options A single tool that – *with the right options* – gets you all the way there. ### info → filter One tool gathers the information, the next works on it, giving you the information you need. ### info → filter → aggregate The information is gathered, transformed, and condensed. You loose the details, and get the big picture. ### info → aggregate → filter The big picture is too big – let's ignore some of it ### info → action Change files, send info, trigger web events, send info to an api ### measure Get duration, or what files are open, storage consumption, memory pressure ### compare files, or results from different sources or at different times ### monitor and act on event alert a human / capture data / take action in response to an event ### act regularly / act later capture information autonomously for later analysis ### check with user process a list with a human in the loop ### template / expand generate text on request ### generate input + call a routine generate a range of data, pass it through code, return output ## Parts Directory *Note – this is for the exercise. I'll keep a directory* [*here*](https://www.workroom-productions.com/parts-for-tools/)*, and may get rid of this list.* Here's a few lists to get started. Most of these are installed by default. ### context ***where*** will it do its work – your machine, a browser, the internet? ***what*** will it work on – files, processes, environments, downloaded data? ### Info – gather / produce #### files and text | Command | Description | | ------- | ------------------------------ | | ls | List directory contents | | find | Return files to match criteria | | cat | Concatenate and print files | | df | Report free storage space | | du | Estimate file space usage | | file | Report type of files | | seq | Generate number sequences | | echo | Display text | #### system | Command | Description | | ------- | ---------------------------- | | ps | Report process status | | date | Report system date and time | | env | Report environment variables | #### newtwork / remote | Command | Description | | --------------------------- | ---------------------------------- | | curl | get info from an internet source | | wget | get a file from an internet source | | ping | contact a network address | | ifconfig netstat ss iproute | network information | #### Excel / sheets / notepad / VSCode | excel: paste csv / tsv | import rows, splitting into colums | | copy from somewhere, paste somewhere else| | ### Filter and Change #### filter / rearrange rows | Command | Description | | ------------------------- | ------------------------------- | | grep | Search text for pattern | | sort | Sort lines of text files | | uniq | Report or filter repeated lines | | head | Copy first part of files | | tail | Copy last part of files | | join | Merge files on common field | | diff | Compare two files | | excel: use filters / sort | | #### filter columns | Command | Description | | ---------------------------------------- | -------------------------------------- | | cut | Cut selected fields from lines | | paste | Merge corresponding lines of files | | awk | Pattern scanning and processing | | basename dirname | Extract filename / directory from path | | excel: drag + drop / cut or copy + paste | | #### change characters | Command | Description | | ---------------------------- | ----------------------------- | | tr | Translate characters | | sed | Stream editor | | expand / unexpand | Convert tabs to / from spaces | | printf | Format and print text | | excel: global find / replace | | | excel: replace on import | | ### Aggregate and Combine | Command | Description | | ------------ | -------------------------- | | wc | Count lines, words, bytes | | paste | Merge lines of files | | join | Join files on common field | | sort | Sort and merge files | | sum cksum | Checksum | | test | Evaluate expressions | | excel: pivot | | ### Take action #### on files | Command | Description | | ----------------- | ------------------------------------------- | | \> \>> | write to / append to a file | | mv cp cp | Move or rename / copy / remove files | | mkdir rmdir | Make / remove directories | | chmod chown chgrp | Change file permissions / ownership / group | | ln | Link files | | touch | Change file timestamps | #### trigger and act on processes | Command | Description | | ------- | ------------------------- | | kill | Terminate processes | | sh | Shell command interpreter | | nohup | Run immune to hangups | | xargs | do a command on all items | #### remote action | `curl` | send info to an internet endpoint | | `ping` | check response | | `ssh` | log into a remote | ### Set or change context | Command | Description | | ------- | ---------------------------- | | cd | Change working directory | | env | Set environment for commands | ### Redirection commands | Command | Description | | --------- | ------------------------------ | | tee | Duplicate standard output | | \| (pipe) | pipe output to input | | \> | Output redirection - overwrite | | \> | Output redirection – append | | < | Input redirection | ### Start a human interface | Command | Description | | ------- | -------------------------- | | less | Page through files | | more | Display files page by page | | top | Display running processes | | vi | Visual text editor | | ed | Line text editor | ### Measuring | Command | Description | | ------------- | -------------------------------------------------------- | | time | Measure command execution time | | du | Measure disk usage | | df | Measure filesystem space | | ps \+ options | various measures of processes i.e. memory / CPU / uptime | ### Scheduling / repeating | Command | Description | | ------- | --------------------------- | | cron | Schedule periodic tasks | | at | Execute commands later | | watch | execute a command regularly | ### Workroom Playtime 037: xargs tool for testers URL: https://www.workroom-productions.com/workroom-playtime-037-xargs-tool-for-testers/ Last updated: 2025-11-12T00:39:24.000Z This week's [Workroom PlayTime](https://www.workroom-productions.com/workroom-playtime/) is [Thursday 6 November at **4pm** London time](https://this-ti.me/?uts=1762444800&tz=Europe%2FLondon&name=Workroom+PlayTime+037). That's *4*pm. ***4***. We'll gather on Zoom to work with [xargs for testers](https://www.workroom-productions.com/xargs-for-testers/) . `xargs` allows you to run the same command on lots of things. We'll work in pre-built environments, so you don't need to bring anything but a browser. This will become part of a [tutorial at Agile Testing Days](https://agiletestingdays.com/2025/session/the-art-of-crafting-your-custom-tools/) in Potsdam in November. These exercises are for everyone, for free. [Subscribers](https://www.workroom-productions.com/#/portal/signup) get to play together every week. If you can see a **joining info** section below, then that gives you access to the Zoom etc. for the workshop. If you'd like to bring a friend, you can do that. I'll say yes until I've got too many. We'll be pretty tool-focussed over the coming weeks as I try out elements for that [ATD tutorial](https://agiletestingdays.com/2025/session/the-art-of-crafting-your-custom-tools/). I plan to cover: - `038` 13 November: parts for tools - `039` 20 November: «something fun» - `040` 27 November: from ATD! We've set a date, so **we'd like to invite you to a trial run of** [**Testing Transparently**](https://agiletestingdays.com/2025/session/testing-transparently/) – the keynote that [Elizabeth Zagroba](https://elizabethzagroba.com) and I are working on for ATD. We'll do it on Zoom at [9:30am (UK time) on Friday 14 November](https://this-ti.me/?uts=1763112600&tz=Europe%2FLondon&name=TestingTransparently+runthrough). Ping me for an invite if you've not already heard from us. I'll run a chunk of the Crafting Custom Tools tute with Bart Knaack and Huib Schoots soon, too. **Reply to this to catch that train.** Or email me / comment here depending on how you're reading... ↓ *signed-in subscribers will see joining info and links below* ↓ _This post is for subscribers only._ ### xargs For Testers URL: https://www.workroom-productions.com/xargs-for-testers/ Last updated: 2025-11-06T16:36:53.000Z xargs to automate test tasks _This post is for subscribers only._ ### Custom Asserts – what and why URL: https://www.workroom-productions.com/custom-asserts/ Last updated: 2025-10-31T11:03:54.000Z #### Terminology Gerard Meszaros’ comprehensive 2006 [xUnit Patterns](https://dl.acm.org/doi/book/10.5555/1076526) book uses the term **custom asserts*, Chai’s docs pull them into the collective term **helpers*, [Hamcrest](https://hamcrest.org) calls them **matchers*, and I’ve often called them **custom matchers*. Here, we’ll use Meszaros’ term – it’s oldest, and the description is the deepest. In automated tests, keywords like`assert`or`expect` or `should` wrap round a comparison between some expectation, and some measurement taken during the test. There are typically `assert`s like `toBeEqualTo` and `toBe`, with decorations like `not.toBe` and complications like `toContain`. A *custom assert* is written for the system under test, and usually wraps around one or more of these simple checks. It’s an abstraction to help the tests be understood by humans. It needs an understandable name, and a place to live. ## Handy Things A custom assert is useful because: - it can give better messages - it gives a name to more-complex checks and abstracts them to a single thing - it can be reused in several tests, across different teams - it can be reused outside unit tests (sometimes) - tests using well-named custom asserts are shorter and more descriptive – and so their intent may be more obvious - using shared custom asserts allows updates in one place to be used in several, making the tests easier to evolve as the system evolves - when checks fail because of a change to a shared custom assert (and they will, which can hurt), that failing can expose misunderstandings between teams, and can expose problems with the test architecture. - they can be tested independently of the system-under-test, so can be tuned or made more reliable ## Engineering? If your test automation is littered with copy-pasted bitrot, Custom asserts are one gateway to ‘engineered’ test code because they require some engineering up-front to let testers 1) build code 2) share code 3) put their shared code into change control 4) plug into whatever enables reuse in their team. The results and potential of that engineering can be seen swiftly across a team, letting you make a small experiment into an achievable and costed improvement. For funsies, set up a single file that can be accessed by tests running on anyone's machine. If you can't do that, take a moment and think of how handy that might be. Doesn't have to be live, can be a local repo. Stuff a custom matcher in it for something that you already use in more than one place. Refactor a test or three to use it. Show a colleague. Use it on their machine. Keep going if it seems useful. Some of you may smirk. Good for *you*. Some poor sod is starting today, as an SDET in a company whose name you know, and finds that plenty of tests have all the steps, but *no* checks that any values are right. The barest of happy-path confirmatory validation (vom-firmatory?) relies on the system failing obviously, and nothing as subtle as checking that something has a value that can be predicted. ## Making custom asserts Custom asserts are typically simple, encapsulated, and have really good examples – making them good candidates for generation via LLMs. We need to build them, rather than use libraries, because they typically express something specific about the system we’re testing – but they seem to be a pain to get right because they’re also different across different unit test platforms. An LLM assist can get you off the blocks. Here are some [exercises](https://www.workroom-productions.com/custom-asserts-exercise/) to get you started. ## Some sources [Custom Asserts](https://learning.oreilly.com/library/view/xunit-test-patterns/9780131495050/ch21.html#ch21lev1sec3) in xUnit Test Patterns: Refactoring Test Code by Gerard Meszaros. Comprehensive, aspirational, reasoned – and in Java. [Mastering Assertions in Software Testing](https://blog.bluetrail.software/the-power-of-precision-mastering-assertions-in-software-testing-21ec6a5fa69e) on [Alexis Monroy](https://alexisml.medium.com/?source=post%5Fpage---byline--21ec6a5fa69e---------------------------------------)’s blog. Short, understandable, purposeful Python. ...more to come. ### Secrets at my Fingertips URL: https://www.workroom-productions.com/secrets-at-my-fingertips/ Last updated: 2025-12-16T13:13:56.000Z *Tap it, unwrap it – how I send secrets with biometrics* When I'm setting up kit, **I need secrets**. Typically keys to identify myself as someone authorised. Keys for provisioning servers, for talking to LLMs, for setting up DNS, for using `ssh`. And when I say 'I', I mean my *tools* – so those keys need to be stored somewhere, not typed in when needed. As they're secrets, they need to be stored in an encrypted way – and I don't mean the encryption that munges my laptop's whole SSD, but encryption specifically for the secrets file. So **`secrets.yml` needs a password.** We should take that password seriously, of course: should we remember it or scribble it on a hidden postit? Clearly neither. And we're *collaborative* engineers, so when we need to share that password, and those secrets, we need to share in a way that means the secret isn't generally known and access can be killed off. As it happens, I keep the password to the secrets file in a password manager, [1Password](https://1password.com). So I can already share it with care, change it in a way that I can check up on, restrict and enable as needed. It's not Sailpoint, but it works for me and my circle. Here's the enabler: **1Password has a** [**command-line interface op**](https://developer.1password.com/docs/cli/get-started/)**, and unlocks with biometric convenience.** So, in Ansible, when I want to edit or open my `secrets.yml`, I don't need to type the password, but I use this parameter: `--vault-password-file <(op item get 'Ansible secrets' --fields label=password --reveal)`. That uses `--vault-password-file` to pass the field `password` from the 1Password item `Ansible secrets` as the password for the secrets file. Which means I can put the command in something openly stored, have the kit request my identity when it first needs it, and have the decrypted secrets in memory and not in a shabby test file somewhere obscure\*. In practice: **I use the command, verify with a fingerprint, and off it goes.** 💡 ****More general use...** I can use 1password's secrets more generally by passing them on the commandline. If I pass them as a named environment variable on the same line as the command, 1password asks for my finger, puts the secret in the environment variable, and that variable evaporates once the command is sent. **Example*: If I want to pass n8n's encryption key to a docker container as I start it up (the alternative being to keep it in plaintext in the container's data volume, which seems risky), then I prepare by putting `N8N_ENCRYPTION_KEY=${N8N_ENCRYPTION_KEY}` in the `docker-compose.yaml` , and run `N8N_ENCRYPTION_KEY=$(op read "`[op://Private/n8n\_docker/encryptionKey"](about:blank)`) docker compose up -d` on the commandline – the kit asks for my finger, and the key is visible in container only as long as it lives, rather than in the data as long as any backup persists. **Note that I'm using* *`op read`* *rather than* *`op item get`* *, because it's simpler.* '\* unless, like a pillock, I do `ansible view secrets.yml` , forget, crash something, scroll up my history file and just see all the keys there. So, y'know, don't do that. ### Workroom Playtime 036: Custom Asserts URL: https://www.workroom-productions.com/workroom-playtime-036-custom-asserts/ Last updated: 2025-10-29T08:58:45.000Z This week's [Workroom PlayTime](https://www.workroom-productions.com/workroom-playtime/) is [Thursday 30 October at **4pm** London time](https://this-ti.me/?uts=1761840000&tz=Europe%2FLondon&name=Workroom+PlayTime+036). That's *4*pm. ***4***. We'll gather on Zoom to work with [Custom Asserts](https://www.workroom-productions.com/custom-asserts-exercise/) – which bring reusable deep checks to your test automation. We'll work in pre-built environments, on working software with a test harness, so you don't need to bring anything but a browser. This will become part of a [tutorial at Agile Testing Days](https://agiletestingdays.com/2025/session/the-art-of-crafting-your-custom-tools/) in Potsdam in November. These exercises are for everyone, for free. [Subscribers](https://www.workroom-productions.com/#/portal/signup) get to play together every week. If you can see a **joining info** section below, then that gives you access to the Zoom etc. for the workshop. If you'd like to bring a friend, you can do that. I'll say yes until I've got too many. We'll be pretty tool-focussed over the coming weeks as I try out elements for that [ATD tutorial](https://agiletestingdays.com/2025/session/the-art-of-crafting-your-custom-tools/). I plan to cover: - `037` 6 November: `xargs` for testers - `038` 13 November: parts for tools - `039` 20 November: «something fun» - `040` 27 November: from ATD! We've set a date, so **we'd like to invite you to a trial run of** [**Testing Transparently**](https://agiletestingdays.com/2025/session/testing-transparently/) – the keynote that [Elizabeth Zagroba](https://elizabethzagroba.com) and I are working on for ATD. We'll do it on Zoom at [9:30am (UK time) on Friday 14 November](https://this-ti.me/?uts=1763112600&tz=Europe%2FLondon&name=TestingTransparently+runthrough). Ping me for an invite if you've not already heard from us. I'll run a chunk of the Crafting Custom Tools tute with Bart Knaack and Huib Schoots soon, too. **Reply to this to catch that train.** Thank you for reading! Cheers – James ↓ *signed-in subscribers will see joining info and links below* ↓ _This post is for subscribers only._ ### Exercise: Custom Asserts URL: https://www.workroom-productions.com/custom-asserts-exercise/ Last updated: 2025-11-06T23:44:33.000Z *For these exercises, we'll use Github Codespaces. You will need a github account. You will need to run this on your own account, which has ?60 hours a month free codespaces.* In this exercise. we'll use LLMs to help us build custom asserts. You'll use your own access to an LLM, and we'll work in environments with a JavaScript test harness, a Python test harness, and code which could do with a custom assert or two. We'll look at the tiny systems, identify what would help, ask LLMs to build working asserts, see what they give back, and compare notes. Here's more on [custom asserts](https://www.workroom-productions.com/custom-asserts/). ## Environment I've built something that has all the code, test harnesses and plumbing for you to try this in JavaScript and Python. You'll access your own copy of that environment via VisualStudio Code in your browser. ## Exercise 0 – your own environment Go to the [Custom Asserts repo](https://github.com/workroomprds/WorkroomPlayTime036%5FCustomAsserts) , and you should see a `readme` page – click on the "open in Github Codespaces" button, wait and respond reasonably to Github's questions. Further instructions are in the readme – once you're in, you'll need to run something on the commandline to install all the infrastructure dependencies, and more to run tests. You can browse the software under test while the dependencies are being installed. The environment has two exercises, each in both Javascript and in Python. I'm *still* building the examples, so if you set up early, you won't have them. ### Infrastructure The tests for my JavaScript examples here use [Chai](https://www.chaijs.com/) (assertion library) and [Mocha](https://mochajs.org/) (unit test framework). My Python examples use bare Python and [pyTest](https://pytest.org/) (unit test framework). You'll need to go to your own LLM to ask for code – you'll find examples below. #### Use your own LLM? You don't **have* an LLM – but you may have access to one. If not, you may use the 'free' version of [ChatGPT](https://chatgpt.com) or use a burnable email for [Claude](https://claude.ai/), or you can use copilot in the sidebar of the browser-based VSCode window. #### Python and assertion libraries Assertion libraries seem not as prevalent in the Python community as [Chai](https://github.com/chaijs/chai) and others are in the JS community. Perhaps that is because it is simple and common to build custom asserts directly in Python. Pytest leans heavily on Python’s own assert and has a hook `pytest_assertrepr_compare` for detailed messages. Python does have assertion libraries – [pyHamcrest](https://github.com/hamcrest/PyHamcrest) being the standout example. ## Pick a codebase, get the tests running Different tech has different tests. I've tried to make this exercise work in Python and in Javascript / Node. Both are readable: you might want to work in your less-familiar technology. Or you might want to compare. You'll find Python in `./python_project` and JavaScript / Node in `./javascript_project`. I've tried to give the files and functions similar names. ## Exercise 1 – get the tests running Get the tests running. Have a look at the `calculator` code and tests to see how they work. ## Exercise 2 – checking XML tags We'll build a custom assert to check XML tags. #### Context – `MT103` \-> `PACS.008` An `MT103` is a Swift transaction that acts as a proof of payment. It has recently been superseded by `PACS.008` which is XML-based. As an intermediary step in preparation for the change, some organisations reformatted the old record into XML, passing that through their live systems. Shared test asserts to do specific XML verification are useful when coding that change. - `MT103_gen.js` and `MT103_gen.py` have a function `generateMT103` which returns a simple XML example record. - `./test/test_MT103_gen.js` and `./tests/test_MT103_gen.py` both test that function. We're going to use this as a dummy to build a custom assert. We want the custom assert to check that all the XML tags are present, and that there are no extras. For the sake of plumbing build the custom assert in a file called `customAsserts.js`/`.py`, and call the method `containXmlTags`. You'll need to register it before using it. Using the names above, you could ```javascript // Register the custom assertion const containXmlTags = require('./customAsserts'); chai.use(containXmlTags); ``` ```python3 from customAsserts import containXmlTags ``` And, of course, you'll need to add the test... Once you've added it, you can test it by changing the generation code `MT103_gen`. We'll compare and contrast our experiences. ### Examples of generated test code *My approach – for greatest transparency, I'm copying the code and chucking it straight into Claude 4.* > Given the following code, and the following tests, make me a shareable custom assert called `containXMLtags` in a file `customAsserts.js`. The custom assert should check that all tags for this example MT103 are present, and none are extra. > \[I pasted the code in here\] *which returned:* Sure, here's the `customAsserts.js` file with the `containXMLtags` custom assert: ```javascript const { DOMParser } = require('@xmldom/xmldom'); const containXMLtags = (actual, expected) => { const parser = new DOMParser(); const doc = parser.parseFromString(actual, 'text/xml'); // Check for parse errors const parseError = doc.getElementsByTagName('parsererror'); if (parseError.length > 0) { return { message: () => 'The provided XML string is not valid.', pass: false, }; } // Check root element exists if (!doc.documentElement) { return { message: () => 'The provided XML string does not have a root element.', pass: false, }; } // Check for expected tags const expectedTags = [ 'MessageType', 'SenderBIC', 'ReceiverBIC', 'TransactionReference', 'ValueDate', 'Amount', 'Currency', 'BeneficiaryName', 'BeneficiaryAccount', 'BeneficiaryAddress', 'RemittanceInformation', ]; for (const tag of expectedTags) { if (!doc.getElementsByTagName(tag).length) { return { message: () => `The expected tag '${tag}' is not present in the XML.`, pass: false, }; } } // Check for extra tags const extraTags = Array.from(doc.getElementsByTagName('*')) .map((node) => node.nodeName) .filter((tag) => !expectedTags.includes(tag)); if (extraTags.length > 0) { return { message: () => `The following unexpected tags were found in the XML: ${extraTags.join(', ')}`, pass: false, }; } return { message: () => 'The provided XML string contains all expected tags and no extra tags.', pass: true, }; }; module.exports = { containXMLtags }; ``` To use this custom assert in your tests, you can import it from the `customAsserts.js` file: ```javascript const { expect } = require('chai'); const { containXMLtags } = require('./customAsserts'); describe('MT103_gen', () => { let generateMT103; before(() => { try { ({ generateMT103 } = require('../MT103_gen')); } catch (err) { generateMT103 = null; } }); it('should generate valid XML', () => { const result = generateMT103(); expect(result).to.satisfy(containXMLtags); }); it('XML should be labelled a SwiftPayment record', () => { const result = generateMT103(); expect(result).to.satisfy(containXMLtags); }); }); ``` This does both more (more checks) and less (it's not generic) than we'd asked for. I might re-prompt, asking for the test to contain the list of tags. *Note: these are examples – you'll get different stuff, and it might be better. I've not yet used* this *example.* ## Exercise 2 – checking a distribution We'll build a custom assert to check that a certain type of record is in the 'right' proportion to other records. #### Context: a simulation In the simulation, items need to be common, rare, and very rare. There is lots of code to make the items; we need a check for all the different options to let us know if some items are too rare, or not as rare as they should be. - `candidate_gen` generates candidates - `test_candidate_gen` does very little We're going to use this as a dummy to build a custom assert. We want the custom assert to check that 8-12% of the candidates are of type `high`, 0.5 to 1.5 % are `mighty` and the rest (and no more than 91%) are `base`. Again, you'll need to register your custom assert, and use it in a test. For the sake of plumbing build the custom assert in a file called `customAsserts.js`/`.py`, and call the method `checkDistribution`. Adjust the candidate generation to test it – note that the dummy *should* fail. It has too few records to have a `mighty`. So you'll need to make different distributions to properly test the assert. Tested *test* code? Whatever next!? ### Examples of generated test code > I want to build a custom assert to check candidate proportions. The assert should take parameters. In this case, 8-12% of the candidates should be of type high, 0.5 to 1.5 % are mighty and the rest (and no more than 91%) are base. Example candidate generation and existing test code follows. The custom assert should be called \`checkDistribution\` and will be used in a file called `customAsserts.py` *we'll build examples in the workshop.* ### Load Time in One Line URL: https://www.workroom-productions.com/load-time-in-one-line/ Last updated: 2025-10-23T12:53:04.000Z A one-liner to get the load time for a page might be... `curl -w "%{time_total}\n" -o /dev/null -s workroom-productions.com` Unpack this into four parts – the first to produce information, two more to *stop* information, and the web address: 1. `-w "%{time_total}\n"`: This uses `-w` (or `--write-out` ) to pick out the total time taken to complete the request from the information about the transfer. 2. `-o /dev/null`: This option throws away what `curl` has retrieved by redirecting the output to the bottomless pit `/dev/null` . 3. `-s` : Silent mode, which stops the progress meter from being output to the terminal. 4. From here, of course. I've not specified page, nor scheme( i.e. `https` ) This one-liner needs you to know that `curl` captures this information. To my knowledge it's the only common tool which does. But `curl` isn't the only way to get a page – `wget` can, and `lynx` can, and you can use `time` to measure how long it takes. Indeed you could use `time` to measure how long the `curl` takes.... Here's a useful [stackoverflow answer](https://stackoverflow.com/questions/12714584/wget-time-measurement#19113627) ### Workroom Playtime 035: build 3 throwaway tools in 15 minutes URL: https://www.workroom-productions.com/workroom-playtime-035-build-3-throwaway-tools-in-15-minutes/ Last updated: 2025-10-21T15:40:46.000Z (it's for the ***23rd*** – I know what it says in the email summary... I messed up) This week's [Workroom PlayTime](https://www.workroom-productions.com/workroom-playtime/) is [Thursday 23 October at **3pm** London time](https://this-ti.me/?uts=1761228000&tz=Europe%2FLondon&name=Workroom+PlayTime+035). We'll gather on Zoom to build [3 throwaway tools](https://www.workroom-productions.com/exercise-build-throwaway-tools/). This will become part of a [tutorial at Agile Testing Days](https://agiletestingdays.com/2025/session/the-art-of-crafting-your-custom-tools/) in Potsdam in November. These exercises are for everyone, for free. [Subscribers](https://www.workroom-productions.com/#/portal/signup) get to play together every week. If you can see a **joining info** section below, then that gives you access to the Zoom etc. for the workshop. If you'd like to bring a friend, you can do that. I'll say yes until I've got too many. We'll be pretty tool-focussed over the coming weeks as I try out elements for that [ATD tutorial](https://agiletestingdays.com/2025/session/the-art-of-crafting-your-custom-tools/). I plan to cover: - `036` 30 October: bespoke matchers via an LLM - `037` 6 November: `xargs` for testers - `038` 13 November: parts for tools - `039` 20 November: «something fun» I'll also run a couple of longer things for subscribers; a chunk of the Crafting Custom Tools tute with Bart Knaack and Huib Schoots, and a rehearsal of the Testing Transparently keynote with Elizabeth Zagroba. **Want to know more? Reply to this.** Thank you for reading! Cheers – James ↓ *signed-in subscribers will see joining info and links below* ↓ _This post is for subscribers only._ ### Exercise: Build Throwaway Tools URL: https://www.workroom-productions.com/exercise-build-throwaway-tools/ Last updated: 2025-10-23T14:03:26.000Z *10-15 mins building and talking, 5 mins talking about building.* Build env if you need one: Ephemeral tools are built and thrown away. We can of course build them and keep them and share them – and then they're *no longer ephemeral*. So here we'll practice building tools with very little investment. We're going to build tools to suit several briefs. You'll use your knowledge, and work together, and at the end you'll have a working rubbish tool. I've got a working version of each if you want hints. Then we'll think about how we built, and how we might build tools that we actually need. ## Exercise 1 – page load time *We want a one-line command-line tool to return page load times, so that we can catch it as a repeatable metric.* **Working example:** [Load Time in One Line](https://www.workroom-productions.com/load-time-in-one-line/), which builds on a big tool, and removes all but the salient data. ## Exercise 2 – check dependencies *We want a tool which, given a list of strings, checks to see which are used in nearby files, and which strings are unused.* Examples at . `git clone ` **Working example:** [checkrefs.sh](https://www.workroom-productions.com/tiny-tool-checkrefs-sh/), which is a small and inspectable shell script that opens the door to a command-line one-liner. ## Exercise 3 – recent events *We want to see the most-recent entries in the most-recent logs.* **Working example:** [recent events](https://www.workroom-productions.com/recent-events/) – this command-line one-liner uses `xargs` to run commands on each row of an output. ## Exercise 4 – spreadsheet *We want to discover what characters can't easily be pasted in google sheets. Heaven knows why...* **Working example:** The linked sheet is an example of 5000 checks. - A is a list of numbers - B is a character representing the number - C is a copy/paste of B - D is a check of B vs C, and has a filter. [CopyPaste Gamut![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/icon/spreadsheets_2023q4.ico)Google Docs![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/thumbnail/AHkbwyJNtp14f8XbhwJ4z079j_xvLVjDZUphil0j_cEP7HFvdkYyWF1IGVTjQWtnbtqmGjYFplXXcAdTWRSla9tHseEdJaddnJ-Tb4AKGONnPxr2GyLjOEvf-w1200-h630-p)](https://docs.google.com/spreadsheets/d/1fd7w255C3KfMKt8TC5aCktvO7VopJ3Pg16MTra2%5FPMk/edit?gid=0#gid=0) ## Exercise 5 – coming ## Wrap up *5 mins* Let's share how we built. #### How I build tiny tools If I'm thinking about data transformation, I'll often have a vague idea of input and output. If I'm thinking about repeated actions, I'll often have an idea about what to change / iterate, and what to keep the same. I'll know whether I want that for efficiency, for accuracy, or to do something I can't otherwise do. I've got a basic idea of command-line redirection, scripts, spreadsheet coding, scripting, and my current set of tools. This helps me to see opportunities for tooling. I've used LLMs to get me to the right place with some of these, but those interactions often give too much, or too little. Sometimes, I ask LLMs about tool **capability*. I know I can (generally) handle getting information into or out of a tool – but perhaps the capability I need is lesser-known. I'll use stackoverflow to see if anyone else has a similar problem – and the range of solutions may be more interesting than the solutions themselves. I'll want to consider quality by sanity checks, by inspection, by trying in the small, by trying edge cases, by asking a colleague or an LLM to parse. But my preferred approach to tiny tooling is to have something where problems are obvious without being painful, and run the tool several times looking at what it does. I rarely go test-first, but I often go example-first. Tools should have a purpose – let's share how we recognise our needs, and how we reckon that a tool might help. More on [Ephemeral Tools](https://www.workroom-productions.com/ephemeral-tooling/), and the page of [tiny tools](https://www.workroom-productions.com/tag/tiny-tool/). ### Recent Events URL: https://www.workroom-productions.com/recent-events/ Last updated: 2025-10-21T15:34:38.000Z I wanted to see the most-recent entries of the most-recent logs. `ls -1td /var/log/*.log | head -3 | xargs -I {} sh -c 'tail -10 "{}"'` This uses: - `ls` to make a list of files (with paths) sorted by recency - `head` to clip the top 3 - `xargs` to run those three into `tail` to pick up the bottom 10 lines. Wrinkle 1 – listing with the paths means `xargs` doesn't need to be passes the path of the files. Wrinkle 2 – `{}` is something uncommon to use as a pattern – no more. But it's a common pair to use, when finding things to sub for in commandline things. Presumably no use in javascript code... This alternative separates the logs – easier to use. `ls -1td /var/log/*.log | head -3 | xargs -I {} sh -c 'echo "=== {} ==="; tail -10 "{}"'` ### Tiny Tool: file types in a directory URL: https://www.workroom-productions.com/file-types-in-a-directory/ Last updated: 2025-10-21T15:21:49.000Z I needed to understand directories with far too many files – first stop, were they all equivalent types? I asked an LLM for help – after an iteration or two, it came up with this, which I've set up with `/var/log` to work as an example. Change `/var/log` to `.` to do the local directory, and you won't need to `sudo`. `sudo find /var/log -type f -print0 | xargs -0 file --mime-type | cut -d: -f2- | tr -d ' ' | sort | uniq -c | sort -nr` What's it doing? - `find` all files, and output them as a null-delimited list - use `file` on each to pick out the mime type - `cut` the filetype and and `tr`im spaces - use `sort` and `uniq` to get a `-c`ount of how many of each type - sort that list putting biggest first. ### Workroom Playtime 034: `jq` for testers URL: https://www.workroom-productions.com/workroom-playtime-034-jq-for-testers/ Last updated: 2025-10-14T23:03:21.000Z This week's [Workroom PlayTime](https://www.workroom-productions.com/workroom-playtime/) is [Thursday 16 October at **3pm** London time](https://this-ti.me/?uts=1760623200&tz=Europe%2FLondon&name=Workroom+PlayTime+034). We'll gather on Zoom to play with [jq for testers](https://www.workroom-productions.com/exercise-jq-for-testers/). This is another in the series of [UNIX tools for testers](https://www.workroom-productions.com/unix-tools-for-testers/), and will become part of a [tutorial at Agile Testing Days](https://agiletestingdays.com/2025/session/the-art-of-crafting-your-custom-tools/) in Potsdam in November. These exercises are for everyone, for free. [Subscribers](https://www.workroom-productions.com/#/portal/signup) get to play together every week. If you can see a **joining info** section below, then that gives you access to the Zoom etc. for the workshop. If you'd like to bring a friend, you can do that. I'll say yes until I've got too many. We'll be pretty tool-focussed over the coming weeks as I try out elements for that [ATD tutorial](https://agiletestingdays.com/2025/session/the-art-of-crafting-your-custom-tools/). I plan to cover: - `034` 16 October: `jq` for testers - `035` 23 October: build 3 throwaway tools in 15 minutes - `036` 30 October: bespoke matchers via an LLM - `037` 6 November: `xargs` for testers I'll also run a couple of longer things for subscribers; a chunk of the Crafting Custom Tools tute with Bart Knaack and Huib Schoots, and a rehearsal of the Testing Transparently keynote with Elizabeth Zagroba. **Want to know more? Reply to this.** Thank you for reading! Cheers – James ↓ *signed-in subscribers will see joining info and links below* ↓ _This post is for subscribers only._ ### Exercise: jq for testers URL: https://www.workroom-productions.com/exercise-jq-for-testers/ Last updated: 2025-10-16T13:59:10.000Z A set of testing-relevant exercises to help testers use, and see the use of, `jq` – a tool to filter and manipulate json. #### Notes on `jq` **(maybe on a page)* JSON is a text-based record format as a hierarchy – it's not rows and columns. A JSON record is typically an object `{}` or a list `[]`, and everything inside is a `key: value` – and those values can themselves be objects `{}` (named parts in any order) or lists `[]` (ordered rows), giving the hierarchy. The basic unit is `key : value` , one per line. `jq` lets you filter based on name (`.foo` shows you a top level value thta has the key `foo`), lets you drill down in to a hierarchy (`.foo.bar` shows you the value for the property `bar` under `foo`), and lets you iterate or access specifics in a list with `[` *`n`*`]` (`.[].foo` will, if you've given `jq` a list, give you the value of the property `foo` on every item on the list, while `.[0].foo` will give you the first. `jq` stacks commands – you'll send information from left to right with `|`, and you'll often send to one of `jq`'s collection commands e.g. `.[].foo | unique | sort`, or to a selection command `.foo[] | select(.bar == "moon")` Ultimately, it's a powerful filtering language for extracting, exploring and creating data. I'll not be rewriting the [manual](https://jqlang.github.io/jq/manual/), and the [cookbook](https://github.com/jqlang/jq/wiki/Cookbook) shows the complex tasks the tool can manage. ****Reading a query:** - if you see a `jq` with information in all the `[]`, it’s like dot notation – it’s effectively an address. - If a `[]` is empty, it’s iterating, and if you see a `|` it’s filtering or aggregating, sometimes to another filter, sometimes to a keyword like `sort` or `select(` . - If you see a `..` it’s traversing the whole tree, and a `?` means it’s finding stuff, or checking stuff and handling errors. - If you see `.` then that’s picking out everything, and `.foo` picks out the thing called `foo` at whatever level the `jq` reader is at. `,` separates and aggregates queries. - If there’s a `[` and `]` around something that produces information, then that information is being wrapped up as an array. If it’s wrapped in `{` and `}` you’re getting an object. For the exercises, we'll use online `jq` tools. As with regex, I reckon that we don't need to be expert with `jq`, if we have tools to help build and explore queries. All the following tools take in json and a query, and show you the results in real time. - [JQ / Playground](https://jq.port.io) is a page that lets you ask for `jq` queries in plain language, and to see their effects in real time. - [JQ Play Offline](https://jiehong.gitlab.io/jq%5Foffline/?query=.) is a page that shows you the results of queries more readably, and lets you work with command-line options - [The JQ Playground](https://play.jqlang.org) lets you get your json from an http endpoint – and is from the makers of `jq`. It runs slightly less swiftly. `jq` is primarily a command line tool: These pages are handy for building or testing a query, and for learning through playing, but you'll want the command line version to do work. The tools abstract away the command-line syntax, which goes `jq -c '.fruit.color,.fruit.price' fruit.json` – wrapping the command in `''`, acting on a file, and giving an option. However, `jq` is not always available, because it's not one of the core xNIX toolset. Its official page is [jqlang.org](https://jqlang.org), and ## Exercise 1 – exploring data with `jq` *10 minutes* We'll use the data below, which is a generated JSON output from an imaginary system that logs events to an `events` object, and logs traces to a `traces` object. - use `keys` to show the top-level keys - use `traces[]` to show everything in the traces array - use `.traces[].traceId` to show the traceIDs. - Aside: Compare with `.traces[] .traceId` and `.traces[] | traceId` - Aside: Compare with `.traces | map(.traceId) | unique` and `[ .traces[] | .traceId ] | unique` - use `events[]` to show everything in the events array - use `events[0]` to show the first event in the events array - use `.events | map(.traceId) | unique` to show all the unique traceIDs in the events array. - use `.traces | map({(.traceId): [.spans[].name]}) | add` to show all event names for each trace. - use `.events[] | select(.type == "mobile_request") | .data.device` to see all devices for mobile requests. This will, I hope, let you be familiar with accessing by name, with iterating and picking – and familiar with the data itself. ## Exercise 2 – building queries *10 mins* Imagine something you'd like to know as a tester from this data. Use an LLM of your choice, or [JQ / Playground](https://jq.port.io) above, to find a query. Try it out. Parse the query you got. Share with the rest of the group – they may try it out, parse it themselves, check that the data is right, or see if they can get the same data in a different way. ## Exercise 3 – Generating data with `jq` Wrap a `jq` command in `[]` or `{}` to generate an array or object. ### Sample Data ```json { "traces": [ { "traceId": "123456789", "duration": 2000, "spans": [ { "id": "span1", "name": "web_request", "startTime": 1634491200000, "endTime": 1634491201000, "tags": { "http.method": "GET", "http.url": "/api/users" } }, { "id": "span2", "name": "database_query", "startTime": 1634491201000, "endTime": 1634491201500, "tags": { "db.query": "SELECT * FROM users WHERE id = 1" } }, { "id": "span3", "name": "cache_lookup", "startTime": 1634491201500, "endTime": 1634491201800, "tags": { "cache.key": "user_profile_1" } } ] }, { "traceId": "987654321", "duration": 3000, "spans": [ { "id": "span1", "name": "web_request", "startTime": 1634491202000, "endTime": 1634491203000, "tags": { "http.method": "POST", "http.url": "/api/payments" } }, { "id": "span2", "name": "payment_processing", "startTime": 1634491203000, "endTime": 1634491204000, "tags": { "payment.amount": 100.00, "payment.method": "credit_card" } }, { "id": "span3", "name": "fraud_check", "startTime": 1634491204000, "endTime": 1634491204500, "tags": { "fraud.score": 10 } } ] }, { "traceId": "456789123", "duration": 1500, "spans": [ { "id": "span1", "name": "mobile_request", "startTime": 1634491205000, "endTime": 1634491205500, "tags": { "http.method": "GET", "http.url": "/api/products" } }, { "id": "span2", "name": "product_catalog_lookup", "startTime": 1634491205500, "endTime": 1634491206000, "tags": { "product.category": "electronics" } } ] }, { "traceId": "321654987", "duration": 2500, "spans": [ { "id": "span1", "name": "web_request", "startTime": 1634491207000, "endTime": 1634491208000, "tags": { "http.method": "POST", "http.url": "/api/orders" } }, { "id": "span2", "name": "inventory_check", "startTime": 1634491208000, "endTime": 1634491208500, "tags": { "product.id": "prod123", "quantity": 5 } }, { "id": "span3", "name": "order_processing", "startTime": 1634491208500, "endTime": 1634491209500, "tags": { "order.id": "order456", "order.total": 50.00 } } ] }, { "traceId": "159753", "duration": 1000, "spans": [ { "id": "span1", "name": "mobile_request", "startTime": 1634491210000, "endTime": 1634491210500, "tags": { "http.method": "GET", "http.url": "/api/notifications" } }, { "id": "span2", "name": "notification_fetch", "startTime": 1634491210500, "endTime": 1634491211000, "tags": { "notification.type": "promotion", "notification.message": "25% off sale!" } } ] }, { "traceId": "753159", "duration": 1800, "spans": [ { "id": "span1", "name": "web_request", "startTime": 1634491212000, "endTime": 1634491212500, "tags": { "http.method": "GET", "http.url": "/api/profile" } }, { "id": "span2", "name": "user_profile_fetch", "startTime": 1634491212500, "endTime": 1634491213000, "tags": { "user.id": "user123", "user.email": "user@example.com" } }, { "id": "span3", "name": "settings_fetch", "startTime": 1634491213000, "endTime": 1634491213800, "tags": { "settings.theme": "dark", "settings.language": "en" } } ] }, { "traceId": "951357", "duration": 2200, "spans": [ { "id": "span1", "name": "mobile_request", "startTime": 1634491215000, "endTime": 1634491215500, "tags": { "http.method": "POST", "http.url": "/api/feedback" } }, { "id": "span2", "name": "feedback_submit", "startTime": 1634491215500, "endTime": 1634491216000, "tags": { "feedback.rating": 4, "feedback.comment": "Great app, but could use more features." } }, { "id": "span3", "name": "support_ticket_create", "startTime": 1634491216000, "endTime": 1634491217200, "tags": { "ticket.id": "ticket123", "ticket.priority": "medium" } } ] }, { "traceId": "357951", "duration": 1700, "spans": [ { "id": "span1", "name": "web_request", "startTime": 1634491218000, "endTime": 1634491218500, "tags": { "http.method": "GET", "http.url": "/api/analytics" } }, { "id": "span2", "name": "data_aggregation", "startTime": 1634491218500, "endTime": 1634491219000, "tags": { "data.source": "user_activity", "data.timeframe": "last_7_days" } }, { "id": "span3", "name": "report_generation", "startTime": 1634491219000, "endTime": 1634491219700, "tags": { "report.id": "weekly_analytics", "report.format": "pdf" } } ] }, { "traceId": "753951", "duration": 2800, "spans": [ { "id": "span1", "name": "mobile_request", "startTime": 1634491220000, "endTime": 1634491220500, "tags": { "http.method": "PUT", "http.url": "/api/settings" } }, { "id": "span2", "name": "user_settings_update", "startTime": 1634491220500, "endTime": 1634491221000, "tags": { "setting.language": "es", "setting.notification_preferences": "email" } }, { "id": "span3", "name": "email_notification", "startTime": 1634491221000, "endTime": 1634491222800, "tags": { "email.recipient": "user@example.com", "email.subject": "Your settings have been updated" } } ] } ], "events": [ { "timestamp": 1634491200000, "type": "user_login", "data": { "userId": "user123", "ipAddress": "192.168.1.100" }, "traceId": "123456789" }, { "timestamp": 1634491201000, "type": "database_query", "data": { "query": "SELECT * FROM users WHERE id = 1" }, "traceId": "123456789" }, { "timestamp": 1634491201800, "type": "cache_hit", "data": { "cacheKey": "user_profile_1" }, "traceId": "123456789" }, { "timestamp": 1634491202000, "type": "payment_initiated", "data": { "paymentId": "payment123", "amount": 100.00 }, "traceId": "987654321" }, { "timestamp": 1634491204000, "type": "fraud_check_complete", "data": { "fraudScore": 10 }, "traceId": "987654321" }, { "timestamp": 1634491205000, "type": "mobile_request", "data": { "userId": "user456", "device": "iPhone" }, "traceId": "456789123" }, { "timestamp": 1634491205500, "type": "product_catalog_lookup", "data": { "productCategory": "electronics" }, "traceId": "456789123" }, { "timestamp": 1634491207000, "type": "order_placed", "data": { "orderId": "order456", "totalAmount": 50.00 }, "traceId": "321654987" }, { "timestamp": 1634491208500, "type": "inventory_check", "data": { "productId": "prod123", "quantity": 5 }, "traceId": "321654987" }, { "timestamp": 1634491210000, "type": "mobile_request", "data": { "userId": "user789", "device": "Android" }, "traceId": "159753" }, { "timestamp": 1634491210500, "type": "notification_fetched", "data": { "notificationType": "promotion", "notificationMessage": "25% off sale!" }, "traceId": "159753" }, { "timestamp": 1634491212000, "type": "user_profile_viewed", "data": { "userId": "user123", "userEmail": "user@example.com" }, "traceId": "753159" }, { "timestamp": 1634491213000, "type": "user_settings_fetched", "data": { "theme": "dark", "language": "en" }, "traceId": "753159" }, { "timestamp": 1634491215000, "type": "mobile_request", "data": { "userId": "user456", "device": "Android" }, "traceId": "951357" }, { "timestamp": 1634491215500, "type": "feedback_submitted", "data": { "rating": 4, "comment": "Great app, but could use more features." }, "traceId": "951357" }, { "timestamp": 1634491216000, "type": "support_ticket_created", "data": { "ticketId": "ticket123", "priority": "medium" }, "traceId": "951357" }, { "timestamp": 1634491218000, "type": "web_request", "data": { "userId": "user789", "ipAddress": "192.168.1.101" }, "traceId": "357951" }, { "timestamp": 1634491218500, "type": "data_aggregation", "data": { "dataSource": "user_activity", "timeframe": "last_7_days" }, "traceId": "357951" }, { "timestamp": 1634491219000, "type": "report_generated", "data": { "reportId": "weekly_analytics", "reportFormat": "pdf" }, "traceId": "357951" }, { "timestamp": 1634491220000, "type": "mobile_request", "data": { "userId": "user123", "device": "Android" }, "traceId": "753951" }, { "timestamp": 1634491220500, "type": "user_settings_updated", "data": { "language": "es", "notificationPreferences": "email" }, "traceId": "753951" }, { "timestamp": 1634491221000, "type": "email_notification_sent", "data": { "recipient": "user@example.com", "subject": "Your settings have been updated" }, "traceId": "753951" } ] } ``` ## ### Workroom Playtime 033: Exploratory Interfaces 2 URL: https://www.workroom-productions.com/workroom-playtime-033/ Last updated: 2025-10-07T22:33:32.000Z This week's [Workroom PlayTime](https://www.workroom-productions.com/workroom-playtime/) is [Thursday 9 October at **3pm** London time](https://this-ti.me/?uts=1760018400&tz=Europe%2FLondon&name=Workroom+PlayTime+033). We'll gather on Zoom to play with [Exploratory Interfaces alongside unit tests](https://www.workroom-productions.com/exploratory-interfaces-alongside-unit-tests/). You don't need to have done [Exploratory Interfaces 1](https://www.workroom-productions.com/playtime-002-exploratory-interfaces-i/). I've written something if you want to read my evolving thoughts around [Exploratory Interfaces](https://www.workroom-productions.com/exploratory-interfaces/). These exercises are for everyone, for free. Paying [Subscribers](https://www.workroom-productions.com/#/portal/signup) get to play together every week. If you can see a **joining info** section below, then that gives you access to the Zoom etc. for the workshop. If you'd like to bring a friend, you can do that. I'll say yes until I've got too many. I'm prepping for an [ATD tutorial](https://agiletestingdays.com/2025/session/the-art-of-crafting-your-custom-tools/) at the moment, and we'll be pretty tool-focussed over the next four weeks as I try out elements. I plan to cover: - `033` 9 October: Exploratory Interfaces 2 - `034` 16 October: `jq` for testers - `035` 23 October: build 3 throwaway tools in 15 minutes - `036` 30 October: bespoke matchers Thank you for reading! Cheers – James ↓ *signed-in subscribers will see joining info and links below* ↓ _This post is for subscribers only._ ### Exploratory Interfaces alongside unit tests URL: https://www.workroom-productions.com/exploratory-interfaces-alongside-unit-tests/ Last updated: 2025-10-09T13:58:48.000Z In this, we'll take some example code – something small-enough that it has unit tests rather than integration tests, and see what it gives us to play with . Then we'll play with it. We'll take [CadenceKeeper](https://exercises.workroomprds.com/timeSince/) as our subject – but we're not going to play with it directly very much. Instead, we'll play with [code](https://exercises.workroomprds.com/timeSince/Timer%5Fv2.js) : [Unit tests](https://exercises.workroomprds.com/timeSince/jlChecks.html?spec=Timer%20v2) : [config.js](https://exercises.workroomprds.com/timeSince/config.js) : [Exploratory Interface](https://exercises.workroomprds.com/timeSince/timer%5Fexplorer.html) ## Exercise 1: look at the code *5 mins* Take a look at the [code](https://exercises.workroomprds.com/timeSince/Timer%5Fv2.js) – identify its public interface (the ways that other code can interact with it – and the information that other code needs to provide and can expect). Have a swift look at [config.js](https://exercises.workroomprds.com/timeSince/config.js) to see its programatic context and at [CadenceKeeper](https://exercises.workroomprds.com/timeSince/) to see how it's used. #### using an LLM to do some work I asked Claude Haiku 3 `Tell me about the public methods that this exposes`, and it replied... The `Timer_v2` class exposes the following public methods: 1. ****constructor(name, startTime, uniqueId, interval)**: This is the constructor method that initializes a new `Timer_v2` instance. It takes the following parameters: 2. - `name`: A string representing the name of the timer. - `startTime`: An optional `Date` object representing the start time of the timer (defaults to the current time). - `uniqueId`: An optional unique identifier for the timer (defaults to the current timestamp). - `interval`: An optional interval value from the `UPDATE_INTERVAL` object, which determines the update interval for the timer (defaults to `UPDATE_INTERVAL.SECONDS`). 3. ****get elapsedTime()**: This is a getter method that returns the elapsed time in milliseconds since the timer was started. 4. ****setupTick(newInterval)**: This method sets up the timer's tick functionality. It takes an optional `newInterval` parameter to specify a new update interval. If no `newInterval` is provided, it will use the `currentInterval` property. 5. ****setupChangeovers()**: This method sets up the timer's changeovers from seconds to minutes and from minutes to hours. 6. ****stopUpdates()**: This method stops the timer's updates by clearing the interval and timeout IDs. 7. ****static fromObject(obj)**: This is a static method that creates a new `Timer_v2` instance from an object. If the object has an `id` property but no `uniqueId`, the `id` property is used as the `uniqueId`. These public methods allow you to create, manage, and interact with the `Timer_v2` instances in your application. ## Exercise 2 – have a look at the tests *5 mins* Have a look at the [unit tests](https://exercises.workroomprds.com/timeSince/jlChecks.html?spec=Timer%20v2) – is it clear what is being exercised? Not really, but the code of the page might tell you... #### another LLM shortcut I asked \`If the following are the unit tests, what public methods might match up with the tests?\`, and got back... Based on the unit tests provided, the following public methods from the `Timer_v2` class are likely to be covered: 1. ****constructor(name, startTime, uniqueId, interval)**: The constructor is tested in the "Construction" suite, where it checks the creation of timers with default and custom values. 2. ****get elapsedTime()**: The "It should keep track of elapsed time" suite tests the `elapsedTime` getter method, ensuring that it returns the correct elapsed time. 3. ****stopUpdates()**: The "Update Control" suite tests the `stopUpdates()` method, ensuring that it stops the timer's updates by clearing the interval and timeout IDs. 4. ****setupTick(newInterval)**: The "Interval Changes" suite (though currently marked as "xit") tests the behavior of the `setupTick()` method when the timer's age changes, ensuring that the update interval is adjusted accordingly. 5. ****static fromObject(obj)**: The "fromObject" suite tests the static `fromObject()` method, ensuring that it can create a `Timer_v2` instance from a stored object, including handling the legacy `id` property. Additionally, the "Event Handling" suite tests the timer's response to the `timerReset` event, which likely involves the `setupTick()` and `startTime` properties. The other public methods, such as `setupChangeovers()`, are not explicitly tested in the provided unit tests, but they may be covered implicitly or in other test suites not shown here. ## Exercise 3 – explore the unit *10 minutes* I had an irritating problem that I saw in [CadenceKeeper](https://exercises.workroomprds.com/timeSince/) and couldn't reproduce with the tests. So I asked for an exploratory interface, and played with the code that way... Go play with the [Exploratory Interface](https://exercises.workroomprds.com/timeSince/timer%5Fexplorer.html). Think about how it helps, and how it complements confirmatory unit tests and code analysis. 💡 An **exploratory interface* allows me to explore something. The commandline is certainly an exploratory interface, but this is a custom-built UI, so it's a harness and observatory and lets me trigger actions and iterate over things that might be iterable #### Using an LLM to build an exploratory interface Surely it is ridiculous to build a complex UI and then throw it away. It certainly **was* ridiculous. Not any more. Here's the prompt that got me close to this interface... I want an HTML page, using components in `../components` to let me explore `Timer_v2.js.js`. I want it to have sections containing a mockup UI for - elements that the code interacts with, - test UI elements \*\* to let me trigger all events that it listens to, \*\* to interface (input and output) with all functions that it exposes, \*\* to listen to and log any event that it may send \*\* to show the internal data and state of the object(s) – including any newly-made objects. I envisage a set of buttons for listened-to events and input areas if those events need a payload, and a set of collapsible sections for each and every exposed function, with any inputs needed for their payload, and the outputs processed to make them readable. Where an input has default values, I need an alternate interface to pass in specific values. Offer drop-downs if an input has a limted choice, and also offer an alternate to allow free entry of any data. If an input takes a range, offer an alterntive with a parametric slider between reasonable limits. Pay attention to the config file `config.js` , using values from that file – and also put a section containing a list of the values used into the interface, allowing them to be changed. In terms of look and feel: - Output **data* should be in a monospaced font. - Sections should be collapsible and re-orderable by dragging. - Show the data exposed by the objects and functions in real time, updated every 100ms – and add a pause+countdown button to stop that refresh for 5s to allow copy. ## Conclusion... Let's share and talk as we play. ### Workroom Playtime 032: Redirection Operators for Testers URL: https://www.workroom-productions.com/workroom-playtime-032/ Last updated: 2025-10-07T22:41:00.000Z This week's [Workroom PlayTime](https://www.workroom-productions.com/workroom-playtime/) is [Thursday 2 October at **3pm** London time](https://this-ti.me/?uts=1759413600&tz=Europe%2FLondon&name=Workroom+PlayTime+032). We'll gather on Zoom to play with exercises around [Redirection Operators](https://www.workroom-productions.com/exercise-redirection-operators/) – that's `|`, `>`, `>>`, `<` and a few related bits that you'd see on a unix commandline... These exercises are for everyone, for free. Paying [Subscribers](https://www.workroom-productions.com/#/portal/signup) get to play together every week. If you can see a **joining info** section below, then that gives you access to the Zoom etc. for the workshop. If you'd like to bring a friend, you can do that. I'll say yes until I've got too many. I'm prepping for an [ATD tutorial](https://agiletestingdays.com/2025/session/the-art-of-crafting-your-custom-tools/) at the moment, and we'll be pretty tool-focussed over the next four weeks as I try out elements. I plan to cover: - `032` 2 October: Redirection operators ( `|` etc.) for testers - `033` 9 October: Exploratory Interfaces II - `034` 16 October: `jq` for testers - `035` 23 October: build 3 throwaway tools in 15 minutes Thank you for reading! Cheers – James ↓ *signed-in subscribers will see joining info and links below* ↓ _This post is for subscribers only._ ### Exercise: Redirection Operators URL: https://www.workroom-productions.com/exercise-redirection-operators/ Last updated: 2025-10-02T13:58:23.000Z A set of testing-relevant exercises to help testers use, and see the use of, redirection operators on the commandline. To move information between commandline commands, use `|` to link inputs to outputs. Use `<` to take the input from a file, then use `>` and `>>` to put the output into a file. More info at [Redirection Operators for Testers](https://www.workroom-productions.com/redirection-operators-for-testers/). Use environments at Password is `password`, please choose a name to go to a home page, from there pick code server, and you should be in vscode. From there, open your home folder as a file browser, and open the terminal, and get the files needed with `git clone ` ## Exercise 1 – count the unique values in a table We'll use `items.csv` Use `cut -d',' -f4 items.csv` to see all the country codes. Use `cut -d',' -f4 items.csv | sort | uniq -c` to pipe that through `sort` and through `uniq -c` . You'll see the count of unique country codes. Try using an input redirect: `cut -d',' -f4 < items.csv | sort | uniq -c` . - You can move that redirect to the front of the command: `< items.csv cut -d',' -f4 items.csv | sort | uniq -c` . - You can use `cat` and a `|` to do (much) the same: `cat items.csv | cut -d',' -f4 items.csv | sort | uniq -c` . By now, you'll be frustrated that one of those codes is not a country code? Use `tail -n +2` to cut out the top line. #### how I did that... `< items.csv | tail -n +2 | cut -d',' -f4 | sort | uniq -c` Let's use `sort -nr` and `head -3` to see just the three countries with the greatest supply. #### how I did that... \`< items.csv | tail -n +2 | cut -d',' -f4 | sort | uniq -c | sort -nr | head -3\` Use `>` to put that output somewhere permanent. ## Exercise 2 – split up `find`'s output `find` has different outputs. It sends names of files found to *stdout*, and errors to *stderr*. Generally, *stdout* goes to the terminal (as a reply to your command), and so does *stderr*. If you're looking for one file on a disk, and you've got a heap of permission problems, you'll see mainly permission problems. Let's use the following, without `sudo` `find / -name items.*` compare with `find / -name items.* 2>/dev/null` The `2>` sends all the permission problems to `dev/null`, which is unix's bottomless pit, so you won't see permission problems at all. Use `>` to send stdout to a file (leaving stderr visible) Use `&2>` to send stderr to a file (leaving stdout visible) ## Exercise 3 – reading a command What does this do? ``` find /var/log -type f -print0 | xargs -0 file --mime-type | cut -d: -f2- | tr -d ' ' | sort | uniq -c | sort -nr ``` *(sorry that it doesn't wrap...)* You can use it and try cutting bits out to see experimentally, or ask an LLM. #### I reckon it... - lists all the files in `/var/log` – and lists them safely and with `null` delimiters - uses `xargs` to split at the nulls - passes them to the `file` command, asking for the type - cuts out the filetype column - strips whitespace - makes a count of type - sorts by most common type ## Exercise 4 – last 10 lines of recent logs This uses redirections – and also `xargs`, which runs code on each line of input - list files (with paths), sorted by recentness - take the top 3 - show the most recent 10 line of use those files-with-paths `sudo ls -1td /var/log/* | head -3 | xargs -I {} sh -c ' sudo tail -10 {} '` do this to add a descriptive line between log files `sudo ls -1td /var/log/* | head -3 | xargs -I {} sh -c 'echo "=== {} ==="; sudo tail -10 {}'` Note the double `sudo`... one for the `ls`, and a different one for the `tail` inside the `xargs` block. --- sprue from here... ## Exercise 4 – streaming into `tail` ## Exercise 5 – overwriting and appending ### Redirection Operators for Testers URL: https://www.workroom-productions.com/redirection-operators-for-testers/ Last updated: 2025-10-02T13:06:05.000Z Redirection operators (`|`, `>`, `>>`, `<`) let you chain Unix commandline commands together, and move information in and out of the chain. - `|` chains commands i.e. sends the output of one to the input of the next. More precisely, it sends the *stdout* of one the the *stdin* of the next, sending *stderr* to the terminal. - `<` sends input into commands – so you can (generally) - `*doathing* < file.txt` and - `< file.txt *doathing*` - which works similarly to the common chain `cat file.txt | *doathing*` - `>` and `>>` and `>&` redirects the output. (`>` overwrites a file while `>>` appends to a file, and `>&` redirects to a file descriptor – typically a stream). 💡 If the `>` has a space before it, it redirects **stdout*, leaving **stderr* to throw errors at you in the terminal. 💡 While `>` and `>>` are related, `|`(pipe) and `||` (OR) are not. `&`(file descriptor) and `&&` (AND) are not, either. ## Outputs and File descriptors *tl;dr `2>` sends stderr to a file and `>&2` sends an output to stderr* Commands output (what you expect) to *stdout*, which is typically the terminal window, and (problems) to *stderr*, which is also the terminal window. These are streams, and they're mixed as they are produced. `find`'s splurge of mingled filenames and permission errors is a perfect example of this. You can separate and redirect these. *stdout* is file 1 within the command, so you catch *stdout* in `*file.txt*` with `1> *file.txt*` . *stderr* is file 2, and you catch *stderr* in `*fubar*` with `2> *fubar*`. Numbered `file descriptor`s let you refer to these on the output: *stdin* `&0`, *stdout* `&1`, *stderr* `&2` . Note that while `N>` means send N somewhere, `>N` means send to a file called `N`, and you need `>&N` to tell the machine to use the descriptor. Common patterns: - `> *something*` sends *stdout* to *something,* leaving *stderr* on-screen in your terminal. - `2>*something*` sends *stderr* to *something*. - `2>&1` sends *stderr* to *stdout* - `&>*something*` redirects both *stdout* and *stderr* to *something* - `> /dev/null` or `1> /dev/null` to throw *stdout* away, `2> /dev/null` to throw *stderr* away. ⚠️ `curl` 's approach to **stderr* is unusual. If **stderr* is the terminal (i.e. if you're using it manually) it will report !progress (dynamically)! and errors. If not, (i.e. you're using it within a command) it'll just report errors. ### The numbers mean... - **0** represents standard input *stdin* - **1** represents standard output *stdout*, `>` is equivalent to `1>`. `&1` is equivalent to `/dev/stderr` - **2** represents standard error *stderr. `&2`* is equivalent to `/dev/stdout`. Higher numbers only make sense if they’ve been opened explicitly ## More thoughts What does `2>>&1` do? Nothing: it is bad syntax While `2>&1` sends errors to stdout, the order is not guaranteed: streams may be interleaved, but one does not overwrite the other. ### No Workroom PlayTime on Thursday 25 Sept URL: https://www.workroom-productions.com/no-workroom-playtime-on-thursday-25-sept/ Last updated: 2025-09-24T22:11:21.000Z Apologies – we'll go next week. Also: Two longer things are incoming (hint: prep for Agile Testing Days and more) ...mentioning [ATD](https://agiletestingdays.com/pricing/) – I have a 'personal speaker discount code' `25-150-JamLyn` that gives you €150 off the ticket price. I hear that it's a *stackable* discount – so if you're hovering over that '[buy](https://agiletestingdays.com/pricing/)' button, do it now. ### Workroom Playtime 031: Puzzle 17 URL: https://www.workroom-productions.com/workroom-playtime-031/ Last updated: 2025-09-17T20:50:13.000Z Join me on Zoom, on **Thursday 18 Sept at 3pm London time** ([local time for you](https://this-ti.me/?uts=1758204000&tz=Europe%2FLondon&name=Workroom+PlayTime+031)) to play with [Puzzle 17](https://blackboxpuzzles.workroomprds.com/newtech/puzzle17.html). These exercises are for everyone, for free. Paying [Subscribers](https://www.workroom-productions.com/#/portal/signup) get to play together every week. This week's is open to subscribers and friends: You'll need to be a signed-in subscriber to see the Zoom link below. Or just ask me. Thank you for reading! Cheers – James ↓ *signed-in subscribers will see joining info and links below* ↓ _This post is for subscribers only._ ### Workroom Playtime 030: Completeness URL: https://www.workroom-productions.com/workroom-playtime-030-completeness/ Last updated: 2025-09-10T16:07:07.000Z We'll gather on Zoom / Miro, on **Thursday 11 Sept at 4pm London time** ([local time for you](https://this-ti.me/?uts=1757602800&tz=Europe%2FLondon&name=Workroom+PlayTime+030)) to run [Completeness (is a wicked problem)](https://www.workroom-productions.com/completeness-is-a-wicked-problem/). You'll need to be a signed-in subscriber to see the links. Or just ask me. These exercises are for everyone, for free. Paying [Subscribers](https://www.workroom-productions.com/#/portal/signup) get to play together every week. This week's is open to subscribers and friends: Subscribers will see joining info below. Pop your name on the Miro board if you're coming. Thank you for reading! Cheers – James ↓ *signed-in subscribers will see joining info and links below* ↓ _This post is for subscribers only._ ### Completeness (is a wicked problem) URL: https://www.workroom-productions.com/completeness-is-a-wicked-problem/ Last updated: 2025-09-10T16:00:08.000Z *10 minutes exploration and analysis, 10 minutes conversation.* This exercise is to give us a moment to think about completeness – and a moment to consider how we work with [wicked problems](https://en.wikipedia.org/wiki/Wicked%5Fproblem). We'll explore [Converter](http://exercises.workroomprds.com/converter%5Fv2a/)v2 (part 1) Together, we'll consider what a 'complete' set of tests might look like, for this system. What makes that set 'complete'? **Debrief:** let's talk about what 'completeness' means, to you – and how that differs between us, and what that means in teams. --- Here's something on wicked problems ### Workroom Playtime 029: Imagined | Real URL: https://www.workroom-productions.com/workroom-playtime-029-imagined-real/ Last updated: 2025-09-03T20:46:24.000Z We'll gather on Zoom / Miro, on **Thursday 3 Sept at 3pm London time** ([local time for you](https://this-ti.me/?uts=1756994400&tz=Europe%2FLondon&name=Workroom+PlayTime+29)) to run [Imagined | Real](https://www.workroom-productions.com/imagined-real/). You'll need to be a signed-in subscriber to see the links. Or just ask me. These exercises are for everyone, for free. Paying [Subscribers](https://www.workroom-productions.com/#/portal/signup) get to play together every week. This week's is open to subscribers and friends: Subscribers will see joining info below. Pop your name on the Miro board if you're coming. In other news: No news. Everyone is very tired. News tomorrow. Or later. Who knows. Slight bug, I can't attach a picture to this post before emailing it. Thank you for reading! Cheers – James ↓ *signed-in subscribers will see joining info and links below* ↓ _This post is for subscribers only._ ### Imagined | Real URL: https://www.workroom-productions.com/imagined-real/ Last updated: 2025-09-04T14:00:12.000Z From *Exploratory Testing Workshop – Exploration II – Work.* This version for *Workroom PlayTime.* As testers, we regularly hold two comparable things in mind: something imagined, and something real. When trying to plan or complete work, we'll often try to work our way through one, looking at the other. We'll imagine a list of checks, and see whether they're fulfilled by the real software. Or we'll look at the methods exposed by an API, see what each does, and judge whether it's reasonable or weird. In this, we'll work with common artefacts used in planning and doing testing, and think together about how they fit into these patterns. *5-10 minutes* Look for the oval containing **imagined* | *real** Above it, there should be a set of stickies, each listing something you might use while testing a system. Collectively, sort the stickies into either side of the circle. For example *Execuable files* actually exist. *SLA*s describe what the system should be. If you find something fits both, or if you find something on the wrong side, duplicate it and make a note. Debrief – can we see patterns? Can we see iterables on both sides? What else might we add? What doesn't fit? What is ambiguous? *5 minutes* Considering the iterables – how might we iterate? ![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/2025/09/image.png) ### Workroom Playtime 028: My Story Chart URL: https://www.workroom-productions.com/workroom-playtime-028-my-story-chart/ Last updated: 2025-08-27T22:25:10.000Z We'll gather on Zoom / Miro, on Thursday 28 August at 4pm London time ([local time for you](https://this-ti.me/?uts=1756393200&tz=Europe%2FLondon&name=Workroom+PlayTime+028)) to run [My Story Chart](https://www.workroom-productions.com/my-story-chart/). You'll need to be a signed-in subscriber to see the links. Or just ask me. These exercises are for everyone, for free. Paying [Subscribers](https://www.workroom-productions.com/#/portal/signup) get to play together every week. This week's is open to subscribers and friends: Subscribers will see joining info below. Pop your name on the Miro board if you're coming. In other news: - No news: Back to London, no time for news. Thank you for reading! Cheers – James ↓ *signed-in subscribers will see joining info and links below* ↓ _This post is for subscribers only._ ### My Story Chart URL: https://www.workroom-productions.com/my-story-chart/ Last updated: 2025-08-27T22:05:23.000Z To play with telling a story about testing, illustrated with a picture that can be interpreted as values. ## Exercise *Drawing: 5 mins* Think of a testing story – something about testing that has a beginning and an end. Doesn't have to be true. True may be easier. Draw a chart of the story, picking out two or three values that changed meaningfully over the period of the story. Label the chart. Make it public. *Sharing: 5-10 mins* Have a swift look at all the charts We'll go round the group. Everyone will tell their story – initially with a sentence, then with a narrative. Tell the story rather than describe the chart. *Concluding: 5 mins* Share an insight or a difficulty that you understood from how your story worked for you. What would you keep, or change, about your story or your chart? *Extensions:* - ask the storyteller about another value – they can add it to the chart if - talk about why you (as the storyteller) chose those values and those labels - what happened before and after the time on the chart? Did the values exist meaningfully? Did they change? - could you have tracked these values as they happened? - could you have other interpretations of the same story with the same chart? - how did the chart help? - can you tell an entirely different story with the same chart? Try it! #### More about charts A graph is a picture that can be interpreted as values; a chart shows how things have changed over time. Typically the time shown is the same for all values, but the height of one value / line has a different meaning from another. Think of how something changed over time – did it get bigger or smaller? Was the change rapid or gradual? Did it keep changing? Draw a horizontal line to represent the time that your story happens over. Draw another line showing how that something changed over time – height relates to how this thing changed. It's fine for the wobbly line to go through the horizontal line (perhaps you're drawing a picture of money in an account). It doesn't make sense for the line to have two heights at the same time, so avoid doubling back or looping. ### Workroom PlayTime 027: Being Random URL: https://www.workroom-productions.com/workroom-playtime-027/ Last updated: 2025-08-19T20:49:05.000Z We'll gather on Zoom / Miro, on Thursday 21 August at 4pm London time ([local time for you](https://this-ti.me/?uts=1755788400&tz=Europe%2FLondon&name=Workroom+PlayTime+027)) to run [Being Random](https://www.workroom-productions.com/being-random/). These exercises are for everyone, for free. Paying [Subscribers](https://www.workroom-productions.com/#/portal/signup) get to play together every week. This week's is open to subscribers and friends: Subscribers will see joining info below. Pop your name on the Miro board if you're coming. In other news: - No news: I'm by the sea. Lovely. Last few days. Thank you for reading! Cheers – James ↓ *joining info and links below* ↓ _This post is for subscribers only._ ### Exercise: Being Random URL: https://www.workroom-productions.com/being-random/ Last updated: 2025-08-21T13:46:52.000Z *Used in* [*Workroom PlayTime 027*](https://www.workroom-productions.com/workroom-playtime-027/) Randomness is hard to assess. While testing, we commonly need random data – sometimes to be meaningless, sometimes to be more 'real', sometimes to avoid disrupting behaviour with unintended patterns. This exercise lets us explore different ways that patterns can be present and absent in a string of text. #### More to read and play with See [Randomness Test](https://en.wikipedia.org/wiki/Randomness%5Ftest) for information on how to check for randomness, and play with (in the exercise below) Machine-made randomness and human-made randomness are different. Humans can make meaningful strings that algorithms judge to be random, while human attempts at randomness often fail algorithmic assessments. Algorithms can make endless random output, but that randomness may be constrained to work only within the bounds of the algorithm. See [Moravec's paradox](https://en.wikipedia.org/wiki/Moravec%27s%5Fparadox) for more on tasks that algorithms find hard but people find easy Play with the [ARC-AGI daily puzzle](https://arcprize.org/play) for direct experience with something that is designed for people to manage but which AIs find tricky. The exercise involves assessing and producing random text. ## Examples / questions How is `abbabbbbababbababbbbabababaaababababbabba` random? Compare with `` or `` how about ``? how about `ropemediumfillsadtrout`? Or `азсекасбамджеймс`? Let's talk about these examples, and reflect on what we understand by random. ## Exercise Use . get rid of the video overlay, chose "manual ... input" and start hammering 1s and 0s. Hit "start now" to test. Try to type in a random string of 1s and 0s 😥 **My software for the exercise is not ready for primetime.* ## Reflect What insights can you share? You've used randomness in testing – did it matter, to the tests, whether 'random' was random? How did you check? --- ?*\[placeholder for now – needs instructions, text input, submission mechanism, analysis\]* \-- JL Notes: We need random data *to data that is effectively meaningless to act as a control (i.e. text field content), distributed evenly but not regularly across a range (, or distributed around some value (as contrasted to being bang on)* *JL Notes: We test for randomness with* JL notes: we generate randomness with (entropy, algorithms, keyboard bashing). We need to be aware of distribution, of sample size, of Type a random string in the box below. Hit `assess` to see an algorithmic assessment of how random it is. Input Text: Assess Analysis Results: Results will appear here... ### Exercise 1 Try to make a string that seems random to the assessor ### Exercise 2 Try to make a string that seems random to the assessor, but is clearly not random. You're welcome to explain the trick that makes it deterministic to you. ### Exercise 3 What randomnesses, and what patterns, is the assessor looking for? Explore to discover, share your findings and examples. ## Reflect ### Workroom PlayTime 026: Shared Inspiration URL: https://www.workroom-productions.com/workroom-playtime-026-shared-inspiration/ Last updated: 2025-08-13T12:54:29.000Z We'll gather on Zoom / Miro, on Thursday 13 August at **4pm** London time ([local time for you](https://this-ti.me/?uts=1755183600&tz=Europe%2FLondon&name=Workroom+PlayTime+026)) to run the playful exercise [Shared Inspiration](https://www.workroom-productions.com/playful-exercises-about-testing/). These exercises are for everyone, for free. Paying [Subscribers](https://www.workroom-productions.com/#/portal/signup) get to play together every week. This week's is open to subscribers and friends: Subscribers will see joining info below. Pop your name on the Miro board if you're coming. In other news: - Tour's done, and as of yesterday I'm on holiday. Woohoo! So it's good news: there's no news. Thank you for reading! Cheers – James ↓ *joining info and links below* ↓ _This post is for subscribers only._ ### No Workroom PlayTime until 14 August URL: https://www.workroom-productions.com/no-workroom-playtime-until-14-august/ Last updated: 2025-07-30T22:37:28.000Z I've stepped away to be with family, and to sing. Tomorrow I'm travelling all day. You'll know that I had hopes of doing Workroom PlayTime 26 at a different time this week – hasn't happened, and now can't happen. Cheers, all. May your tests fail delightfully. Thank you for reading! James ### Workroom PlayTime 025: curl URL: https://www.workroom-productions.com/workroom-playtime-025-curl/ Last updated: 2025-08-13T12:52:52.000Z We'll gather on Zoom / Miro, on Thursday 25 July at **4pm** London time ([local time for you](https://this-ti.me/?uts=1753369200&tz=Europe%2FLondon&name=Workroom+PlayTime+025)) to play with data transfer tool [cURL](https://everything.curl.dev). The exercises aren't up yet... I'll update this page when they are. These exercises are for everyone, for free. Paying [Subscribers](https://www.workroom-productions.com/#/portal/signup) get to play together every week. This week's is open to subscribers and friends: Subscribers will see joining info below. Pop your name on the Miro board if you're coming. In other news: - I published one more tiny tool, to [clear up your LinkedIn feed](https://www.workroom-productions.com/clear-up-your-linkedin-feed/). Not testing related, but feels neatly empowering to use. - *Advance notice:* Workroom PlayTime for Thursdays 31 July and 7 August will certainly be shunted about – I'm on stage for one and in the air for the other. They'll move, but I can't yet see when I can move them *to.* Thank you for reading! Cheers – James *I announced this Workroom PlayTime just a few hours before the scheduled time. Here are some thoughts on that.* *What happened?* Sometime around 1 this morning (it was dark, I had my eyes closed), I recognised that I had not sent an email to invite subscribers. *Realised* is the wrong word: there were no surprises here. The slot for Workroom PlayTime was in the family and work diaries, I've arranged stuff to avoid those slots, and I'm expecting to run it. Also, I had not set up the exercises, the support page, the environment, the miro board... or invited *you.* These two incompatible truths were present in one mind, together. I have been trying to put an experience together that explores how we as testers need to be able to hold incompatible truths: It's interesting to be in the middle of a practical example. No excuses here, but it's an interesting observation about minds, and I wanted to share. I do, of course, need to apologise for being late to announce. I'm *often* late to announce, despite planning to plan, and despite experimenting with changes so that this doesn't happen in the same way another time. So if that's how it works and I can't find a way to change it, I'll need to accept the situation rather than fight its nature. One reason to offer Workroom PlayTime is to explore how it works: I'll give that some more thought over the summer. ↓ *joining info and links below* ↓ _This post is for subscribers only._ ### cURL Exercises for Testers URL: https://www.workroom-productions.com/curl-exercises-for-testers/ Last updated: 2025-07-24T11:51:01.000Z *placeholder* ## Information about an endpoint Try `curl -I "https://www.workroom-productions.com/curl-exercises-for-testers/"` compare with `curl -I "https://agiletestingdays.com/2025/session/the-art-of-crafting-your-custom-tools/"` ## Using with other tools curl uses different output streams: - **stdout**: HTTP response headers/body - **stderr**: Progress information, error messages When curl detects: - **Terminal output**: Shows headers only, no progress meter - **Piped output**: Shows progress meter to stderr, sends headers to stdout See what happens when piping to other tools: `curl -I "https://www.workroom-productions.com/curl-exercises-for-testers/" | grep "status"` Note the unexpected *extra* info compared with `curl -I "https://www.workroom-productions.com/curl-exercises-for-testers/"` Manage this with `-s` Compare with `curl -sI "https://www.workroom-productions.com/curl-exercises-for-testers/" | grep "Status"` ## Information about a connection try `curl -vI "https://www.workroom-productions.com/curl-exercises-for-testers/"` ## Speed try `curl -w "Total time: %{time_total}s\n" -s -o /dev/null "https://www.workroom-productions.com/curl-exercises-for-testers/"` This uses the `-w` writeeout option, which exposes information about the transaction: | Variable | Description | | ---------------------- | ----------------------------- | | %{content\_type} | Content-Type of response | | %{http\_code} | HTTP status code | | %{http\_connect} | HTTP connect response code | | %{local\_ip} | Local IP address | | %{local\_port} | Local port number | | %{num\_connects} | Number of connections made | | %{num\_redirects} | Number of redirects | | %{remote\_ip} | Remote IP address | | %{remote\_port} | Remote port number | | %{size\_download} | Bytes downloaded | | %{size\_header} | Bytes of headers downloaded | | %{size\_request} | Bytes sent in request | | %{size\_upload} | Bytes uploaded | | %{speed\_download} | Download speed (bytes/sec) | | %{speed\_upload} | Upload speed (bytes/sec) | | %{time\_appconnect} | SSL/TLS handshake time | | %{time\_connect} | Connection establishment time | | %{time\_namelookup} | DNS lookup time | | %{time\_pretransfer} | Pre-transfer time | | %{time\_redirect} | Redirect time | | %{time\_starttransfer} | Time to first byte | | %{time\_total} | Total time | | %{url\_effective} | Final URL after redirects | ### Clear up your LinkedIn feed URL: https://www.workroom-productions.com/clear-up-your-linkedin-feed/ Last updated: 2026-02-10T20:07:14.000Z *Worked in 2025\.* *In 2026, I needed a new bookmarklet, because LinkedIn changed from meaningful class names to obscure names. And more. Took about 20 minutes while I was multitasking audio stuff, giving the old bookmarklet and some new examples to Claude Sonnet 4.5\. You can do it, too – do ping me if you want the new one.* Here's my bookmarklet to remove various LinkedIn posts. Use Method 1 from [Wikipedia on Bookmarklets](https://en.wikipedia.org/wiki/Bookmarklet) to install. ```javascript javascript:(function(){let e=document.querySelectorAll('div.occludable-update');for(let i=0;is with a class of "occludable-update", where a contained with class "update-components-header__text-view" contains text "likes this" ``` Claude gave me options and minimised the simplest to go in a bookmarklet. A bit of fiddling from me (once I'd used it) to get rid of more, and we have it. Claude gave me another option to keep running it for 60s throwing away newly-loaded qualifying posts. I'll put that below if it's useful to me without being too disruptive. This is a response to this:[ Top of my feed: a post by an MP calling for "mass deportations"](https://www.linkedin.com/posts/jameslyndsay%5Ftop-of-my-feed-a-post-by-an-mp-calling-for-activity-7351914217274241024-TJQa?utm%5Fsource=share&utm%5Fmedium=member%5Fdesktop&rcm=ACoAAAAyiC8BkEoqTadPzV9zHdb5crUj0ZK30y0). ### Workroom PlayTime 024: Puzzle 11 (is an awful toy) URL: https://www.workroom-productions.com/workroom-playtime-024-puzzle-11/ Last updated: 2025-07-16T16:27:15.000Z We'll gather on Zoom / Miro, on Thursday 3 July at **4pm** London time ([local time for you](https://this-ti.me/?uts=1752764400&tz=Europe%2FLondon&name=Workroom+PlayTime+024)) to play together with [Puzzle11](https://nwtg.workroomprds.com/puzzle11.html). Puzzle 11 is simple, and it can also be an excruciating experience. We'll play with the puzzle, and talk about what circumstances and characteristics make it unpleasant. It is, perhaps, an interesting exercise in what makes an exercise *not* work. These exercises are for everyone, for free. Paying [Subscribers](https://www.workroom-productions.com/#/portal/signup) get to play together every week. This week's is open to subscribers and friends: Subscribers will see joining info below. Pop your name on the Miro board if you're coming. In other news: - I've recently published two tiny tools: [checkRefs](https://www.workroom-productions.com/tiny-tool-checkrefs-sh/) and [summarise history](https://www.workroom-productions.com/summarise-history-prompt/). Subscribers get first look at the LLM prompts and failures that helped me get to useful tools. - Over the last couple of Mondays, I ran a 90-minute online workshop [Test-Driven Generation](https://www.workroom-productions.com/online-test-driven-generation-workshop/) and a [small conversation and demo group on ephemeral environments](https://www.workroom-productions.com/ephemeral-environment-workgroup/) for teaching testing (and, by extension, for testing). If you're sad that you missed them, I'll run them again in October: tell me that you're interested. - As the summer moves on, my plans get ever-more out-of-shape. Thank you for bearing with me. I plan to carry on with Workroom PlayTime while I'm travelling with family, but expect bigger changes. *Advance notice:* Workroom PlayTime for Thursdays 31 July and 7 August will certainly be shunted about – I'm on stage for one and in the air for the other. They'll move, but I can't yet see when I can move them *to.* Thank you for reading! Cheers – James ↓ *joining info and links below* ↓ _This post is for subscribers only._ ### Tiny Tool: Summarise History Prompt URL: https://www.workroom-productions.com/summarise-history-prompt/ Last updated: 2025-07-18T14:58:35.000Z ```shell history -300 | llm -m claude-3.5-sonnet "write a consise narrative of my recent history – comment only on the actions and on particular patterns and groupings. Don't comment on me. Don't summarise at the end. pay particular attention to file location from cd to give or change context. A narrative is sequential, and looks like: «Called Ansible playbook xxx several times. Edited file yyy. Changed dirextory to zzzz. Switched focus to llm tooling.»" ``` This one-liner summarises my recent actions on the commandline. I've used it at the end of a session of exploring / configuring / coding to remind me what I've been doing. Here's an example of what it produces: > Started by looking at file headers and random data using head/tail commands on /dev/urandom. Created and edited files boo.yog and doo.txt, experimenting with various cat, head, and tail options. Monitored system logs using tail -F on various log files in /var/log/. > Made extensive use of Ansible, running multiple playbooks particularly user\_setup\_and\_info.yml with the "7july" group parameter. Accessed secrets using 1Password (op) for vault passwords. Worked with Docker, checking versions and installing via Homebrew. > Changed directory to ansibleSetup, worked with droplet\_only\_setup.yml playbook multiple times. Created and managed vault-encrypted secrets files. Created a checkRefs.sh script and made it executable, then used it to process template files in the templates directory. > Finally, switched focus to exploring llm tooling, experimenting with different ways to analyze command history using llm with various models and parameters, particularly trying to generate narrative summaries of recent activities. The tool pulls together: - the shell history (all the commandline commands I've entered this session), accessed through the `history` command\*. I estimate how many entries I've made that I want to summarise. - Simon Willison's `llm` tool, which lets me experiment with my choice of large language model and gives me a transparent way to iterate on my prompt. - A text prompt, which could be better, but is better than I started with I built the tool because I find that I need a reminder of what I've been doing – most recently I needed to explore a novel-to-me build tool while assessing the development experience, and found that I needed more reminders than I'd imagined. Browsing the history helps, but I need to re-interpret commands and gloss over the long ones. Having an acceptable summary helps me remember, which gives me a good place to start when looking over the history. It also gives me a good place to start if I need to summarise what I've been doing. Making the tool took about 15 minutes, half of which on getting to the bottom of which history was in use, and the rest fiddling with the prompt and auditioning model. Writing about making it has taken an hour. I imagine I'll continue to use it when I'm exploring something using the commandline, and when I'm building. Changes: I'd love it to work on timestamped stuff, so that I can see sessions or ask for 'in the last two hours', or persuade the LLM to recognise bursts, or (maybe) combine two parallel histories in two simultaneous terminals. To do that I'll need to find a way to get VSCode to timestamp its history. It seems possible with bash and zsh history, so I'm hopeful. I'd like it to run every time I leave the commandline for (say) more than 15 minutes, and pop the summary somewhere I can see it. I'd like it to be much more clear about changes of context and directory. I need to use it in situations where I've been swinging about a bit. I want to see what it makes of several old terminal windows (with their own history) so I can see what I was doing! I particularly want to try combining with other timestamped logs, especially with dictated notes, system event sniffers, and trend monitors. Subscribers get to see how the prompt evolved, and what other the other models produced. \`\* The `history` command? I'd assumed that history would be in the file referenced by `$HISTFILE`, which might be `/.zsh_history` or `/.bash_history` – but both these are wrong: I've been working in the terminal in VSCode. VSCode keeps history in `~/Library/Application Support/Code/User/History/` on a Mac, referenced by a randomish number. So `history` it is – what a handy shim. _This post is for subscribers only._ ### Tiny Tool: checkRefs.sh URL: https://www.workroom-productions.com/tiny-tool-checkrefs-sh/ Last updated: 2025-11-25T22:35:20.000Z #### Addendum: why not `xargs`? Why not, indeed? I revisited this tiny tool as I built a short workshop on `xargs` and rethought (again, with an LLM). This one-liner is the end product: `ls -1 ./templates | xargs -I{} sh -c 'echo "## {}"; grep -Flsir "{}" ./* ' ` That substitutes the shell script below for `xargs -I{} sh -c 'echo "## {}"; grep -Flsir "{}" ./* '` What does this do? - `ls -1 ./templates` makes a list of the filenames in the `templates` directory - `|` sends the list to xargs - `xargs -I{} sh -c` works through that list of filenames, runs the shell script in the single quotes `'...'`, and substitutes any `{}` in the shell script for the filename. - `echo "## {}";` Sends the filename prefixed with `##` to the terminal - `grep -Flsir "{}" ./* '` seeks the filename in any of the files in the current directory, with `-Flsi` options as below. Note the new `r` recursive option, so it track tracks down the hierarchy, and the starting point and pattern `./`**.* I could use `.` for the current directory, but this pattern lets me change directories up (`../`) and down (`./target/*`). At the moment, this picks up the template directory, too. ```bash while read -r referenced_String; do matches=$(grep -Flsi "$referenced_String" *) echo "## ${referenced_String}" if [ -n "$matches" ]; then echo "$matches" else echo "unreferenced" fi done ``` Used with `ls -1 ./templates | sh ./checkRefs.sh`, this will check the files in the current directory to see which ones use a file from `templates` To try it out, clone this for an example set of files, the shell script and command. #### command `sudo git clone ` I used this when building playbooks, to see which templates were used by files in the current directory – and which were unused. If a file referenced two templates, I could see that, too. I built the tool because I did the work by hand a couple of times, and that became slow and error-prone as the number of files and strings increased. I couldn't find a way to make a tool (either on the commandline or within my IDE) do it. So it's tool time. Making it took about 15 minutes, and taught me some stuff – time well spent. Writing about making it has taken an hour. I imagine I'll use it generally: I'll certainly use it on my current project again. Subscribers get to see how I used LLMs Qwen3 and Claude3.7Sonnet. ## Building I built it / breathed it into existence by asking Qwen3 (local, thinking) and Claude3.7 (cloud, powerful, thinking off). - Qwen3 ended up with a script close to the one above. - Claude3.7 built something more robust, with better output, but a bit harder to read. - I took Qwen's, added the "unreferenced" from Claude and a `grep` option, and fiddled the `referenced_String` to help me understand next time. I checked its output against something I'd already done by hand – it matched. To check that a string referenced multiple times turned up several times, and at the same time checking the behaviour for a file referencing several strings (which wasn't going to be checked by the set I was using by hand), I temporarily popped a file listing of the `templates` directory into the current directory – so *every* string was in twice, and one file had them all. That took less time to do than to describe, and worked to my satisfaction. ## Details of prompting The prompt I used started with was : ```text Given a list of strings (i.e. a file of newline-delimited filenames) I want to search through a directory of files to see where those filenames are used, and whether any filename is unused. ``` I had a couple of conversations with Qwen, which was fixating on seeing what files were in both directories, and got very confused in its thinking (particularly with a list of filenames to search in a list of named files). Those conversations helped me to refine by asking for a shell script, by clarifying the difference between the strings to be searched for and the files to search, and by clarifying the way I'd use it in terms of input and output. Watching Qwen also gave me time to pause and think – but it's a mistake to imagine that one's thinking is improved while watching an LLM chew its own tail. The eventual conversation starting prompt ended up being: ```text Given a list of strings (i.e. a file of newline-delimited filenames) I want to search through a directory of files to see which files contain those strings/filenames, and whether any string/filename is unused. I want a shell script, taking stdin as input. I will probably pipi ls -1 to the input. I want the output organised (for now) by the input list in the form ## «string» ``` For this, Qwen got (literally) lost in thought over several minutes. It eventually output code in its 'think' response, threw up into the conversation title (weird...) and failed to output anything else at all. But the code looked OK. I ran the same prompt with Claude3.7Sonnet, which gave me 50+ lines of guff. I asked several times to make the offered code terse and readable: by using `grep` rather than `find`, by *removing* robustness, by *removing* a "count" function, and by *removing* bullet points in the output. I asked it to compare its eventual output with Qwen's / mine, and it outlined differences that... led me back to the simpler one. I say "simpler", but the `grep -Flsi` bit is complex, and more complex than suggested by either. Qwen went for `grep -Fls`, Claude went even simpler but did suggest `-i` to make the command case-insensitive (I want to see if something's used, not if it's used well, so that's better). I can imagine that I'll wonder what they meant, and I imagine I'll ask an LLM, so... Here's Claude to explain: #### In bullet points, talk me through the options used in `grep -Flsi` - ****F**: ****Fixed strings** \- Treats the pattern as a fixed string rather than a regular expression. This means special regex characters like `.`, `*`, `[`, etc. are treated literally. - ****l**: ****List filenames only** \- Instead of showing the matching lines, grep only outputs the names of files containing matches. Each filename is printed once, even if multiple matches exist in the file. - ****s**: ****Suppress error messages** \- Silences error messages about nonexistent or unreadable files. Useful for avoiding error output when grep encounters permission issues or other errors. - ****i**: ****Case-insensitive** \- Makes the search ignore case differences between the pattern and the file contents. For example, searching for "hello" would match "Hello", "HELLO", etc. Together, these options make grep return just the filenames where the exact string appears (ignoring case), without showing any error messages. (LLM-generated content, judged reasonable by me) ### Workroom PlayTime 024 – not this week URL: https://www.workroom-productions.com/workroom-playtime-024-not-this-week/ Last updated: 2025-07-09T22:35:11.000Z Probably no Workroom Playtime this week. In other news: - Last Monday I ran [Test-Driven Generation](https://www.workroom-productions.com/online-test-driven-generation-workshop/) for subscribers. I'll do it again in September / October, this time at US-friendly hours. - Next Monday [14 July 4:30 - 6pm London time](https://this-ti.me/?uts=1751877000&tz=Europe%2FLondon&name=Test-driven+Generation) I'll run a [small conversation and demo group on ephemeral environments](https://www.workroom-productions.com/ephemeral-environment-workgroup/) for teaching testing (and, by extension, for testing). It's free and it is for people who are already teaching testing. Ping me for a space. Thank you for reading! Cheers – James ↓ *joining info and links below* ↓ _This post is for subscribers only._ ### Monday 7/7 Workshop on Test-driven Generation. URL: https://www.workroom-productions.com/monday-7-7-workshop-on-test-driven-generation/ Last updated: 2025-07-03T10:30:54.000Z Play with tests, code and LLMs on Monday 7 July. _This post is for paying subscribers only._ ### Workroom PlayTime 023: LLM comparison I URL: https://www.workroom-productions.com/workroom-playtime-023/ Last updated: 2025-07-02T23:48:13.000Z We'll gather on Zoom / Miro, on Thursday 3 July at **4pm** London time ([local time for you](https://this-ti.me/?uts=1751554800&tz=Europe%2FLondon&name=Workroom+PlayTime+023)) to play together with testing stuff. (apologies – this said 4 July for a few hours) ![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/2025/07/image-2.png) We'll do the following exercise, where we compare several LLM's responses to the same prompt. We pick that prompt to be something in which we are already expert, so that we have something to judge against. We compare in several ways. I hope that the exercise will help use explore *judgement*. [LLM comparison – local knowledgeCompare LLM output – something obscure that you know well.![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/icon/wpl-60x60-logo-semitrans-14.png)Workroom ProductionsJames Lyndsay![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/thumbnail/066E8C8E-A47A-4E83-A10D-546B35422DF1_1_105_c-1-1.jpeg)](https://www.workroom-productions.com/llm-comparison-local-knowledge/) These exercises are for everyone, for free. Paying [Subscribers](https://www.workroom-productions.com/#/portal/signup) get to play together every week. This week's is open to subscribers and friends: Subscribers will see joining info below. Pop your name on the Miro board if you're coming. In other news: - Next Monday [7 July 9:30-11am London time](https://this-ti.me/?uts=1751877000&tz=Europe%2FLondon&name=Test-driven+Generation) I'll run [Test-Driven Generation](https://www.workroom-productions.com/online-test-driven-generation-workshop/) – I've run that workshop in various forms here, at AgileTesting Days and at EuroSTAR. You'll generate code that passes tests, and see how weird that is. It's for paying subscribers – if you already are, ping me or put your name on the board for a space. If you're not a paying subscriber... [subscribe](https://www.workroom-productions.com/#/portal/signup)! - The following Monday [14 July 4:30 - 6pm London time](https://this-ti.me/?uts=1751877000&tz=Europe%2FLondon&name=Test-driven+Generation) I'll run a [small conversation and demo group on ephemeral environments](https://www.workroom-productions.com/ephemeral-environment-workgroup/) for teaching testing (and, by extension, for testing). It's free and it is for people who are already teaching testing. Ping me for a space. Thank you for reading! Cheers – James ↓ *joining info and links below* ↓ _This post is for subscribers only._ ### Workroom PlayTime 022: Testing Story URL: https://www.workroom-productions.com/workroom-playtime-022/ Last updated: 2025-06-30T13:58:24.000Z We'll gather on Zoom / Miro, on Thursday 26 June at **4pm** London time ([local time for you](https://this-ti.me/?uts=1750950000&tz=Europe%2FLondon&name=Workroom+PlayTime+022)) to play together with testing stuff. We'll tell a collaborative story about testing, and see what happens. Here's more (not much more) about [Testing Story](https://www.workroom-productions.com/playful-exercises-about-testing/#testing-story). These exercises are for everyone, for free. Paying [Subscribers](https://www.workroom-productions.com/#/portal/signup) get to play together every week. This week's is open to subscribers and friends: Subscribers will see joining info below. Pop your name on the Miro board if you're coming. In other news: - Here's a new [exercise comparing LLMs](https://www.workroom-productions.com/llm-comparison-local-knowledge/) – this time on something that you know well (it's illustrative example is something *I* know well). It's a companion piece to [Different LLMs do Different Things](https://www.workroom-productions.com/different-llms-do-different-things/), which looked at code. We'll run it as an exercise soon. - Here are last week's [exercises for cat, head and tail](https://www.workroom-productions.com/exercises-for-cat-head-and-tail/), and a page on [how cat, head and tail might be used](https://www.workroom-productions.com/cat-head-and-tail-for-testers/) by testers. - I've updated my [page for the command-line tool series](https://www.workroom-productions.com/unix-tools-for-testers/) to move it (mostly) into public. It will be past of my contributions to a [tutorial](https://agiletestingdays.com/2025/session/the-art-of-crafting-your-custom-tools/) with [Bart](https://agiletestingdays.com/speaker/bart-knaack/) and [Huib](https://agiletestingdays.com/speaker/huib-schoots/) about tools for testers which we're building for a debut at [Agile Testing Days](https://agiletestingdays.com). It's a full-day interactive, and we hope you'll come away with tools that you've built for yourself and for your needs. [Sign up](https://agiletestingdays.com/register/)! - FINALLY, announcing two workshops (with tentative dates and times). There are barriers to entry: the test-driven generation one is free to *paid* subscribers, the environments one has prep work and I'm offering it for free to people who are already teaching. Ping me if you want to put your name in for any of those – I've got some names already but don't know whether anyone can make those times. I will run these again at *different* times soon – so tell me if you want one in your timezone. - 1) a [re-run of the Test-Driven Generation](https://www.workroom-productions.com/online-test-driven-generation-workshop/) EuroSTAR session, tentatively for [7 July 9:30-11am London time](https://this-ti.me/?uts=1751877000&tz=Europe%2FLondon&name=Test-driven+Generation) - 2) a [small conversation and demo group on ephemeral environments](https://www.workroom-productions.com/ephemeral-environment-workgroup/) for teaching testing (and, by extension, for testing) on [14 July 4:30 - 6pm London time](https://this-ti.me/?uts=1751877000&tz=Europe%2FLondon&name=Test-driven+Generation). Thank you for reading! Cheers – James ↓ *joining info and links below* ↓ _This post is for subscribers only._ ### Online Ephemeral Environment workgroup URL: https://www.workroom-productions.com/ephemeral-environment-workgroup/ Last updated: 2025-06-25T20:13:27.000Z *This is a placeholder page.* I'm going to run a small irregular group for conversation and demos around ephemeral environments for teaching testing (and, by extension, for testing) on [14 July 4:30 - 6pm London time](https://this-ti.me/?uts=1751877000&tz=Europe%2FLondon&name=Test-driven+Generation). You'll need be already be someone who teaches testing in public. I'll kick off with an introduction to the scripts and tools I used to deliver a hands-on session with command-line server access, LLMs, test tools and more – all through a browser window, to 80 people, over conference wifi. I'll post materials here soon. For now, here's [Ephemeral Environments](https://www.workroom-productions.com/ephemeral-environments/) and [Browser-based VSCode for Workshops](https://www.workroom-productions.com/browser-based-vscode-for-workshops/). ### Online Test-Driven Generation workshop URL: https://www.workroom-productions.com/online-test-driven-generation-workshop/ Last updated: 2025-07-06T16:06:29.000Z I'm going to run my hands-on [test-driven generation](https://www.workroom-productions.com/test-driven-generation-a-hands-on-experience-at-eurostar-2025-2/) workshop online for *paying* subscribers. Subscribe to come to the workshop. The workshop will be online on [7 July 9:30-11am London time](https://this-ti.me/?uts=1751877000&tz=Europe%2FLondon&name=Test-driven+Generation) . We'll be on Zoom. It will be a hands-on workshop. You don't need to install anything. You do need something that you can type on, that has a modern browser. You will, with guidance: - generate code from tests - adjust the tests to see what happens to the code - explore the weirdness - see how the moving parts fit together And you will take away: - a hands-on experience of working with an LLM on code and tests - an understanding of one (or several) ways to guide an LLM in a non-chat-based, tools-based. agentic way. Subscribe to the site from this page (pink icon) to tell me you're coming to the workshop, and I'll take it from there. If you're already a paid subscriber, you'll see joining details below. _This post is for paying subscribers only._ ### LLM comparison – local knowledge URL: https://www.workroom-productions.com/llm-comparison-local-knowledge/ Last updated: 2025-07-03T21:41:31.000Z *Used in* [*Workroom PlayTime 023*](https://www.workroom-productions.com/workroom-playtime-023/) This exercise lets us explore judgement and how our own expertise and aesthetic influences judgement. The exercise involves comparing several LLM's responses to the same prompt. We pick that prompt to be something in which we are already expert, so that we have something to judge against. We compare in several ways. ## Exercise Take something you know about, but which has few sources of information. Pick something that has *verifiable* facts – particularly facts that you have a primary source for. Ask several LLMs about it. Share your prompt before you explore. - Use the same prompt for all the LLMs. - Use one prompt – don't get into conversation. - Try (this may be harder as tech changes) to avoid having any personal history between you and the LLM that might influence its answer. - Try not to let answers cross-polinate – if you're using a tool that lets you switch LLM mid-conversation, it may pick up the previous LLM's answer as part of your prompt. - Avoid allowing the LLM to query and parse the web to get your answer, unless that's what you're looking to test. Do try the same LLM several times of you like. Share the answers as they come in. Compare answers between different LLMs– and keep an eye out for the mechanisms *you* use to compare. Put your hat of cynicism on and consider whether the LLM has offered verifiable facts, general sentiment, and to what extent it has echoed your prompt. Can you see shared facts – and are they right? Can you distinguish between information you asked for, and what else LLMs might be giving you to reflect your tone or otherwise cold-read what might satisfy? Can you see things that look right, but are inconsistent when take together? If you're using them, are there differences between the conclusions of 'thinking' models and non-'thinking' – between small local models and huge remote ones? Compare the LLMs' answers against what you know, first-hand. What's certainly wrong and certainly right? What's probably wrong, and what's surprisingly right, and how have you verified its statements? Here's something I did. ## My expertise: Lettsom Gardens I used [Msty](https://msty.app/) to run the same prompt simultaneously on these LLMs: a query about [part of my local area](https://www.lettsomgardens.org.uk/?page=2) that I know fairly well. > What do you know about Lettsom Gardens and the surrounding area in London? Let's remember that LLMs are making it all up, all of the time – but sometimes their fantasy is directed and constrained by their training, system prompts, what they've retrieved or recently mentioned, and more. Note: none of these were doing retrieval from the web. All were fresh conversations. The local models take up only 2-4 Gb – treat whatever they gave me as the product of *very* lossy compression. In summary: - Local models (as in local to my laptop, not trained in local knowledge) **Qwen3** and **DeepSeekR1** entirely made up almost every detail of history and features. In terms of geography, they placed it in the wrongly, and were geographically illiterate about London. **Llama 3.2** made excuses, offered something that it indicated was tentative and tangential, and stopped. Good for Llama 3.2. - **GPT4o-mini** located it within a couple of miles, got the right reason for the name in §1, then imagined a handful of features and described the (wrong) neighbourhood. - **Claude 3.7 Sonnet** got the location right, and included 3x as many checkable facts as 4o-mini. Almost all of those facts match my local knowledge. - **GPT-4** got the location and several facts right, made up a pond and a hardwood forest, and spent the rest of its short answer extolling mostly-vague virtues of nearby Camberwell. I used my knowledge as a local resident; not only what I know from being a key-holder and user of the garden, but what I know by direct experience, what I've read / heard and remembered, what I refreshed from the garden's website, from wikipedia and from pages via the council and the local school. The information returned by the LLMs is detailed below. I don't want it turning up in training data, so I've paywalled it. _This post is for paying subscribers only._ ### Exercises for cat, head and tail URL: https://www.workroom-productions.com/exercises-for-cat-head-and-tail/ Last updated: 2025-06-19T13:52:23.000Z If we're live, go to and pick an environment. Read about these tools at [cat, head and tail for Testers](https://www.workroom-productions.com/cat-head-and-tail-for-testers/) ## Exercise 1 – exchange experiences Let's talk: What have we seen `cat` `head` and `tail` used for in testing? ## Exercise 2 – demo `cat` , `head` and `tail` to ourselves Open the terminal, and 1) Try this to pick out all the test definitions in a test file `cat ~/code/rs_py/tests/test_* | grep "def test_"` and consider why you might do this rather than `grep "def test_" ~/code/rs_py/tests/test_*` 2) try this to see the tops of all the test files in a directory: `head ~/code/rs_py/tests/test_*` 3) Try this to see the access log for the web server – you'll see it change as the page you're working in makes requests of the server. `sudo tail -F /var/log/nginx/access.log` You need `sudo` to get system access. `-F` follows the log, updating as it changes. 4) `cat` will do some transforms – and especially `cat -tve` will sanitise otherwise unprintable data. Compare `head -n 2 /var/log/wtmp` and `head -n 2 /var/log/wtmp | cat -tve` ## Exercise 3 – oddities The stream \` `/dev/urandom`\` is full of randomness, so `head -c 5 /dev/urandom` returns [stuff](https://en.wikipedia.org/wiki//dev/random). Try it – then try `tail` ... . Try `cat` if you're brave. *What's happening here?* Open two terminals, or a terminal and an editor.On the commandline (of a terminal), `tail -F whatever.txt`. In the other terminal, or in the editor, open `whatever.txt` , add text, save, check the terminal, repeat. *What's happening here?* Just type `cat`. *What's happening here?* ### cat, head and tail for Testers URL: https://www.workroom-productions.com/cat-head-and-tail-for-testers/ Last updated: 2025-06-19T13:52:06.000Z *Ubiquitous commandline tools. Text unfinished, still useful.* These three small tools give you back what you give them, line by line. `head` gives you the top 10 lines, `tail` gives you the bottom 10 lines. And `cat` gives you all the lines in a great waterfall of stuff. We'll come back to `cat`. The context these are initially useful in is the context of handling text files without a neatly scrollable editor, and possibly without the memory to actually hold the whole file. Want to see what's in `my_vast_list.csv`? Use `head my_vast_list.csv` and you can get a clue without reading the whole thing. Want to see the most recent lines glommed on the end of `awful.log`? Use `tail awful.log` and there they are without needing to load the file or scroll through it. You can use `cat this.txt that.txt` to lump two files together: `cat` is made to *concatenate* files. But you'll more often find it used in the wild with just one input. That's because concatenating a file with nothing does nothing to the file, so `cat` hands it back to you, line by line – the file has become a stream. Streams are managed line-by-line, which is memory-efficient. Work scales with more time, rather than more hardware. Unix is very stream-y – all the connections between tiny tools with `|`and `>` are built on streams. `cat` opens the door to processing files with the power of those commandline tools. ## `cat` to stream ## *I keep rewriting this – the basic summary is that `cat` is handy enough to be used all over the place, but that plenty of that use is not necessary good use!* `cat`'s broad usefulness is that it allows something to be processed line by line. Why write `for i = 0 to lines_in(file); do something to line(i); append that to output` if you can write `cat file | something > output`. However, this ubiquitousness also leads to [Useless Use of cat](https://stackoverflow.com/questions/11710552/useless-use-of-cat). You might instead write `something < file >> output`. Here are some examples: - use `cat myresults.md | grep "FAIL"` to pick out all the lines marked "fail" (or use `grep "FAIL" myresults.md` - use `cat myresults.md | grep "FAIL" > fails.md` to put those lines in a file - use `cat test_whatever.py | grep "test_"` to pick out also - use `cat stuff > /tmp/before.txt` to copy a file - AVOID THIS: To swiftly create a file without bothering with an editor, `cat > myfile.txt`, type away (you're putting a stream into `cat` and it's streaming it into a file) and hit ctrl+d when done. It's neat, but it overwrites `myfile.text` irrevocably and I've kicked myself for doing it. However... - `cat` is a regularly used in "[here-documents](https://en.wikipedia.org/wiki/Here%5Fdocument)", which I see more commonly using LLMs than I did in my time doing ops and environments. Look out for `cat < *something*`, followed by lines with an `EOF` at the end. ## `cat` tricks While it has less-used features that can be very handy, - Number the lines with `cat -n boogaloo.txt`. Number just the non-blank ones with `cat -b boogaloo.txt`. - Skip multiple blank lines with `cat -s boogaloo.txt` – and you can combine that with `-n` or `-b` for counts. - Some characters found in files can't be displayed in a terminal – at best your days gets choppy or beepy, at worst your setup is changed in an unspeakable and unfindable way. Pipe output like this `nastyout | cat -tve` to make it sane. - Split that up to look for invisible problems in text. In `cat -tve` above, `-t` shows tabs as `^I`, `-e` shows line-ends as `$` and `-v` shows non-printing characters in a printable (but obscure) way. Use it to spot places where tabs should be spaces, where line-ends aren't where you want, or where some sossidge has furkled your data. - `cat` is made to *concatenate* files. It does the job well. Use it as `cat me.txt you.txt` to fill your terminal, or pipe it somewhere useful `cat us.txt them.txt > everyone.txt` ## Using and abusing `cat` - `cat` regularly pours gigabytes of text into my terminal, typically when I don't know what I'm looking at. `cat *whatever* | less`lets me step through it using the `less` paginator. Or I could just `less *whatever*`, though by some unthinking habit I use `cat`. ## Shared behaviours - You can use `cat`, `head` and `tail` in combination with others, as input or as output. - You can use all these with file patterns - `head -n 3 *` will show you the tops of all the files (with a`==> filename <==` at the top of each). - `tail -n 5 *.log` will show you the last `5` lines in each file in the local directory ending `.log` - `cat *.log` will show the contents of all the files ending `.log`. If they all have the same serial info at the start, you can at a pinch `cat *.log | sort` to interleave them. ### Shared with `head` and `tail` - You can set the size of `head` and `tail` with `-n`, so `head -n 5 awful.log` shows the top 5 lines of `awful.log`. Both default to 10. - You can limit the number of characters output with `-c`: handy if producing something of a particular size. - You can use `-q`uiet to never show filenames (handy if passing to processing) and `-v`erbose to always show filenames (handy if reading and confused) ## Back to `head` and `tail` - You can aim the output of something into the input of one of these. If `grep diddley doo.txt` shows all lines containing `diddley`, then `grep diddley doo.txt | head -n 5` shows the first 5 `diddley`s. - You can `tail -f` a log – and as the log grows (from the bottom, naturally), that changes what you see. `-f` is for ***f***ollow. Use `-F` to retry: good when the file doesn't exist, or may vanish. - Want to work with something of a particular size? Use `head`/`tail` with `-n` for number of lines, `-c` for number of characters. - Want something inbetween two limits? Use `head -n 10 | tail -n 3` to see lines 7 to 10\. More precise – want just line 453? `head -n 453 | tail -n 1` - Just want to see the start of *somecommand*? Throw it to head `somecommand | head` . Just want the last lines of message? Throw it to tail `somecommand | tail`(?jlcheck this for long output). This gets used with `curl` and `wget` when grabbing stuff from a URL, with SQL commands when one hasn't bothered to restrict the size of the output, and more. ## Oddnesses Careful of using `-n` vs `-*somenumber*` i.e. `head -n 3` and `head -3`. They do the same, but `-3` is considered obsolete. Note that `head -3` and `head -n 3` may do the same thing, but `head -n +3` and `head +3` don't do the same things – on some systems I believe that `head -n +3` starts from the third line... on my Mac with a bash shell, it doesn't. I don't know how to do (say) "from the 5th line to the end" without piping a `head` to a `tail` and that means knowing the file length. Note that the man for `tail` says "**\-n** +NUM to skip NUM-1 lines at the start". I don't know what this means – at the start of *what*; the input file, or the result. ## `man` pages ## [cat(1) - Linux manual page![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/icon/faviconV2-4)Linux manual page![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/thumbnail/TLPI-front-cover-vsmall-3.png)](https://www.man7.org/linux/man-pages/man1/cat.1.html) [head(1) - Linux manual page![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/icon/faviconV2-3)Linux manual page![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/thumbnail/TLPI-front-cover-vsmall-2.png)](https://www.man7.org/linux/man-pages/man1/head.1.html) [tail(1) - Linux manual page![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/icon/faviconV2-2)Linux manual page![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/thumbnail/TLPI-front-cover-vsmall-1.png)](https://www.man7.org/linux/man-pages/man1/tail.1.html) --- because it does a simple thing: it writes out what you give it. If you give it *one* file, it lists it back for you. Think of these tools as acting on something as *rows* – you can give each a file, a bunch of files, or a stream. - `cat` shows (and can transform) all the rows, in order. - `head` shows rows from the top, in order - `tail` shows rows from the bottom, in order That's kind-of obvious for files, and that's not the only way to think of these tools. ### Workroom PlayTime 021: cat, head and tail URL: https://www.workroom-productions.com/playtime-021/ Last updated: 2025-06-19T13:53:02.000Z We'll gather on Zoom / Miro, on Thursday 19 June at 3pm London time ([local time for you](https://this-ti.me/?uts=1750341600&tz=Europe%2FLondon&name=Workroom+PlayTime+021)) to play together with testing stuff. We'll play with the command-line tools `cat`, `head` and `tail`. If you're on the email list, you'll know that I shipped the date before the content... links to follow. These exercises are for everyone, for free. Paying [Subscribers](https://www.workroom-productions.com/#/portal/signup) get to play together every week. This week's is open to subscribers and friends. Subscribers will see joining info below. Pop your name on the Miro board if you're coming. In other news: - Here's [Weird Word 2](https://www.workroom-productions.com/weird-word-2/) about an odd iOS moment – can we send a transcribed voice message about M&S? We can... but people on iOS devices won't receive it. - Here's a video 'Live at EuroSTAR' with Leandro Melendez (Señor Performo) and I talking about testing and playing and stuff. [#eurostar | 🤠Leandro Melendez (Señor Performo)Live from EuroSTAR Conference '25 in Edinburgh We have Señor Performo chatting with James Lyndsay about toy making and the importance of Playing=Working We are streaming live from #EuroSTAR 2025! Check: https://lnkd.in/dFtSUVdM![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/icon/al2o9zrvru7aqj8e1x2rzsrca-1)LinkedIn🤠Leandro Melendez (Señor Performo)![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/thumbnail/1749128816439)](https://www.linkedin.com/video/live/urn:li:ugcPost:7336377934074568705/) - Dates for the following will need to be in July, probably in the first week or two, as I don't see a letup to give me time to set up otherwise. 1) a session on Exploratory Interfaces 2) a re-run of the EuroSTAR session 3) a small group on those environments. Ping me if you want to put your name in for any of those. Cheers – James _This post is for subscribers only._ ### PlayTime 019: Puzzle36 URL: https://www.workroom-productions.com/playtime-019-puzzle36/ Last updated: 2025-06-12T09:38:02.000Z We'll gather on Zoom / Miro, on Thursday 12 June at 3pm London time [(local time for you](https://this-ti.me/?uts=1749736800&tz=Europe%2FLondon&name=Workroom+PlayTime+19)) to play with [Puzzle36](https://www.workroom-productions.com/puzzle-036/). I know last week's was serial number `020` – I accidentally skipped `019`. Using `019` is an indication that uniqueness is more useful to me here than ordering. It also appears that I abhor a gap. Not that you'd know that from my puzzle numbering, but that's another story... These exercises are for everyone, for free. Paying [Subscribers](https://www.workroom-productions.com/#/portal/signup) get to play together every week. This week's is open to subscribers and friends. Subscribers will see joining info below. Pop your name on the Miro board if you're coming. In other news: - Last week's double-length "deep dive" interactive session at [EuroSTAR](https://conference.eurostarsoftwaretesting.com/event/2025/test-driven-generation-a-hands-on-experience/) in Edinburgh was great: I had a full room of people generatingcode to pass tests. - I made a tool to set up around 120 environments for the workshop. My approach is reusable, built on open-source tools, and [here's a page that will shortly share methods and scripts](https://www.workroom-productions.com/ephemeral-environments/). I'm delighted that it worked. - I wrote about [what those environments cost to deliver](https://www.workroom-productions.com/workshop-costs/) – in terms of money for LLMs and servers, and also in terms of emissions and water use. - I ran last week's Workroom PlayTime from a table in the centre of the conference – worked well for conference-goers, less-well online. - I wrote about [what it cost](https://www.workroom-productions.com/workshop-costs/) to run the workshop in terms of money, power, water and emissions for the LLM use and environments. - I also wrote about my [experiences with different large language models](https://www.workroom-productions.com/different-llms-do-different-things/) for the workshop. - I'm now actively trying to find June (possibly July) dates for 1) a session on Exploratory Interfaces 2) a re-run of the EuroSTAR session 3) a small group on those environments. Ping me if you want to put your name in for any of those. And, from a week or two ago, some housekeeping... - I've hardly been publishing here, but I'm certainly writing. So I want to put the following topics in front of you: *Developing a testing aesthetic for systems that are trained to satisfy us* – *things LLMs are good at for testers* – *planning for clarity, not control* – *lessons (and tools) from factchecking* – *testing work that needs ephemeral tools* – *exploratory interfaces* – *ideas need marshalling, not generating* – *power laws and other distributions*. While I seem only to write productively when I feel the urge, if you urge me to polish one of these I'm more likely to publish it! - The program for [Agile Testing Days](https://agiletestingdays.com) is out. I'm delighted to reveal that I'll be doing several things; a scriptless high-wire [Testing Transparently](https://agiletestingdays.com/2025/session/testing-transparently/) keynote with Elizabeth Zagroba, a [Crafting Custom Tools](https://agiletestingdays.com/2025/session/the-art-of-crafting-your-custom-tools/) tutorial with Bart Knaack and Huib Schoots, and a followup [session](https://agiletestingdays.com/2025/session/quick-trick-automation/). - I'm actively seeking new clients for my [teaching](https://www.workroom-productions.com/teaching/) and [consulting](https://www.workroom-productions.com/consultancy/) practice – you could [drop a meeting into my diary](https://savvycal.com/workroomprds) if you'd like to explore some possibilities! Cheers – James _This post is for subscribers only._ ### Ephemeral Environments URL: https://www.workroom-productions.com/ephemeral-environments/ Last updated: 2025-06-11T19:58:19.000Z I've been working on something for people who teach testing in public. I want to share this, as I think it's useful. Shortly, I'll clean it up to make it open-source, and I'll teach people how to use it and set it up. In order to provide these environments, you'll need: A DigitalOcean account (for the server), a Cloudflare account (for DNS), Ansible on your local machine, a public/private key pair, a github repo or two, and a sense of what's possible. --- It's important to me that people can work in a real development environment. That means, to me, that they should: - Have a file system that they can use as they wish - Have a command line with no restrictions to let them take action - Have a file editor to let them make changes - Have a web server to let them see output in a 'nice' way (especially coverage) We can't expect people to use their own machines – their work machine is locked down, or they don't feel happy about installing novel tools. Or they've left it behind. Or it's a tablet. We can't expect them to install a load of tools or to rely on temporary licenses. We can't expect them to download loads of stuff over conference / hotel wifi – let alone install it in the first 30 minutes of a 45-minute talk. We can't expect *everyone* to download and install the tools beforehand. We can't do tech support, in the workshop, for all the kit in the room. --- Clearly, this is idealistic – there is *no* solution. Buut.... As a tester, I know something about orchestration. I can leverage that, and with the help of several LLMs and standing on the shoulders of vast and competent open-source projects, I have leveraged it into something I've used at conference scale. I'd like to share it with you. - It runs in a browser, needs no downloads, is conference wifi-friendly, offers everyone their own VSCode in the browser. - Through the relatively-familiar world of VSCode, it offers access to a user on a DigitalOcean server. That means that people get a browsable file system and a proper commandline. - The scripts set up whatever is wanted on the server. Currently, I've got bits to set up the comandline tool `llm` with plugins and keys to access various models, Python dev env with Pytest tests (and coverage) serving human-readable stuff over Flask, a Javascript dev env with Jest tests (and coverage) serving stuff over Nginx. - To make the environments accessible, each server has useful URL working over https with a viable certificate, a cover web page linking to 10 user pages, and each user page links to the VSCode browser and to a few web-served endpoints. - The scripts 1) set up a droplet image to be reused 2) make copies of that and configure them. Once you've got a droplet image from script (1), you can run (2) several times in parallel – they take 20-30 minutes to stand up an environment. --- Alternatives [GitHub CodespacesGitHub Codespaces gets you up and coding faster with fully configured, secure cloud development environments native to GitHub.![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/icon/pinned-octocat-093da3e6fa40-4.svg)GitHub![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/thumbnail/features-codespaces-social.jpg)](https://github.com/features/codespaces) [CodeSandbox: Instant Cloud Development EnvironmentsCodeSandbox is a cloud development platform that empowers developers to code, collaborate and ship projects of any size from any device in record time.![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/icon/favicon-6.ico)CodeSandbox![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/thumbnail/og.jpg)](https://codesandbox.io) ### Weird Word 2 URL: https://www.workroom-productions.com/weird-word-2/ Last updated: 2025-06-11T18:20:51.000Z *An observation that might tell us something about what lies beneath the tools we use directly.* Here's why you can't send a voice-to-text message to your iOS-using friend, from your iPhone, about M&Ms. [Cracking The Dave & Buster’s Anomaly | Rambo CodesGui Rambo writes about his coding and reverse engineering adventures.![](https://static.ghost.org/v5.0.0/images/link-icon.svg)Rambo Codes![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/thumbnail/opengraph.png)](https://rambo.codes/posts/2025-05-12-cracking-the-dave-and-busters-anomaly) What's interesting to me here is that this is working properly, but 'properly' highlights an incompatibility between the needs of a security tool and an AI and a corporate agreement to use tradenames in a brand-aware way. And this from the firm that got us to accept glyphs. One of the things I love about dysfunction is how much it tells us about how things are working. Did it ~~work~~ fail for me? Not quite: in the UK, saying "Dave and Busters" transcribes as "Dave and Busters", not "Dave & Buster's". However, "M&Ms" failed properly. I suspect there's a bit of UK centric brand-translating going on – and when I talked with Mrs Workroomprds (she's not actually Mrs Workroomprds – she kept her maiden name), she smartly suggested "M&S". And, indeed, if you 1) open messages 2) choose the option to send voice messages 3) say "let me tell you about M&S" and 4) hit send, your own device will neatly transcribe your voice and send it along with the audio – but the other end will not receive message or audio. If you're doing it for yourself, you don't want to be sending a message, or transcribing audio, but sending an audio message – and those are automatically transcribed. From: [Cracking The Dave & Buster’s AnomalyGuilherme Rambo reports on a weird iOS messages bug: The bug is that, if you try to send an audio message using the Messages app to someone who’s also using …![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/icon/favicon-5.ico)Simon Willison’s WeblogSimon Willison](https://simonwillison.net/2025/Jun/5/cracking-the-dave-busters-anomaly/) which led to [Cracking The Dave & Buster’s Anomaly | Rambo CodesGui Rambo writes about his coding and reverse engineering adventures.![](https://static.ghost.org/v5.0.0/images/link-icon.svg)Rambo Codes![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/thumbnail/opengraph-1.png)](https://rambo.codes/posts/2025-05-12-cracking-the-dave-and-busters-anomaly) which referenced [The Dave and Busters AnomalyA small group of Americans becomes convinced they’ve discovered something strange about their iPhones: a forbidden phrase the phone will refuse to transmit. A crack…![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/icon/d622c860429d0e3dda437ba1e219c6d7.png)Search Engine![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/thumbnail/MmWYC2UplkmMjBdbfITkuTnXTLNVS88CgNwhLLBeUFo-)](https://www.searchengine.show/the-dave-and-busters-anomaly/) ### Different LLMs do Different Things URL: https://www.workroom-productions.com/different-llms-do-different-things/ Last updated: 2025-06-20T20:47:55.000Z [Bart](https://agiletestingdays.com/speaker/bart-knaack/) and I built a tool and a [workshop](https://www.workroom-productions.com/guiding-ai-code-with-tests-workshop/) to let testers experiment with generating code from tests. I built the tool to use Claude 3.5 Sonnet, and used [Simon Willison’s llm library](https://llm.datasette.io/en/stable/) to exchange messages with it. I picked Claude 3.5 Sonnet because it produced code that ran when I asked for code – but did I *need* it? I already have everything I need to give me a first-pass answer. `llm` lets me switch models while keeping everything else much the same. Indeed, that's one of the key reasons I used it. I tried switching models; this is a summary of what I found. *Can I substitute GPT-4, 4o-mini or Qwen3 for Claude3.5Sonnet? tl;dr: Nope.* ## Qwen3 – local, competent, recidivist I rejoiced when I found I could run Qwen3 on my own machine: I could aim my code towards a coding capable LLM without going off to an intermittent and pricy provider. I've got Ollama on my 32Gb M2Max, and use [msty](https://msty.app) to manage and work with models. Qwen3 runs *fine*. Qwen3 writes code with the right syntax, fits the tests, finds viable alternatives, and surprised me by being faster locally than any competent LLM in the cloud. It's the 8B variant, so has a context window of 128K tokens. When I put `/no_think` in the prompt, it reliably offers output without reasoning. I use it regularly for general coding queries going via [msty](https://msty.app) if I'm on a train or fancy a change. But Qwen3 never passed all the tests, ever – it would fix one bug, and make another. And I couldn't share it easily in the workshop. I tried local models DeepSeekR1 and Llama3.2, got code that didn't run, and stopped trying. ## 4o-mini – special cases, blind to structure I messed about with 4o-mini. It did two things badly enough to stop me pretty soon: - One suite of tests checks a bunch of dates. They're all Easter Sunday, and that's not mentioned explicitly. Three times out of four, Claude and GPT-4 recognised that the dates were from a pattern, and gave me a complex algorithm that reflected the pattern, typically naming it with something close to one of the accepted ways of calculating Easter. 4o-mini always gave me back code that responded to each specific date; there was no generalisation, not algorithm. - I gave it a test (and hints) to assert that its proposed python code could be used as a module in a package. It never once wrote code that was a module, failing over and over again in the same way – so none of the other tests even ran. ## GPT-4 – monotonic, small context I tried GPT-4\. It wrote code that passed most of the tests – but when the script passed back the failing tests and asked for corrections, GPT-4 gave me new code that failed most of the same tests in the same way. In particular, it kept getting module / function signatures wrong. GPT-4 also failed in a way that would have been entirely comprehensible in early 2024, but felt odd in mid 2025: it said “no” because I’d asked for too much. GPT-4's 'context length' (the amount of input that it will pay attention to) is [8,192 tokens](https://openai.com/index/gpt-4-research/). The second iteration within a conversation typically output not code, but a message along the lines of: «This model's maximum context length is 8,192 tokens. However, your messages resulted in *8570* tokens.». I typically let a 'conversation' run for three tries. If I want my tool to use GPT-4, I either need smaller tests / code / failures / rules, or I need to duck out of the conversation after fewer tries. ## Perils of 'thinking' models 'Thinking' models get in the way. When you’re asking for code, only code, no fences or braces or explanations, you’re fundamentally frustrated by a model that starts replies with ‘So what I think the person is asking for is…’, or returns several hunks of neatly-fenced code interspersed with explanations. I tried a couple of approaches; asking the thing not to think (as in the `/no_think` instruction to Qwen3), asking for just the output, parsing the output to pick up only code delimited by three backticks. I typically spent more time wrangling the output that I wanted, and went back – perhaps temporarily – to something that was reliable. Aside: If you're used to [copilot](https://copilot.microsoft.com/) or [cody](https://sourcegraph.com/cody) or [cursor](https://www.cursor.com), you might notice that those tools do a special (and fragile) thing for you; they insert changed code into the right place. My tools don't have those smarts – they generate one whole file and substitute the lot. ## Claude 3.5 Sonnet Let's recognise a bias: I started with Claude 3.5 Sonnet, I've used it the most, and I continue to use it. There is every chance that I've built the workshop's prompts and processes to this model's strengths and managed its weaknesses. Switching my choice of Large Language Model in one place in the code changes none of the infrastructure implied by the rest of my tools. Still... I've seen Claude 3.5 Sonnet change architecture on the basis of one test, switching from basic to packaged Python, and from ESM (uses `import`) to CJS (uses `require`) Javascript. I've see it change fundamental approaches as its first approach fails to pass tests, progressing from no code to 3 failing tests to 1 failing test, back to 8 failures – then to code that passes the lot on the next attempt. I've seen it ticktock between code that fails one or other test, then find a way through, but I've rarely seen it fail the same test over and over again without the test itself being problematic. I've seen it merrily run through 8 consecutive attempts in one conversation without touching the edges of its massive [200K context window](https://docs.anthropic.com/en/docs/about-claude/models/overview#model-comparison-table). Claude 3.7 seems to keep those qualities (if one ensures thinking is off). Its output context window is 64K (to 3.5's 8K) so I'd switch for larger outputs. I've not yet tried Claude 4. #### An alternative comparison .. I used [Msty](https://msty.app) to run the same prompt simultaneously on these same LLMs: a query about [part of my local area](https://www.lettsomgardens.org.uk/?page=2) that I know well. This was interesting, so I made an [exercise](https://www.workroom-productions.com/llm-comparison-local-knowledge/) for you to play along. \> What do you know about Lettsom Gardens and the surrounding area in London? Let's remember that LLMs are making it all up, all of the time – but sometimes their fantasy is directed and constrained by their training, system prompts, what they've retrieved or what you've recently mentioned, and more. Note: none of these were doing retrieval from the web. All were fresh conversations. The local models take up only 2-4 Gb – treat whatever they gave me as the product of **very* lossy compression. I don't want to share their outputs publicly, because I don't want to pollute the data, so I'll[ put them behind the paywall of the site](https://www.workroom-productions.com/llm-comparison-local-knowledge/). In summary: - Local models (as in local to my laptop, not trained in local knowledge) ****Qwen3** and ****DeepSeekR1** entirely made up almost every detail of history and features. In terms of geography, they placed it in the wrongly, and were geographically illiterate about London. **Llama 3.2** made excuses, offered something that it indicated was tentative and tangential, and stopped. Points to Llama 3.2! - ****GPT4o-mini** located it within a couple of miles, got the right reason for the name in §1, then imagined a handful of features and described the (wrong) neighbourhood. - ****Claude 3.7 Sonnet** got the location right, and included 3x as many checkable facts as 4o-mini. Almost all of those facts match my local knowledge. - ****GPT-4** got the location and several facts right, made up a pond and a hardwood forest, and spent the rest of its short answer extolling mostly-vague virtues of nearby Camberwell. ## Conclusion Here's what I needed from the LLM for my workshop, and ways I pushed when the LLMs didn't satisfy those needs. - a large context window so that I could give it plenty to chew on – I could with some parsing have shrunk my requests, but GPT4's 8K is no match for Claude's 200K and Qwen's 128K. - easily wrangle-able output – I could have managed this with more parsing - enough variation in its patterns that it could offer a variety of approaches – I could have managed this (perhaps) by tuning the 'temperature' (i.e. randomness) of the models in my request - enough coding (and testing) patterns that it could infer code from tests – I could offer more-explicit hints via names, comments and rules files. - a publicly-accessible API with a plugin for the `llm` tool – I wondered about running something [open-source on replicate](https://replicate.com) to avoid sharing API keys. Claude 3.5 Sonnet is good-enough for the workshop – but 3.7 aside, I've not yet found a viable alternative. ### Workshop Costs URL: https://www.workroom-productions.com/workshop-costs/ Last updated: 2025-06-10T11:12:27.000Z **tl;dr – less than a coffee morning** *Costs for the API and environments (under $20), emissions (<1kg CO2), power (1.5kWh) water (5.5l)* On June 4, I ran a longish workshop where I gave a big crowd of testers unlimited access to my Anthropic and ChatGPT API keys. Each participant had a chunk of a server with pre-installed software, VSCode for the web, Simon Willison's LLM, webpages served via Flask and Nginx, and a handful of Python and JavaScript tools. It took *months* to set up, but people want to know what it cost to deliver. These are good questions to answer. Here are my thoughts: ## Imagining limitations I wanted to be virtual, because I didn't want anyone downloading or installing anything. I wanted to be low-bandwidth, because conference wifi. I chose to use [VSCode in the Browser](https://github.com/coder/code-server), via DigitalOcean servers, set up via Ansible before the workshop. My DO account lets me run around 50 servers; experimenting earlier indicated that a reasonable sized server stayed responsive with 5-8 people, and seemed to have headroom to go to rather more. Our space had 11 tables. I gave each a group: each group shared a server; each server was set up to have 10 environments. The API keys for the workshop were all from my Anthropic / OpenAI accounts, which imposed different limitations. Participants (mostly) used Anthropic's Claude 3.5 Sonnet. My account has 1000 requests / min and 80K input tokens / 16K output (from [Rate limits](https://docs.anthropic.com/en/api/rate-limits#tier-2)) for Claude 3.5\. I thought we'd not be likely to get to the request limit, but it seemed likely that we'd hit the token limit. Prior to the workshop, I reckoned a typical request in my workshop would have 2000-2500 input tokens and \~400 output tokens (from [by token-calculator.net](https://token-calculator.net/)) So, each minute, my workshop might squeeze 40 input requests from 80K input tokens, and 40 outputs from 16K output tokens. EuroSTAR told me we'd be 88 people. They were running scripts, not making requests, and those scripts would make up to three requests, typically over a 60-90 seconds. If everyone set off at the same time, that might mean an expected peak around 200 requests in a minute, and a less-likely peak somewhere north of that. I needed to spread out the peaks somehow, or guide people to expect failure – I had a spiel, and a trick with a timer. ## Actual limitations I ultimately forgot to do either. I went from table to table throughout the workshop, and only one or two people mentioned that they had seen a limit. All the groups seemed to get stuck in. After the workshop, I saw that over 400 requests had been limited; 366 inputs and 47 outputs. My participants may not have expected failure, but they did *accept* failure. ## LLM Costs Sonnet 3.5 is $3/M in, $15/M out; a 2000-token request is ¢0.6, and a 400 token output is ¢0.6\. One participant wrote: > having attempts linked to tokens with a tangible cost makes me feel more reticent to keep spamming the remake command. (or it would if I were the one paying ;) ) According to usage stats, my Anthropic account used about 2.8M input tokens and maybe 0.4M output tokens on June 4\. That's about $15 dollars total. Anthropic billed me $13 for Jun04 – let's call it a tenner. For those of you in the room, most of the groups used $1-2 of Claude, but group05, 07 and 10 used $0.5 or less, while group09 used $2.11\. The workshop had access to OpenAI, but GPT4o-mini was poor at rewriting code to pass tests, and GPT-4 was s l o w. Total OpenAI usage seemed to be around 50 requests across 8 groups, using 0.4M tokens. Thirsty group09 had their share, but thrifty groups05, 07 and 10 didn't seem to make a dent. Perhaps their participants used their own tokens... Workshop participants will remember that we ran out of (Anthropic) tokens after an hour or so. Handily, that was easy to find out about, and fast to fix. It wouldn't have happened at all if I'd remembered to top up the account beforehand, or had published the realtime cost on screen as planned. ## LLM Consumption Reading [How Hungry is AI? Benchmarking Energy, Water, and Carbon Footprint of LLM Inference](https://arxiv.org/abs/2505.09598), I find that Claude 3.7 Sonnet consumes 2.7 Wh for a query with 1k tokens input / 1k output, and 5.5Wh for 10k in / 1.5k out. From the same paper, it consumes around 20ml water and emits around 2.5g CO2 for the larger query. Taking those values and scaling (by 2.8e6/10e3 =280 ) one might imagine that the workshop consumed 5.5 \* 280 = 1.5kWh in energy and 5.6l water, emitting 700g CO2\. Let's put that into personal terms: That's as much liquid as two big supermarket bottles of milk, as much energy as is needed to boil water for 70 cups of tea, as much carbon as is emitted on a swift car ride to the shop. With domestic electricity in the UK around 25p/kWh, perhaps 40p of my £10 LLM cost is power – though only Anthropic knows what the actual bill might be. This of course ignores the vast costs of training the models in the first place, the copyright theft needed to get the language models good-enough to be sensible (ingesting open-source software without language isn't enough), the risks inherent in training the lunatic chatbots to make us satisfied, and far more that I don't (yet) comprehend. ## Environments My environments were DigtalOcean droplets. They're around 4¢ / hour. I had 15 of them, and they ran for around 8 hours each – that's $5 or so. ### PlayTime 020: Code is Cheap: Tests are Valuable (at EuroSTAR25) URL: https://www.workroom-productions.com/playtime-020-code-is-cheap-tests-are-valuable-at-eurostar25/ Last updated: 2025-06-05T11:37:55.000Z These exercises are for everyone, for free. Paying [Subscribers](https://www.workroom-productions.com/#/portal/signup) get to play together every week. This week's is open to subscribers and friends. Subscribers will see joining info below. Pop your name on the Miro board if you're coming. We're at 3:30, Edinburgh time. Here's the [local time for you](https://this-ti.me/?uts=1749133800&tz=Europe%2FLondon&name=Workroom+PlayTime+020). We've got someting that can generate code from tests, and will guarantee to pass the tests. The idea of the exercise is to throw away the code, remove a test, re-generate the code and see if we can detect which test has been removed by exploring. Then we'll throw the code away again, damage the tests further, and try again. We'll use [cic.workroomprds.com](https://cic.workroomprds.com/) There are 10 names – we'll pick unique ones, follow the link, and open two tabs; one for *code-server* (so we can change the tests) and one for *rs\_js* (so we can play with what we'e made). ![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/2025/06/image-2-1.png) In the code-server tab, your password is `password`. Click through the setup. If asked, open the repository, trust the code. Open the terminal and the file browser (top-right, two icons, hit the one with a side bar split, and the one with a bottom-bar split), switch to "terminal" in the bottom. Change to the right directory – you need to be in `/code/rs_js` . In the terminal, `cd ~/code/rs_js`. Note: `~` works as if it is `/home/«your chosen name»`. Activate the python venv: `source ~/llm-env/bin/activate` We'll stop here until everyone is sorted... To edit the tests: edit `test/jest/relativeSizes.test.js` in the editor and save. If there's a dot by the name, it's not saved and won't be picked up. To delete the code: right-click `relativeSizes.test.js` and pick delete / move to bin from the menu To re-generate the code: in the terminal, check that the venv is active (`(llm-env)` on the commandline) and run `./makeNewJSFromTests.sh relativeSizes.js` To pick up the new code – refresh your browser page "from origin" if at all possible. We can all try each others's changes. We'll start by working with our own and collaborating. Cheers – James _This post is for subscribers only._ ### PlayTime 018: Raster Reveal URL: https://www.workroom-productions.com/playtime-018-raster-reveal/ Last updated: 2025-06-06T07:58:34.000Z We'll gather on Zoom / Miro, on Thursday 29 May at 3pm London time ([local time for you](https://this-ti.me/?uts=1748527200&tz=Europe%2FLondon&name=Workroom+PlayTime+018)) to run [RasterReveal](https://www.workroom-productions.com/raster-reveal/). These exercises are for everyone, for free. Paying [Subscribers](https://www.workroom-productions.com/#/portal/signup) get to play together every week. This week's is open to subscribers and friends. Subscribers will see joining info below. Pop your name on the Miro board if you're coming. [*RasterReveal*](https://www.workroom-productions.com/raster-reveal/) gives you a picture to play with. As you play with it, you can't help but imagine what's in the picture – and that influences your play. You'll find that you decide what to do next without the need for a distinct decision. It's interesting, as testers, to feel that happening, and to see what happens in your mind as the information coalesces, or as it disperses into mist. It's interesting to feel confirmation and refutation. The exercise allows that, without the cognitive load of testing. You can probably tell that it's one of my favourites – and I'm delighted to say that it [turned up kind-of accidentally](https://www.workroom-productions.com/making-the-raster-reveal-exercise/). In other news: - I'm at [EuroSTAR](https://conference.eurostarsoftwaretesting.com/event/2025/test-driven-generation-a-hands-on-experience/) in Edinburgh next week, running a double-length "deep dive" interactive session where we'll *generate* code that passes tests – and we'll consider the weirdness that results. Come say Hello, if you're there. I'm so pleased to be going to EuroSTAR! As an experiment, I'll run next week's Workroom PlayTime from the conference – we'll do a short chunk of my deep dive into exploring generated (and checked) code. - I've hardly been publishing here, but I'm certainly writing. So I want to put the following topics in front of you: *Developing a testing aesthetic for systems that are trained to satisfy us* – *things LLMs are good at for testers* – p*lanning for clarity, not control* – *lessons (and tools) from factchecking* – *testing work that needs ephemeral tools* – *exploratory interfaces* – *ideas need marshalling, not generating* – *power laws and other distributions*. While I seem only to write productively when I feel the urge, if you urge me to polish one of these I'm more likely to publish it! - The program for [Agile Testing Days](https://agiletestingdays.com) is out. I'm delighted to reveal that I'll be doing several things; a scriptless high-wire [Testing Transparently](https://agiletestingdays.com/2025/session/testing-transparently/) keynote with Elizabeth Zagroba, a [Crafting Custom Tools](https://agiletestingdays.com/2025/session/the-art-of-crafting-your-custom-tools/) tutorial with Bart Knaack and Huib Schoots, and a followup [session](https://agiletestingdays.com/2025/session/quick-trick-automation/). - I'm actively seeking new clients for my [teaching](https://www.workroom-productions.com/teaching/) and [consulting](https://www.workroom-productions.com/consultancy/) practice – you could [drop a meeting into my diary](https://savvycal.com/workroomprds) if you'd like to explore some possibilities! Cheers – James _This post is for subscribers only._ ### PlayTime 017: Whatever Next? 2 URL: https://www.workroom-productions.com/playtime-017-whatever-next-2/ Last updated: 2025-05-22T00:17:59.000Z This week, we'll run another playful exercise. We'll gather on Zoom / Miro, on Thursday 23 May at 3pm London time ([local time for you](https://this-ti.me/?uts=1747922400&tz=Europe%2FLondon&name=Workroom+PlayTime+17)). Joining info below. Pop your name on the Miro board if you're coming. These exercises are for everyone, for free. Paying [Subscribers](https://www.workroom-productions.com/#/portal/signup) get to play together every week. Cheers – James _This post is for subscribers only._ ### PlayTime 016: Whatever Next? URL: https://www.workroom-productions.com/playtime-016-whatever-next/ Last updated: 2025-05-14T14:38:19.000Z This week, we'll run a more playful exercise. We'll gather on Zoom / Miro, on Thursday 15 May at 3pm London time ([local time for you](https://this-ti.me/?uts=1747317600&tz=Europe%2FLondon&name=Workroom+PlayTime+016)). Joining info below. Pop your name on the Miro board if you're coming. These exercises are for everyone, for free. Paying [Subscribers](https://www.workroom-productions.com/#/portal/signup) get to play together every week. Cheers – James _This post is for subscribers only._ ### Playful Exercises about Testing URL: https://www.workroom-productions.com/playful-exercises-about-testing/ Last updated: 2026-06-11T14:22:04.000Z Ideas for exercises, in various states of completion. Crucially, they're based in ideas of collaborative play. This is an attempt to write exercises where no one has a pre-existing answer. They're primarily designed for online use, for a handful of playful testers, and should last \~20 mins. _This post is for subscribers only._ ### PlayTime 015: Bring me a Letter URL: https://www.workroom-productions.com/playtime-015-bring-me-a-letter/ Last updated: 2025-05-07T15:34:22.000Z The feedback exercise involves finding a strategy that allows you to reliably arrive at a goal, using examples and feedback. Then we'll see how that strategy survives problematic feedback. We'll gather on Zoom / Miro, on Thursday 8 May at 3pm London time ([local time for you](https://this-ti.me/?uts=1746712800&tz=Europe%2FLondon&name=Workroom+PlayTime+015)). Joining info below. Pop your name on the Miro board if you're coming. These exercises are for everyone, for free. Paying [Subscribers](https://www.workroom-productions.com/#/portal/signup) get to play together every week. Cheers – James _This post is for subscribers only._ ### Bring Me a Letter URL: https://www.workroom-productions.com/bring-me-a-letter/ Last updated: 2025-05-08T19:47:34.000Z This is an exercise about feedback – it's loosely inspired by Johanna Rothman's [Bring me a Rock](https://www.jrothman.com/newsletter/2002/01/volume-5-number-2-project-toolbox-recognizing-the-bring-me-a-rock-schedule-game/). The machine will want something. You can give it an example, from your limited collection. It will tell you whether your example satisfies its need – and how your example compares with what it wants. It takes a moment to get feedback on the comparison, and it takes a little longer before it will accept another example. There's no guarantee that the request can be satisfied by what you've currently got. You can make a new collection of examples, and you're likely to need several new collections. You can make a new collection immediately. You'll need to come up with a strategy to reliably arrive at an acceptable example. That probably means you'll need to try a few approaches. We'll work on part I first, to develop that strategy. In part I, you *always* get honest feedback. Then we'll see how your strategy works when the feedback has specific problems in part II, 1-3\. And you can see whether you can detect the problems with the feedback in part III. For Workroom PlayTime, we'll spend 10 minutes to get to our strategies. Then we'll share, and either talk about how we got there, or try our strategy against a pathology. - Build a strategy by playing with Bring me a Letter [part I](https://exercises.workroomprds.com/bringmeathing/bmarp1.html) - See how it works when there are problems with feedback: Bring me a Letter [part II - 1](https://exercises.workroomprds.com/bringmeathing/bmarp2-1.html), [part II - 2](https://exercises.workroomprds.com/bringmeathing/bmarp2-2.html), [part II - 3](https://exercises.workroomprds.com/bringmeathing/bmarp2-3.html) - Can you detect the problem with the feedback? Bring me a Letter [part III](https://exercises.workroomprds.com/bringmeathing/bmarp3.html) A key role of testing is to provide feedback on the system that is available to test. There's more on feedback elsewhere... I've written that feedback needs to be swift, relevant and true. I'll add links... later. ### PlayTime 014: Puzzle 33 URL: https://www.workroom-productions.com/playtime-014-puzzle-33/ Last updated: 2025-05-07T21:49:43.000Z As with most of the BlackBox Puzzles, [Puzzle033](https://blackboxpuzzles.workroomprds.com/puzzle33/)'s behaviour reflects a system that can be described in a sentence. Our exercise will be to explore, to make a model of what we find, and to test our models. We'll share our work, compare our approaches, and see where we get. We'll gather on Zoom / Miro, on Thursday 1 May at 2:30pm London time ([local time for you](https://this-ti.me/?uts=1746106200&tz=Europe%2FLondon&name=Workroom+PlayTime+014)). Joining info below. Pop your name on the Miro board if you're coming. These exercises are for everyone, for free. Paying [Subscribers](https://www.workroom-productions.com/#/portal/signup) get to play together every week. Cheers – James _This post is for subscribers only._ ### Clearing Cloudflare's Cache URL: https://www.workroom-productions.com/clearing-cloudflares-cache/ Last updated: 2025-04-28T19:33:10.000Z *Swift note, more for me than for you. But still for you. As a tryout, I thought I'd read it out loud. Let me know if that's nice.* Clearing Cloudflare's Cache 0:00 /470.64 1× I was working on a web page [https://exercises.workroomprds.com/comparing\_generated\_code/D/](https://exercises.workroomprds.com/comparing%5Fgenerated%5Fcode/D/src/js/main.js), and found that ~~hacks on~~ changes to the underlying javascript [src/js/main.js](https://exercises.workroomprds.com/comparing%5Fgenerated%5Fcode/D/src/js/main.js) didn't seem to stick. Initially, I noticed because the app on the page gave me an error message, right there in the output, and the app wasn't usefully responsive. Digging into the error, I saw that a necessary function was being called without parameters – and it was doing that because I'd uploaded an old version of some code. Or, rather, I'd uploaded that version to the wrong place. It was an evening of mistakes. I re-uploaded the right file (which might, just might, have had a timestamp before that of the wrong one...). The problem persisted. I looked a the file via the browser's devtools, and saw the old version. I looked at the file via ftp, and found the new version. I'd not made another mistake in my upload. I requested the file directly via its url in the browser, and got an old version. I requested the file with a suffix `?a=1` and got the new version. I suspected a cache. I looked at the behaviour on a different device, and on a different browser, and saw that the error persisted. I looked at the file via devtools on that device, and saw the old version of code. I suspected a non-local cache. I changed the time on the file / re-uploaded the file. Nothing made an immediate difference. I wondered what cacheing was on, that didn't pay attention to file changes. I took stock. There are several caches – likely more than I imagine. Caches are obscure by design, and change over time, so this would need to be an opportunistic search, rather than a painstaking process. I considered: - a local cache (in my browser, or on my machine). I looked here first, and stopped considering it seriously when I saw seen the error persist when on a different machine. - a server-side cache (this site is served by apache). I looked here second, and put it to one side because cacheing is complex and I thought I'd try to discount another before looking here. - a CDN cache (this site goes through Cloudflare). I looked here next. I've not looked at Cloudflare's cacheing before: I was interested by the novelty and wondered what it might do to make things simple. On Cloudflare's site, I found cacheing information when I went to look at the configuration for `workroomprds.com`. Within *configuration* I saw *Purge Cache.^* ![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/2025/04/image-1.png) I chose *Custom Purge*, entered the problem URL (you'll see that in my 'recently purged' below), and waited a few seconds as recommended. ![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/2025/04/image-2.png) The problem persisted. I did the same again, and the problem was gone. That may, of course, be a coincidence. The cached file might have been replaced without any direct action by me – and the cache that needed a nudge might not have been Cloudflare's cache. But now I have a note on how to purge Cloudflare's cache. --- Why is this important to me? I have workshops where code can change, on this server, *during* my workshop. I want participants to have a friction-free experience. I don't want to be faffing about with caches, or with unnecessary problems, when I'm trying to give my attention to a room full of people. So I want to give some thought to all the caches that might banjax the experience for me and for them, if I insist on having code that can change while we work in the workshops. Also... I want testers to understand more about cacheing, and I want to give them something to play with. I don't want them to play with my site, but perhaps I can think of something. --- Pottering further through Cloudflare's cache configuration for `workroomprds.com`, I notice that I turned this *off* in 2016\. ![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/2025/04/image-3.png) Perhaps I should turn it *on* for workshops where the code changes. And then manage the server cache and any triggers to local caches... ### PlayTime 013: Three Things URL: https://www.workroom-productions.com/playtime-013-three-things/ Last updated: 2025-04-17T17:33:43.000Z In Workroom PlayTime 013, we'll be playing with three small systems. All pass the same tests – all have different code. We'll spend 20 minutes on it. You need a browser (you've got a browser). We'll gather on Zoom / Miro, on Thursday 24 April at 3:00pm London time ([local time for you](https://this-ti.me/?uts=1745503200&tz=Europe%2FLondon&name=Workroom+PlayTime+013)). Pop your name on the Miro board if you're coming. These exercises are for everyone, for free. Paying [Subscribers](https://www.workroom-productions.com/#/portal/signup) get to play together every week. Cheers – James _This post is for subscribers only._ ### Exercise: Three Things URL: https://www.workroom-productions.com/three-things/ Last updated: 2025-04-23T21:37:37.000Z *To be used as a 20 minute exercise in* [*Workroom PlayTime 013*](https://www.workroom-productions.com/playtime-013-three-things/) I built the same thing three times. > The small system, written as a web page, takes input in the form of an `inputValue` and a `unit`, and a `scale` to contextualise the number and the unit, and converts that number to a `textOutput` expressing that number as something else in the scale, typically of the right size to be comprehensible. Typical scales are time and distance. Starting with the same tests, (mostly) the same instructions, and the same names for the target files, I asked an LLM to iterate until the code it put into the files passed the tests. Here they are, each with their tests. They all take the same [configuration](https://exercises.workroomprds.com/comparing%5Fgenerated%5Fcode/A/src/config/config.json), which sets out the units the things work with. - [Version C](https://exercises.workroomprds.com/comparing%5Fgenerated%5Fcode/C/index.html), [instructions](https://exercises.workroomprds.com/comparing%5Fgenerated%5Fcode/C/rules.md) and [tests](https://exercises.workroomprds.com/comparing%5Fgenerated%5Fcode/C/test/alltests.html). - [Version D](https://exercises.workroomprds.com/comparing%5Fgenerated%5Fcode/D/index.html), [instructions](https://exercises.workroomprds.com/comparing%5Fgenerated%5Fcode/D/rules.md) and [tests](https://exercises.workroomprds.com/comparing%5Fgenerated%5Fcode/D/test/alltests.html). - [Version E](https://exercises.workroomprds.com/comparing%5Fgenerated%5Fcode/E/index.html), [instructions](https://exercises.workroomprds.com/comparing%5Fgenerated%5Fcode/E/rules.md) and [tests](https://exercises.workroomprds.com/comparing%5Fgenerated%5Fcode/E/test/alltests.html). #### Starting at C?? - [Version A](https://exercises.workroomprds.com/comparing%5Fgenerated%5Fcode/A/index.html), [instructions](https://exercises.workroomprds.com/comparing%5Fgenerated%5Fcode/A/rules.md) and [tests](https://exercises.workroomprds.com/comparing%5Fgenerated%5Fcode/A/test/alltests.html). - [Version B](https://exercises.workroomprds.com/comparing%5Fgenerated%5Fcode/B/index.html), [instructions](https://exercises.workroomprds.com/comparing%5Fgenerated%5Fcode/B/rules.md) and [tests](https://exercises.workroomprds.com/comparing%5Fgenerated%5Fcode/B/test/alltests.html). I'd have sworn all the tests passed, but A and B show failures in the HTML handler. So now you have D and E. ## Exercise 1 *10 minutes, microphones on* Pick one. Explore it. Whatever way you like – knowing that they have the same configuration and pass the same tests, but have different code and may hav differing behaviours. What looks wrong? Share. ## Exercise 2 *5 minutes, mics on* Pick another. Explore it, starting with what you saw that was 'wrong' in Ex 1. Differently wrong? What else looks wrong? Share. ## Review *5 minutes, shared notes in Miro* Write down one novel way you'd try if you came across this situation at work. Write down one artefact you wanted to use. Share any conclusions about testing these things which pass the tests --- ## Extension - Assess the tests - Look at the instructions - Write automation to see differences between C, D and E - Use `diff` or similar to see differences between the underlying code - Try it yourself! Get the test files into a directory. Ask your LLM of choice to make code to pass the tests in `test-whatever.js`, then give it feedback based on `show-test-whatever.html` . Give it `rules.md`, too – and whatever source you might imagine. You'll need to make (JavaScript) `main`, `htmlhandler`, `relativeSizes`, HTML page `index`. You'll ask for CSS `styles` if you like (there's no test, but you'll want the LLM to have read the rules, the index page and the htmlhandler. You can either use the units from here, or ask for a JSON `config` file to suit yourself (there's a test for that). Certainly change the rules, the tests and whatever else. ### PlayTime 012: Testing Decisions – Cost of Trouble URL: https://www.workroom-productions.com/playtime-012-testing-decisions-cost-of-trouble/ Last updated: 2025-04-17T13:56:14.000Z In this week's Workroom PlayTime, we'll be [looking at projected value](https://www.workroom-productions.com/exercise-cost-of-trouble/)~~cost~~, and how our testing decisions are shaped by the trouble we imagine we might find. It's the second in my collection of [exercises simulating testing to help us think about decisions](https://www.workroom-productions.com/tag/testing-decisions/). We'll spend 20 minutes on it. All you need is a browser. We'll gather on Zoom / Miro, on Thursday 3 April at 3:00pm London time ([local time for you](https://this-ti.me/?uts=1744898400&tz=Europe%2FLondon&name=Workroom+PlayTime+012)). Feel free to put your name on the Miro board if you're coming. These exercises are for everyone, for free. Paying [Subscribers](https://www.workroom-productions.com/#/portal/signup) get to play together every week. Cheers – James _This post is for subscribers only._ ### Exercise – Cost of Trouble URL: https://www.workroom-productions.com/exercise-cost-of-trouble/ Last updated: 2025-04-17T12:31:06.000Z *To be used in* [*Workroom PlayTime012*](https://www.workroom-productions.com/workroom-playtime/) *on 17 April.* *A 15-20 minute exploration via a simulation* #### Background Based on an older exercise, no longer interactive, posts here: [An Experiment with Probability](https://workroomprds.blogspot.com/2012/11/an-experiment-with-probability.html), [Broken Trucks](https://workroomprds.blogspot.com/2012/11/broken-trucks.html), [Models, lies and approximations](https://workroomprds.blogspot.com/2012/11/models-lies-and-approximations.html), [Enumeration hell](https://workroomprds.blogspot.com/2012/11/enumeration-hell.html), [Diversity matters, and here's why](https://workroomprds.blogspot.com/2012/11/diversity-matters-and-here-why.html), [Modelling super powers](https://workroomprds.blogspot.com/2012/12/modelling-super-powers.html) . Go there to read background – I'll bring it here in the next few weeks and update it. Each of these simulations have a similar core simplifications: - the **simulation* knows what can be found, and it knows what's likely to find each thing, but the **entities looking* don't know either of these things. - A searchable thing has a limited collection of independent things to be discovered. **For testing, think:* *there are only so many problems.* - Discovery is by chance. Different approaches to searching have different chances. A discoverable thing has a low chance of being discovered by any individual approach – but might have a better chance by another approach. *For testing, think:* **Performance testing will find different problems from usability testing.* - There's a limited budget – the more you spend, the more chances you have to find stuff. - Something stays found, but only counts the first time it is found. You, the user, can see and change the number of things that can be discovered, the makeup of the crew that looks for things, and the budget to spend looking. During the simulation, you can see the results of the simulation in terms of the work done, the discoveries found. You can pause and resume the work. You can increase and decrease the budget. Different simulations will throw up different results, and you'll want to keep track of those to see the variation. Background: a simulation works in chunks. In each chunk, all the explorers have a chance to discover a (random) set of discoverables (of fixed size in the simulation). A discoverable's chance of discovery is set by the explorer's active skill, and the chance of the discoverable being discovered by that skill. ## Priming Question ⁉️ How do you feel about spending, and stopping, in a situation where the things you discover can have wildly different values? We'll come back to this in the debrief ## Exercise 1: Variety of values *5 minutes playing and talking* Go to . In this simulation, the dots change colour and size when 'discovered'. > What can you say about the sizes in the simulation? The size represents the value of the discovery. In testing terms it is perhaps the projected cost of the problem, if it had been found in production. ## Exercise 2: Interpreting into testing *5 minutes mainly talking* > What do you think about the 'value' of what one finds, when one finds a problem in testing? ## Exercise 3: When to stop spending *5 mins playing and talking.* *Let's all work with *Ordinary – A* in the "looking at something with X things to find' selection box..* Press the ⎋ button to remake your test subject – it'll keep the parameters from the simulation (JL except the budget - aargh!) Press the ⏍ button to see the values of all the undiscovered things, all at once. Run the simulation several times. > how do you feel about the budget, and what you're leaving uncovered? 😉 Note that this simulation – because it is a simulation – lets you see information which you would would not typically know when discovering things; \* what can be found \* the value of finding those thing You'll see what can be found by pressing the ⏍ button, in the text starting with 'secret' and in the Value of Trouble graph text. 💡 These values in Ordinary A (and in Dense A) are roughly a "power law" distribution – imagine that 1 in 10 is worth 10, 1 in 100 is worth 100, 1 in 1000 is worth 1000 and so on. The values in the B sets follow a "log normal" distribution. This can also have large, rare values, but doesn't show quite the wildness of a power law. Both are common in nature, regularly identified in systems analysis and seem to be plausible distributions for the cost of software failure. In nature, the largest 'things' are often readily visible, or the distribution is limited by environmental factors. In bugs, the largest is often invisible and its cost can exceed the market capitalisation of the organisation which introduced the fault (think the [crowdstrike](https://en.wikipedia.org/wiki/2024%5FCrowdStrike-related%5FIT%5Foutages) problem in 2024) I'll put links to the research when I feel confident that I've got readable sources. _This post is for paying subscribers only._ ### No Workroom Playtime on 10 April URL: https://www.workroom-productions.com/no-workroom-playtime-on-10-april/ Last updated: 2025-04-09T23:34:32.000Z ...but for those of you wanting an explore or a play, here's a [new toy](https://exercises.workroomprds.com/comparing%5Fgenerated%5Fcode/). Actually, *three* new toys: You know I have scripts and prompts to[ generate code to pass tests](https://www.workroom-productions.com/guiding-ai-code-with-tests-workshop/). I set up a test harness with a bunch of tests for a unit converter, and fired up my newest version of the ~~kludge~~ magic loop three times. So that built three similar things to explore, with rather different code (and the same config), which all pass the same tests, none touched by human hands. And that's what's yours to play with. There's no [Workroom PlayTime](https://www.workroom-productions.com/workroom-playtime/) scheduled for this week. This is as planned, but not what I had remembered when I signed off, last week, with "see you next week". 😳 If you want to play with this together, know that the usual zoom room will be open, and I'll pop in if I can at 3pm London time. I may not be able to stay. See you on the 17th with a new exercise on testing decisions – and look out for an announcement coming soon about a longer session on exploratory interfaces. Cheers – James ## Joining info Use this [zoom](https://us02web.zoom.us/j/86905624247?pwd=YahJdjeLxLm0u8SaQTwpiW3HBcUObV.1) room. No Miro. ### PlayTime 011 – Playing with Sort URL: https://www.workroom-productions.com/playtime-011-playing-with-sort/ Last updated: 2025-04-03T13:41:36.000Z In this week's Workroom PlayTime, we'll be [playing with sort ](https://www.workroom-productions.com/playing-with-sort/)– the next in my collection of exercises around [command-line utilities for testers](https://www.workroom-productions.com/unix-tools-for-testers/). We'll spend 20 minutes on it. All you need is a browser. We'll gather on Zoom / Miro, on Thursday 3 April at 3:00pm London time ([local time for you](https://this-ti.me/?uts=1743688800&tz=Europe%2FLondon&name=Workroom+PlayTime+011)). Feel free to put your name on the Miro board if you're coming. These exercises are for everyone, for free. Paying [Subscribers](https://www.workroom-productions.com/#/portal/signup) get to play together every week. Cheers – James _This post is for subscribers only._ ### `sort` Exercises URL: https://www.workroom-productions.com/playing-with-sort/ Last updated: 2025-04-03T13:44:30.000Z Subscribers get to work together on these in Workroom PlayTime 011 on 3 April, 3pm London time, Zoom and Miro. Go to [Workroom PlayTime 011](https://www.workroom-productions.com/playtime-011-playing-with-sort/) for login etc. There is a (very draft) info page at ## Exercises #### Files `ex01` has one single-digit number per line, and is unsorted. `ex03` has data in comma-delimited columns. The first three columns are date, first name, surname. `ex04` contains some month names / abbreviations, in order as far as English months are concerned. `ex04a` contains variants. `ex05` – one number per line like `ex01`, and sometimes a letter. `ex06` contains randomly ordered numbers, with some duplicates, and `ex06b` contains a collection of numbers with some in `1E6` notation. Go to , pick a user, drop through to VSCode in the browser. The files to sort are in `~/sort_exercises`. We'll be working in the terminal, and you should see something like this at the bottom of your window: ![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/2025/04/image.png) ### Exercise 1: Basic use Type `cat ex01` on the command line to see the contents (or look using the file browser). - Type `sort ex01` to see the output on the command line. - Compare `sort ex01` with `sort -R ex01` and `sort -r ex01` 💡 The syntax is `sort «option(s)» «file(s)»` Sort can reverse with `-r` ... and randomise with `-R` ### Exercise 2: Plumbing - Compare `cat ex01 | sort` with `sort ex01` - Use `sort ex01 > output_of_ex02` to sort into a file called `output_of_ex02` - Use `sort ex01 | less` to open the output in a a file reader `less`. Use `q` to exit the editor. 💡 `sort` is all set up to be used with other commands. As a standalone tool, with real data, it is a bit unwieldy – it's **best* used with other tools. ### Exercise 3: Columns Testers need to work with complex data, and need a column sort. Use `sort -t, -k3,3 ex03` to sort it by surname Use `sort -t, -k2,2 ex03` to sort by first name. Use `sort -t, -k3,3 -k2,2 ex03` to sort by surname then first name, and compare with `sort -t, -k3,3 -k2,2r ex03` which reverse the sort of the first name. 💡 Plain `sort` compares whole lines, character by character. Columns need delimiters: `sort` uses space by default, and takes the `-t` option to change. Specify `-t,` to use commas and `-t$'\t'` to use tabs (probably). Use options twice to sort on two columns. Use modifiers to change the type of sort. 💡 Use `-k2,2` to specify a sort on your data's second column. Use `-k2,4` to sort on the second, third and fourth columns. If you specify `-k2` you'll sort on the second column and everything to the left. **It's weird, don't do it.* ### Exercise 4: Checking You can check if something is sorted with `sort -c` – which is handy if you're checking a sort for a test, or pre-qualifying some data. Use `sort -c` on any of the earlier files – note the error shows the line and the content of the first non-sorted entry. Use `sort -c ex04` to see that a problem is on line 2. Use `sort -Mc ex04` to see that the check changes if told to expect to sort *months*, and within that style of sort, it accepts varieties of abbreviation and case. 💡 Use `sort -c` to check whether data is sorted, in various types of sort. Options can stack This exercise produces not a lot of output – here's the contents of `ex04` for interest. ```text January Feb mar April dEcEmBeR ``` ### Exercise 5: Reducing Sort can throw away duplicates. This is handy to see what data is in use (i.e. if you want unique account numbers, a list of this sessions error messages), and is handier using a columns selection. - Compare `sort ex05` and `sort -u ex05` – what's thrown away? - Compare `sort -k1,1 ex05` and `sort -uk1,1 ex05` – what's lost now? - Weird one: Compare `sort -M ex04a` and `sort -Mu ex04a` – what month names are kept? 💡 option `-u` throws away duplicates 'duplicates' depends on the sort `u` goes at the start, `n` at the end, column stuff in the middle... ### Exercise 6: Problems and avoidances Use `sort ex06` to see a problem. Try `sort -n ex06` to avoid it. Try `sort -g ex06b` to see how that works... 💡 `sort`'s default is to sort by character. option `-n` sorts by value 💡 There are other options for other forms, including \* `-d` dictionary sort – good for names i.e.`O'Leary` and `New York`. \* `-f` caseless i.e. `a` before `B` before `c`. \* `-g` scientific numeric i.e. `1E-2` is sorted as `0.01` \* `-h` human numeric sorts `1` before `1K` before `1G` \* `-M` **English* month acronym sorts `jan` before `feb`. Testing: look out for the 'wrong' sort: it may only be revealed by novel data. Other systems may break when the 'wrong' sort is corrected. ### Exercise 7: Sort and merge Try `sort -g ex06 ex01 ex06b` ### Sources [Linux sort Command with Examples](https://phoenixnap.com/kb/linux-sort) Wikipedia [sort (Unix)](https://en.wikipedia.org/wiki/Sort%5F%28Unix%29) **Man pages** [sort(1) - Linux manual page](https://www.man7.org/linux/man-pages/man1/sort.1.html) --- Sprue below - not useful. _This post is for paying subscribers only._ ### PlayTime 010 URL: https://www.workroom-productions.com/playtime-010-2/ Last updated: 2025-03-27T14:51:13.000Z We'll gather on Zoom / Miro, on Thursday 27 March at 3:00pm London time ([local time for you](https://this-ti.me/?uts=1743087600&tz=Europe%2FLondon&name=Workroom+PlayTime+010)) for this weeks' [**Workroom PlayTime**](https://www.workroom-productions.com/workroom-playtime/), which is [*Testing Decisions – Power of Variety*](https://www.workroom-productions.com/power-of-variety/). Feel free to put your name on the Miro board if you're coming. This session will be 20 minutes long rather than 15. These exercises are for everyone, for free. Paying [Subscribers](https://www.workroom-productions.com/#/portal/signup) get to play together every week. Cheers – James _This post is for subscribers only._ ### Exercise – Power of Variety URL: https://www.workroom-productions.com/power-of-variety/ Last updated: 2026-03-05T14:29:18.000Z *A 20 minute exploration via a simulation* #### history Based on an older exercise, no longer interactive, posts here: [An Experiment with Probability](https://workroomprds.blogspot.com/2012/11/an-experiment-with-probability.html), [Broken Trucks](https://workroomprds.blogspot.com/2012/11/broken-trucks.html), [Models, lies and approximations](https://workroomprds.blogspot.com/2012/11/models-lies-and-approximations.html), [Enumeration hell](https://workroomprds.blogspot.com/2012/11/enumeration-hell.html), [Diversity matters, and here's why](https://workroomprds.blogspot.com/2012/11/diversity-matters-and-here-why.html), [Modelling super powers](https://workroomprds.blogspot.com/2012/12/modelling-super-powers.html) . Go there to read background – I'll bring it here in the next few weeks and update it. #### about the simulations Each of these simulations have a similar core simplifications: - the **simulation* knows what can be found, and it knows what's likely to find each thing, but the **entities looking* don't know either of these things. - A searchable thing has a limited collection of independent things to be discovered. **For testing, think:* *there are only so many problems.* - Discovery is by chance. Different approaches to searching have different chances. A discoverable thing has a low chance of being discovered by any individual approach – but might have a better chance by another approach. *For testing, think:* **Performance testing will find different problems from usability testing.* - There's a limited budget – the more you spend, the more chances you have to find stuff. - Something stays found, but only counts the first time it is found. You, the user, can see and change the number of things that can be discovered, the makeup of the crew that looks for things, and the budget to spend looking. During the simulation, you can see the results of the simulation in terms of the work done, the discoveries found. You can pause and resume the work. You can increase and decrease the budget. Different simulations will throw up different results, and you'll want to keep track of those to see the variation. Background: a simulation works in chunks. In each chunk, all the explorers have a chance to discover a (random) set of discoverables (of fixed size in the simulation). A discoverable's chance of discovery is set by the explorer's active skill, and the chance of the discoverable being discovered by that skill. ## Priming Question ⁉️ How do you feel about spending, and stopping, in a situation where you need to discover new things? We'll come back to this in the debrief ## Get familiar with the simulator *5 mins* Go to , and set the simulation to `One tester, one technique` / `very sparse, all similar`. ![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/2026/03/image.png) Tap the **⎋** button – watch the "things to find" text change. Tap the **+** and **\-** buttons to see the budget change. Hit 'start' – see that the budget goes down and that the button now says 'working - pause'. Hit 'pause'. Unpause. Change the budget. Watch the rest of the UI – we'll talk about what it's showing. ## Simulation 1: Diminishing returns *5 minutes* Switch to 'Ordinary - A' as the subject and stay with 'one tester, one technique'. Start the simulation and 'spend' half the budget. At that halfway point, pause, and **decide whether it's worth spending the other half.** Do this a few times, and see what's influencing your decision. Change your team to 'six testers, one technique'. You'll see a list of workers. Run again, several times. What do you see that is different? Has that changed your decision-making? We'll talk as we go. ## Simulation 2: Varied workers *5 minutes* Change your team to "six testers, ten techniques". You'll see your workers have a selection of approaches. Click on the approaches so that they're all different. Run the simulation, several times and see how the stats change. What do you see that is different? Has that changed your decision-making? ## Simulation 3: One worker, several approaches. *optional* Switch back to a single tester with several skills ("one tester, ten techniques"). Run the simulation, switching their skill as they work. What do you see that is different? Has that changed your decision-making? ## Debrief *5 minutes* How do different approaches affect your decisions about budget and stopping How do you feel about spending, and stopping, in a situation where you need to discover new things? ### No Workroom Playtime this week (21 March) either URL: https://www.workroom-productions.com/no-workroom-playtime-this-week-21-march-either/ Last updated: 2025-03-19T13:17:23.000Z .. this week I'm full of headcold. So there's a second gap. I'm heading back to bed. Sorry all. As a bonus, though, I'm working towards a practical exercise on building custom exploratory interfaces for functions under test. It will be longer – 40-60 minutes – and I'll offer that to subscribers first. I'm hoping to get it done by early April\*. And I want to trial my all-new AIvsTDD workshop which will be delivered at EuroSTAR in June. Even thinking about them makes the generalised dull headache recede – so I'll keep a pen and paper nearby. Wishing you the best with your endeavours! Cheers – James \`\* depending on when I can start to concentrate again, how much of my attention my family needs when I'm back at full focus, and all the other demands that we juggle as we go. ### PlayTime 009 URL: https://www.workroom-productions.com/playtime-009/ Last updated: 2025-03-06T11:18:02.000Z We'll gather on Zoom / Miro, on Friday 7 March at 3:00pm London time ([local time for you](https://this-ti.me/?uts=1741359600&tz=Europe%2FLondon&name=Workroom+PlayTime+009)) for this weeks' [**Workroom PlayTime**](https://www.workroom-productions.com/workroom-playtime/), which is [Notation for Puzzle15](https://www.workroom-productions.com/notation-for-puzzle15/). Feel free to put your name on the Miro board if you're coming. This session will be 20 minutes long rather than 15. Last week, I wrote about some changes I'll make to Workroom PlayTime. We're already trying the longer session length, and the persistent Zoom room. I've not yet been able to schedule anything far in advance, and you can see that I've not yet moved from Fridays. I'll send details when I can. There are new articles and a conference announcement to share, too – I hope to send a newsletter at the weekend. These exercises are for everyone, for free. Paying [Subscribers](https://www.workroom-productions.com/#/portal/signup) get to play together every week. Cheers – James _This post is for subscribers only._ ### Exercise: Notation for Puzzle15 URL: https://www.workroom-productions.com/notation-for-puzzle15/ Last updated: 2025-03-07T10:03:51.000Z In this exercise, we'll explore [BlackBox Puzzle15](https://blackboxpuzzles.workroomprds.com/newtech/puzzle15.html). As usual with a BlackBox Puzzle, we'll aim to work out what's going on. My rule of thumb is that I can describe the behaviours in a tweet. As a variant for this exercise, we'll try to think about how we're keeping track of what we see and do, and of the models we make. ## Short – 20 minutes 20 minutes long, sharing encouraged. Do make a choice about how you'll keep notes. You're welcome to choose to keep notes in your usual way (maybe that means no notes), or in a way that is new to you, or to try out several approaches in a group. We'll talk about choices. 7-10 mins exploring We'll talk about how the notes influenced your actions, and about what actions you recorded (or didn't) in your notes. ## Extensions - Share your notes - Tell the story of your testing, using your notes as a memory aid. - think about the relationship between the time you explored in, and the time you needed to think about your exploration, and the time you need to retell it - What models helped? Did any approach mislead you? - Describe how the model of the system came to you. ## Support (to a separate page later) Our perception and our memory are separated by time – and fundamentally influence each other. This exercise may give insight into how that influence works for you. ### Weird Word 1 URL: https://www.workroom-productions.com/weird-word-1/ Last updated: 2025-03-06T16:50:50.000Z *An observation that might tell us something about what lies beneath the tools we use directly.* In my currently-favoured editor, [Bear](https://bear.app), I typed ![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/2025/03/image-2.png) [MacOS spellcheck](https://support.apple.com/en-gb/guide/mac-help/mchlp2299/mac) suggested, as the *sole* substitution: ![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/2025/03/image-4-1.png) Which is an odd option. I was distractible-enough to dig in. That weird word isn't in any of my dictionaries. I went to see if the suggestion was on the internet... maybe as a word from another language, maybe as a plausible misspelling. DDG assumed the suggestion was the word I'd intended: ![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/2025/03/image.png) This carries on for the rest of the page. Of course it does, because ![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/2025/03/image-3.png) So: what am I revealing here? I know that search and lookup isn't straightforward: I'm aware of the lossy magic of [bloom filters](https://en.wikipedia.org/wiki/Bloom%5Ffilter), and I get that there's almost certainly a large language mode in the background. I'm curious. Something puts these three collections of letters in such close proximity (in whatever coordinate space it is using) that it offers them as substitutes for each other. Not *a* substitute from a list, and not *the top* sub. The *only* sub. And as Apple's MacOS spellcheck and DDG's lookups do such similar weird things, that *something* might be shared. Weirdness reveals unknowns; similar weirdness reveals a pattern. Thoughts? 🧠 Maybe there's a technology that I can know about. Maybe I've made a mistake in my thinking. Either way, exposing my ignorance is one way to find out, and to change. 🤹 I've not put the terms in **text* here, and I think that's because I don't want these terms to be indexed together. I believe that search engines pick text from images, so there's a bit of cognitive dissonance here. ### Function is easier than Fit URL: https://www.workroom-productions.com/function-is-easier-than-fit/ Last updated: 2025-04-07T16:24:43.000Z When building the scripts and prompts for the [AIvsTDD workshop](https://www.workroom-productions.com/guiding-ai-code-with-tests-workshop/), I wanted\* to understand the relevant moving parts, and to abstract the ones that were generic. One thing that became very clear was that *asking for code* is easy. Some LLMs respond with code that is fairly-reliably syntactically correct. Code the runs first time felt like a triumph, for me and my butterfingers. A temporary triumph. Getting *code that passes tests* was the aim of the [magic loop](https://www.workroom-productions.com/guiding-ai-code-with-tests-workshop/#my-magic-loop-is-working). With a bunch of shell scripts, that goal was both achievable and interesting. However, getting *code that fits* is harder. When you give an LLM a huge great lump of code, and ask it for a change, how is it meant to tell you the several places that the change needs to be made? Remember that it's building its reply word-by-word, line-by-line. It doesn't 'know' what it's going to write, and doesn't look back over to see where it fits when it's done so. Something needs to work out where it goes: maybe another prompt to the LLM (you could ask what needs to change and where before asking for the change), maybe your deterministic tools (it's common to ask for a format that look s like `diff` output), maybe you. In our workshop, we pulled a fast one, and skipped that problem\*\*. It's interesting to find, then, this experiment from Aider: [GPT Code Editing Benchmarks](https://aider.chat/docs/benchmarks.html#the-benchmark) – looking at the differences between editing a file based on (something like a) a `diff`, and replacing the whole thing. 🤯 Aside: We started with the unit tests, and iterated several times. Aider start with a description, then run the tests and offer a second round based on the test results; the LLM doesn't see the unit tests. But, as they point out, these are all part of the training data, so in a way the LLM already knows what code works. Aider's initial conclusions balanced between editing and replacement. In their follow-on and ongoing experiment [LLM Leaderboards](https://aider.chat/docs/leaderboards/), they seem to be leaning more towards the editing approach. Either way, the firm conclusion is that while LLMs can produce runnable and maybe-satisfying code, it's harder to reliably put that code in the right place. To follow: Patterns of failure in integrating generated code. [Related comment on LinkedIn](https://www.linkedin.com/feed/update/urn:li:activity:7315036977299427329?commentUrn=urn%3Ali%3Acomment%3A%28activity%3A7315036977299427329%2C7315046612119027714%29&dashCommentUrn=urn%3Ali%3Afsd%5Fcomment%3A%287315046612119027714%2Curn%3Ali%3Aactivity%3A7315036977299427329%29) --- \`\* I reckoned I wanted to work from scratch, but that is entirely illusory when your tools rest on the ingested content of the internet. \`\*\* The sharp-eyed in our workshop will notice that we worked that particular misdirection by asking for changes to one file only, and asking for the *whole* file. *Subscribers get to see notes on what's wrong with changing a whole file* _This post is for subscribers only._ ### Thinking in Client URL: https://www.workroom-productions.com/thinking-in-client/ Last updated: 2025-03-06T16:58:06.000Z When I was a kid, the best puzzle magazine in the world was [*Jeux et Strategie*](https://fr.wikipedia.org/wiki/Jeux%5Fet%5FStratégie)*.* It was in French, and I learnt French by reading it. I found that, in the right context, I *think* in French – and I noticed this only after several years of (occasionally) thinking in French. Now I'm an adult, my ability with languages is pretty poor, by European standards. I have bad Bulgarian and awful French. I get the two muddled if I'm out of context. I can, however, **think in *client***. When I'm consulting, local terminology takes priority. Every client – sometimes every team – has different meanings for `end to end testing`, for `quality`, for every element of `independent verification and validation`. I'll always ask, and I'll often be surprised. *Thinking in client* has led to more productive conversations, and to longer relationships. It makes it easier for me to engage without having my buttons pushed, and it makes it easier to switch. When people use words that have a local meaning, I can feel my mind substituting. I imagine that I'm not alone. I imagine that our ability to switch phrasing and jargon evolved as our species' languages developed in family and local groups, and when humans moved from group to group, those with such abilities had greater representation in the DNA of the next generation: they were better at *chatting up*. If you're moving from group to group as a consultant, it's good to be good at this. I reckon that I can choose to use it, and I can develop it, because I recognise it as a facility. So here's your starting point if this seems useful: **If you're consulting with a client, *think in client*.** More below for subscribers: how to build the ability; how it works for me online, in person, and when facilitating; what it means when approaching writing / speaking. _This post is for subscribers only._ ### Submit! URL: https://www.workroom-productions.com/submit/ Last updated: 2025-02-27T12:18:46.000Z ...no. I'd like to *put forward* my idea. I'd like to *apply* to speak at your event. I'd like to *offer* to work with your organisation. I'd like to *propose* a way forward*.*. Of course, that will involve submitting my idea, my paper, my bid, my plan, to your judgement. *I* submit *it. It* is *submitted*. Not me. Yet I am exhorted to *Submit!* as your chosen call to action. *Cower!* *Yield!* *Confirm!* *Capitulate!* How do *you* feel, reading that? So, **Submit?** *Me?* Fuck off. Fuck right off. My *Think in Client* daemon is doing overtime for your choice of words. ### An Example Exploratory Interface URL: https://www.workroom-productions.com/an-example-exploratory-interface/ Last updated: 2025-02-27T10:19:13.000Z *Here it is:* [*exploratory interface for Timer\_v2.js*](https://exercises.workroomprds.com/CadenceKeeper/timer%5Fexplorer.html) ![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/2025/02/image-6.png) The systems we make and test have all sorts of ways that we can explore them. Each way of exploring will reveal different problems. If we want to explore a unit of a system, we might start with unit tests, and write a few more to explore. I tend to adapt unit tests to use bulk input and simple oracles. The unit tests are my exploratory interface. There are other possibilities, so here's an example, and a story. I built a toy; [CadenceKeeper](https://exercises.workroomprds.com/CadenceKeeper/). ![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/2025/02/image-5.png) CadenceKeeper uses several JavaScript source files – a central one is [Timer\_v2.js](https://exercises.workroomprds.com/CadenceKeeper/Timer%5Fv2.js). I wanted to explore that unit on its own; I had an impression that the timers didn't always change their behaviour as they aged, and preferred to experiment and play rather than to predict and observe. Anthropic's [Claude-3.5-sonnet](https://www.anthropic.com/news/claude-3-5-sonnet) built me this [exploratory interface](https://exercises.workroomprds.com/CadenceKeeper/timer%5Fexplorer.html) for `Timer_v2.js`. I gave it the code, and asked for an HTML page which: - shows the constants in use - lets me construct new objects (that's what the code does) with parameterised and default values - lets me select from objects I've made, and shows me its values in real time - lets me trigger and monitor events described in the code - gives me a simple UI for any callable method in the code If memory serves, I had to ask a few times before I had something I felt I could use. All those bits are there, most of them folded up. Using my cheap new tool, I mucked about with the working code, found a couple of handy bugs, fixed them and moved on. I found it swift and interesting to *prompt for* the interface, *dig for* the problems, *write* the right tests to expose them and then *fix* the problems. I found it particularly helpful, when making code-based confirmatory tests, that I had already understood how to trigger the problem and how to look for its symptoms. I would have been unlikely to write an interface like this by hand, because it would have been a [yak](https://projects.csail.mit.edu/gsb/old-archive/gsb-archive/gsb2000-02-11.html) (and a half). Instead, it was an [*ephemeral tool*](https://www.workroom-productions.com/ephemeral-tooling/)*:* I wished it into existence, I used it, I threw it away. And then I got it back out of the bin and posted it here. It didn't work at all, because my code had already moved on, so I rewound the version of Cadence Keeper back to the older one. It probably never did *work*, perfectly – but it worked well-enough for me to sniff out evidence, it works well-enough to let you explore how it feels, and its cost was marginal. The tininess of that cost is what mattered when I made it: I could give my attention to exploration, not to the code nor to the tooling. If you want to use this tool, you'll need to *explore* the tool, and be mindful of the difference of exploring they system it wraps round. Hard for you, easy for me because I'm the builder. Good enough. Expect this page and this example to change. Next time I do this interestingly, I'll capture the details so I can write about them with clarity. And there's an exercise coming, so that you can build your own interfaces. Indeed, work is afoot with several valued colleagues to do this in public at an event near you... ### PlayTime 008 – Imagining Questions URL: https://www.workroom-productions.com/playtime-008-2/ Last updated: 2025-02-27T10:12:44.000Z We'll gather on Zoom / Miro, on Friday 28 February at 3:00pm London time ([local time for you](https://this-ti.me/?uts=1740754800&tz=Europe%2FLondon&name=Workroom+PlayTime+008)) for this weeks' [**Workroom PlayTime**](https://www.workroom-productions.com/workroom-playtime/), which is [*Imagining Questions*](https://www.workroom-productions.com/imagining-questions/). Regulars should note another shift in time. Also, I'll need to put a hard end on this one, as I'm due in a parent / kid art thing at 3:30. I change as I learn: As we're a couple of months in, here are some changes to the format: - I'm going to lengthen the time from 15 to 20 minutes. I find it hard to build something that works well in 15 minutes, and 20 minutes is a standard 'set'. - I'm going to move it from Friday afternoons. I imagine that I'll run it midweek afternoons. - We'll try using the same Zoom room for all Workroom PlayTimes – and it's kind-of-on all the time so we can drop in. - I'm going to schedule (at least) 4 weeks in advance, and move them if I need to (or run them in different timezones if *you* need to), rather than announce less than 48 hours before (which is satisfying for *nobody*). **Help me by letting me know what day / time might work well for you.** These exercises are for everyone, for free. Paying [Subscribers](https://www.workroom-productions.com/#/portal/signup) get to play together every week. --- **To play with this week:** After too many promises over the last few weeks, it's time for me to give you an example of a [generated exploratory interface](https://www.workroom-productions.com/exploratory-interfaces/). As context, here's a one-page toy called [CadenceKeeper](https://exercises.workroomprds.com/CadenceKeeper/). CadenceKeeper uses several JavaScript source files – a central one is [Timer\_v2.js](https://exercises.workroomprds.com/CadenceKeeper/Timer%5Fv2.js). I wanted to explore that unit on its own; I had an impression that the timers didn't always change their behaviour as they aged, and preferred to experiment and play rather than to predict and observe. *(note: if you had this post as an email, I sent it with the wrong link at the start of the § – fixed now)* Anthropic's Claude-3.5-sonnet built me this [exploratory interface](https://exercises.workroomprds.com/CadenceKeeper/timer%5Fexplorer.html) for `Timer_v2.js`. I gave it the code, and asked for an HTML page which: - shows the constants in use - lets me construct new objects (that's what the code does) with parameterised and default values - lets me select from objects I've made, and shows me its values in real time - lets me trigger and monitor events described in the code - gives me a simple UI for any callable method in the code If memory serves, I had to ask a few times before I had something I felt I could use. Using my cheap new tool, I mucked about with the working code, found a couple of handy bugs, fixed them and moved on. I found it swift and interesting to *prompt for* the interface, *dig for* the problems, *write* the right tests to expose them and then *fix* the problems. I found it particularly helpful, when making code-based confirmatory tests, that I had already understood how to trigger the problem and how to look for its symptoms. I would have been unlikely to write an interface like this by hand, because it would have been a [yak](https://projects.csail.mit.edu/gsb/old-archive/gsb-archive/gsb2000-02-11.html) (and a half). Next time I do this interestingly, I'll capture the details so I can write about them with clarity. And there's an exercise coming, so that you can build your own interfaces. Indeed, work is afoot with several valued colleagues to do this in public at an event near you... If you want to encourage me to keep going, [switch to the paid subscription](https://www.workroom-productions.com/#/portal/signup) – and that means you can be sure of a weekly invite to come and play. Cheers – James 🫣 Aside: don't expect the code or story above to be static. This promised content is delayed because my notes got disorganised and because my memory is flaky. I've posted what I've got and I'll replace and change if I find that I've uploaded something broken or if my story doesn't reflect how I did it. In particular, I want to post the **prompt*. _This post is for subscribers only._ ### Imagining Questions URL: https://www.workroom-productions.com/imagining-questions/ Last updated: 2025-02-27T09:47:29.000Z *Used in the* [*Questions, Questions workshop*](https://www.workroom-productions.com/tag/questions-workshop/) *and* [*Workroom PlayTime 008*](https://www.workroom-productions.com/playtime-008-2/) Imagine that you're in one of the scenarios below... - you're meeting a colleague, who needs a new function to be tested - you've just joined a new testing team - you're trying to find out what someone needs to see in your test report - *or make up your own scenario this is meaningful to you* Share your scenario on the board. Read everyone's, group up if you like, go solo if you prefer. *Solo, in pairs or small groups (5 mins)*: Write down the questions you would ask to find out more information. *Analysis (5 mins)*: How would you order your questions? What patterns and gaps can you see? How would you change to respond to that analysis? *Exchange (5 mins)*: Share what you learned. *Optional extension 1*: Share your questions for your scenario *Optional extension 2*: swap scenarios, add questions to each other's sets, notice further patterns and gaps. ### Ephemeral Tooling URL: https://www.workroom-productions.com/ephemeral-tooling/ Last updated: 2025-05-11T15:40:02.000Z There's a place in our world for tools that we *throw away*. Which makes a nonsense of RoI. The investment is negligible, and we all know about divide-by-zero. So we don't seem to talk about them, and we don't necessarily share them, but we use them nonetheless. Ever thrown a `.csv` into a spreadsheet, filtered it to know something then closed without saving? *Ephemeral tool.* Ever piped a log into a grep to look for trouble with `tail -F log.log | grep 'trouble'`? *Ephemeral tool.* Ever pasted a bunch of (sanitised) data into a regex website to build a regex, get immediate feedback on whether the regex is right, then use the regex to look for data – and when you've found the data, chucked away the regex? *Ephemeral tool.* Ever asked an LLM for a web page to make a graphical interface to a new function? Well... only since 2024, at a guess. And when someone changes the function , do you fiddle about with the weird-ish LLM-generated page code to make a matching change, or do you just... regenerate it? You see where I'm going here. Since the advent of LLMs that can barf out fairly-OK code, fairly-OK code has become vanishingly cheap (if one ignores for the sake of rhetoric, the environmental impact of one's query and the social impact of having the tool at all). And if code is cheap, then it might be cheaper to chuck it away and start again than to engage in maintenance. *Ephemeral tool.* --- Automated tests are *not* ephemeral. Indeed, they can be imagined to be more permanent than the system they test. They may confirm that a system continues to be valuable, and a system's *value* can outlast any part of its *code*. Risks, though, should be temporary. I've seen a weird resistance to tooling up to look for trouble, perhaps because the Return part of the RoI is lumpy – generally zilch, occasionally zillions. Once the Investment part is tiny, though, off we go. So we're perhaps at an inflection point; it's time to *talk about* how we do this. --- Testing is a business of experimentation. To set up those experiments, we regularly deal with ephemeral *data*. If your organisation supports it, ephemeral *environments* enable new classes of experiments. Ephemeral *tools* open our access to further experiments – allowing us to trigger and to observe in ways that we can't with our slow and impatient senses. --- Ephemeral tooling isn't for production. That would mean that it's nether ephemeral, nor tooling... but that's semantics. And I learnt a bunch of tiny tools from midnight ops gangs who kept various very-production overnight batches running. These tiny tools offer swift access to novel experiments. The results of those experiments are likely to be as weird or weirder than the more-sold software they might test. We *don't* trust the output. We use the tools to hint to our judgement that something smells funny. Sometimes, we use them to do jobs we can do (because they're faster, more relentless, more precise) so that we can free up our minds – ephemeral tooling lets us to use that power on one-off jobs, too. In both cases, ephemeral tools don't need to have high quality to some end user: their value is in their high *relevance* to us. Easier accessibility and lower cost don't make such tools more useful or more compelling – they make them more feasible. --- Weirdly, while I was planning this, I found an [article I'd written in 2022](https://www.workroom-productions.com/exploration-and-fact-checking/), spun out of an idea from a writer named [Mike Caulfield](https://hapgood.us/), and which I'd published in such a way as to be unfindable. I read it (didn't remember it), jiggled it and re-published it. Imagine my astonishment when this popped up today from his [Substack](https://mikecaulfield.substack.com/): [Is AI-Produced Ephemeral Software the Future of Novice Computing?There are issues, but I think its a likely path![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/icon/icon.svg)The End(s) of ArgumentMike Caulfield![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/thumbnail/https-3A-2F-2Fmikecaulfield.substack.com-2Ftwitter-2Fsubscribe-card.jpg%3Fv%3D-258285940%26version%3D9)](https://mikecaulfield.substack.com/p/is-ai-produced-ephemeral-software?utm%5Fsource=share&utm%5Fmedium=android&r=hbr2&triedRedirect=true) Different take, but how lovely. It ends with this, which reminds me of all the [wrangling](https://www.workroom-productions.com/wrangling-debugging-and-testing/) we do before we ever see a test pass, or fail. > My daughter took an AP computer class about 4 years back, and wrote a simulated card game — half the project was getting the Java development environment to compile the right dependencies. She went in excited about building things, and came out never wanting to touch code again. That stuff (IDE, dependencies, object-oriented design) is all important, but there’s just a lot of people who could at least start to engage with programming if the entry point was a bit easier. --- There's more to come, and I'm hiding it behind a paywall that nobody but me can see. So you can subscribe to see more, but not from this article... _This post is for subscribers on the Paying and Owner tiers only._ ### PlayTime 007 – Sloganise for all subscribers! URL: https://www.workroom-productions.com/playtime-007-sloganise/ Last updated: 2025-02-19T17:37:06.000Z We'll gather on Zoom / Miro, on Friday 21 February at 3:30pm London time ([local time for you](https://this-ti.me/?uts=1740151800&tz=Europe%2FLondon&name=Workroom+PlayTime+007)) for this weeks' [**Workroom PlayTime**](https://www.workroom-productions.com/workroom-playtime/), which is [Sloganise](https://www.workroom-productions.com/sloganise/) from SpeakerPrep. Regulars should note the slight shift in time: we're 30 minutes earlier for this one. If we're all lucky, this week's blurred background will be different: I'll be in [Kew Gardens](https://www.kew.org) with family, and will take a risk of pixellation by coming in by mobile. These exercises are for everyone, for free. Paying [Subscribers](https://www.workroom-productions.com/#/portal/signup) get to play together every week. --- Other things this week: **Things to play with:** It's half term, so everything's a bit half-baked. Nevertheless, here's a short article [Judging by Content, Delivery, and BraveryHow I Judge Talks![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/icon/wpl-60x60-logo-semitrans-9.png)Workroom ProductionsJames Lyndsay![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/thumbnail/photo-1739911013843-0380d6504480.jpeg)](https://www.workroom-productions.com/judging-by-content-delivery-and-bravery/) and a related exercise [Exercise: Abstract JudgementJudging an abstract, and thinking about judgement![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/icon/wpl-60x60-logo-semitrans-10.png)Workroom ProductionsJames Lyndsay![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/thumbnail/photo-1666609563103-76c1c7a3f586)](https://www.workroom-productions.com/exercise-2/) I've still not got all the ducks in a row for the custom exploratory interface I teased you with a few weeks ago. I'm putting together a supporting page on Ephemeral Tools, now, too... about which much more, soon. If you want to encourage me to keep going, [switch to the paid subscription](https://www.workroom-productions.com/#/portal/signup) – and that means you can be sure of a weekly invite to come and play. Cheers – James _This post is for subscribers only._ ### Sloganise URL: https://www.workroom-productions.com/sloganise/ Last updated: 2025-02-19T17:04:54.000Z *A SpeakerPrep exercise to give your audience a peg to hang your idea on.* You'll need a talk, or an abstract for a talk. Your talk will have an overall message or conclusion. Write it down now if it helps you to keep it in mind. Give yourself 10 minutes to come up with a phrase you can say at the start and end of the talk, and that people who have been to your talk might say to each other. Use the approaches below if you need hints, and try it out as you're playing. Work alone or with someone else, or in a group. Do be silly and over-the-top as you're playing with ideas – it will help free your mind, and you don't need to share every step on your journey. At the end: Share your slogans. ### Some approaches Maybe you'll try out different parts of your message. Maybe you'll try a few phrases and polishing down. Maybe you'll try for a one-word essence and build up. Maybe you'll see where the randomness of alliteration and rhyming take you, or maybe you'll seek the nuance that pulls the phrase together. - Alliterate – if words don't turn up, use a thesaurus. - Rhyme – use a rhyming dictionary - Delete filler – does it still work? If not, why not? - Make a headline – clickbait, tabloid, or chapter - Who is doing (or saying) this? Can you remove the reference, make it explicit, or change it? - Reduce your message to one word – then try again for three words, or five words. - Give action - Fiddle with vocabulary and endings - Try lots of alternatives, combine as needed ### Try it out - Think about punctuating your talk with it. - Think about putting it on a t-shirt or hat or sticker. How would it feel to have it on every slide? - Think about what you want the audience to feel as you close the talk. - Think about what you want them to do differently once they've given you their attention. - Think about the way that your title, your conclusion, your list of headings and / or your key takeaways relate to your slogan. When you've tried it out it yourself, try it out with other people in the workshop ### Exercise: Abstract Judgement URL: https://www.workroom-productions.com/exercise-2/ Last updated: 2025-02-19T16:19:11.000Z *An exercise to think about our own judgement, to help us write things that other people will judge.* Pick an upcoming event whose sessions are described online. Use if you don't have one in mind. Read the sessions. Which would you go to? See if you can work out *why* you feel like that. Any patterns in what you read? Is there just one reason? If you have several reasons, how different are they? Any underlying patterns in how you judge? **Publicly:** Share your patterns. That is *not* the end of the exercise. **Privately:** look at your own work – how would you judge it, based on your patterns? How might someone else judge that work? What might you change about it? *Extension: If you've got time, try a different event – or think back to an event you went to, and read into the abstracts for the talks you went to.* *Extension: If you've got the urge, do this with music, or chocolate, or political ideas, or friendships...* --- Background: read [Judging by Content, Delivery, Bravery](https://www.workroom-productions.com/judging-by-content-delivery-and-bravery/) for what I find I judge against. This exercise exists in different forms, typically around building an aesthetic; I'll add links to the ones where we judge chocolate (or beer), generated pictures and code. ### Judging by Content, Delivery, and Bravery URL: https://www.workroom-productions.com/judging-by-content-delivery-and-bravery/ Last updated: 2025-02-19T15:59:55.000Z I'll look at your 300-word abstract, or the one-liner the conference has given you, or your slide deck, or your video, or your performance, and I'll make an irrational judgement\*. Trying to rationalise my judgement, I *think* that I favour the following: - **Content**: if there's something specific that I want to use - **Delivery**: if I'm having fun - **Bravery**: if you're pushing yourself somehow I'll favour something with more of these qualities over talks with fewer, and my mood at the time dictates whether I favour one over another. There are also turn-offs. I don't, at all, favour teases (especially if I'm being *asked* to judge) and assume that if you won't tell me your meaningful content, you don't have much. If even your summary stinks of LLM-generation (or, in earlier days, is littered with letter errors), I'll imagine that you're not going to pay much attention to your delivery, either. And while I admire chutzpah, I abhor bullshit manifesting as bravery. Boo to you. An experience report pretty-much always has usable content somewhere, and can be packed with bravery if candid. An interactive workshop lets me play, a performer makes me delighted, and a proper story catches me like a big old rusty hook. I'll give my attention to a brave new speaker even if all they have are old truths, but an old hand had better have really good stuff or at very least be weird. --- \`\* *Judgement* is a feeling. And it's also a measured dissection of the situation, its context, and the regulations that apply to arrive at a binding decision. These two things are different: That's how language works. I recognise that I *feel* attracted or repelled by a talk, I don't think that I should trust the feeling without some thought, so I try to take it apart to see how it works. ### PlayTime 006 – Explorable Interfaces II URL: https://www.workroom-productions.com/playtime-006-explorable-interfaces-ii/ Last updated: 2025-02-12T13:38:09.000Z We'll gather on Zoom / Miro, on Friday 14 February at 4:00pm London time ([local time for you](https://this-ti.me/?uts=1739548800&tz=Europe%2FLondon&name=Workroom+PlayTime+006)) for this weeks' [**Workroom PlayTime**](https://www.workroom-productions.com/workroom-playtime/), which is on Exploratory Interfaces. We'll play using the Exercise [Exploring Fixed Input](https://www.workroom-productions.com/exercise-exploring-fixed-input/). As support, here's something on [Exploratory Interfaces](https://www.workroom-productions.com/exploratory-interfaces/), which I'll add to as I go. These exercises are for everyone, for free. [Paying subscribers](https://www.workroom-productions.com/#/portal/signup) get to play together. --- Other things this week: **Community stuff:** Subscriber kudos to [Uros Stanisic](https://www.linkedin.com/in/uros-stanisic-78965849/), who recently pointed out a rubbish link in a recent PlayTime page. Remember to get a [submission in to Agile Testing Days](https://agiletestingdays.com/call-for-papers/). I dropped into [MoT's Testing Planet on Quality Engineering](https://www.ministryoftesting.com/the-testing-planet-sessions/exploring-quality-engineering-the-testing-planet-news-episode-08?s%5Fid=19009086) last week, and found handy perspective on the "[Engineer](https://www.engc.org.uk/glossary-faqs/frequently-asked-questions/status-of-engineers/)" title in [Isabel Evans](https://www.linkedin.com/in/isabelevans/)' description of her research into what we do in our jobs vs the titles we get (on Lisa Crispin and Janet Gregory's [Donkeys and Dragons](https://www.youtube.com/@AgileTestingFellowship/videos) vlogcast thing). **Things to play with:** Apart from the new exercise, I've put up a couple of new toys from last week's PlayTime: [VSCode with command line access to a server](https://www.workroom-productions.com/browser-based-vscode-for-workshops/) – accessible via workshop participants' browsers. Ansible scripts on GitHub. [Browser-based VSCode for WorkshopsPutting a cloud server and IDE into workshop participants’ browsers.![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/icon/wpl-60x60-logo-semitrans-7.png)Workroom ProductionsJames Lyndsay![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/thumbnail/photo-1488992783499-418eb1f62d08.jpeg)](https://www.workroom-productions.com/browser-based-vscode-for-workshops/) A fast-moving log, first in a [series of artificial logs for teaching with](https://www.workroom-productions.com/making-logs-for-teaching/). Shell script on GitHub. [Making Logs for TeachingPointer to a github repo of tools that make logs![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/icon/wpl-60x60-logo-semitrans-8.png)Workroom ProductionsJames Lyndsay![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/thumbnail/photo-1604292771492-1aee47c1456b)](https://www.workroom-productions.com/making-logs-for-teaching/) That custom exploratory interface I teased you with a few weeks ago (built with/by Claude to help me exercise a piece of code) is about to land, but that depends on half term and my various conflicting / meshing responsibilities. Argh: adulting. If you want to encourage me to keep going, [switch to the paid subscription](https://www.workroom-productions.com/#/portal/signup) – and that means you can be sure of a weekly invite to come and play. Cheers – James _This post is for paying subscribers only._ ### Exercise: Exploring Fixed Input URL: https://www.workroom-productions.com/exercise-exploring-fixed-input/ Last updated: 2025-02-12T12:35:13.000Z *15 Minutes.* > *This exercise is being used for Workroom PlayTime 006, Friday 14 Feb at 4pm London time, on Zoom. Become a paying subscriber to join us.* We'll explore [Converter\_v3](https://exercises.workroomprds.com/converter%5Fv3%5Fbase/), again. We'll do this together, mics / cameras on, talking. We often explore with inputs that we can vary. That can lead us to limitation. In this exercise, we'll explore system inputs that we can't change. ## Exercise: What's Configurable? *5-10 minutes* Pick the `Metric units: configuration`option to the see the current configuration of the system under test. What looks configurable, to you? We'll share ideas, we'll talk about what test ideas might be in play – and we won't hold back from experimenting as we go. *Debrief: Share a particular part of the config that you want to investigate.* ## Exercise: Input to Explore Configuration *3-5 minutes* We can pick our inputs to explore the effects that our configuration might have. Let's do that. *Debrief: what did we learn from our experiments?* *Extend:* We can also change how our inputs work. James has two approaches to look at behaviours around configuration using bulk input. Ask if you want them to be turned on. ## ### Browser-based VSCode for Workshops URL: https://www.workroom-productions.com/browser-based-vscode-for-workshops/ Last updated: 2025-06-12T16:03:55.000Z [Code Server](https://github.com/coder/code-server) puts [VSCode](https://code.visualstudio.com) in the browser, giving access to a remote server. [Ansible](https://docs.ansible.com/ansible/latest/index.html) lets me configure cloud servers. I can use them together to give workshop participants a configured and familiar development environment, without needing downloads or installs, and to make it available in moments. I used this at EuroSTAR to give \~90 people browser-based zero-install access to a VSCode's file browser and commandline, on servers set up to serve web pages and access LLMs. It worked transparently – people engaged with the work, not the tooling. I'll I've got two scripts. One [builds a cloud server](https://github.com/workroomprds/VSCodeInBrowser/blob/main/droplet%5Fonly%5Fsetup.yml) that I can close down and save for later use (it takes a while, and is flaky, so it's better to do it once and keep it somewhere). The other uses that to [build a live server, sets up VSCode and several users, installs tools and adds files](https://github.com/workroomprds/VSCodeInBrowser/blob/main/user%5Fsetup%5Fand%5Finfo.yml) that might be needed. Here's a github with (a slice of) the current state. [GitHub - workroomprds/VSCodeInBrowserContribute to workroomprds/VSCodeInBrowser development by creating an account on GitHub.![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/icon/pinned-octocat-093da3e6fa40-3.svg)GitHubworkroomprds![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/thumbnail/VSCodeInBrowser)](https://github.com/workroomprds/VSCodeInBrowser) This needs polish around VSCode config – setting up so that particpants don't need to battle a first-use page, making sure the file browser and terminal prompt are pointing to the same place, allowing (reliable) copy/paste from the browser-hosting OS to the OS on the server, pre-installing VSCode extensions, activating a Python virtual environment on the remote server. Why not use [Replit](https://replit.com), [CodeAnywhere](https://codeanywhere.com), or [GitPod](https://replit.com)? I have... and I may write about those (sometimes expensive) adventures elsewhere. ### Making Logs for Teaching URL: https://www.workroom-productions.com/making-logs-for-teaching/ Last updated: 2025-02-07T15:50:21.000Z When I'm testing, I spend time with logs. I don't have a go-to library of log makers to use when I want to teach someone how I use logs. So I've started one of my own: [GitHub - workroomprds/logmaker: a repo to contain various tools to make logs, so that people can play with logsa repo to contain various tools to make logs, so that people can play with logs - workroomprds/logmaker![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/icon/pinned-octocat-093da3e6fa40-2.svg)GitHubworkroomprds![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/thumbnail/logmaker)](https://github.com/workroomprds/logmaker) ## Contents ### Logmaker.sh `logmaker.sh`: a short shell script to generate a rapidly-updating log. The log is size-limited, and meaningless. It's random., but not uniformly random. It's multi-process and headless, so it does have job control – call with `start` to start, and `stop` to stop. It writes to `/tmp/simulated_log.txt` and puts another two files into `/tmp` – use at your own risk. ### PlayTime 005 – `grep` URL: https://www.workroom-productions.com/playtime-005-grep/ Last updated: 2025-02-06T14:32:08.000Z We'll gather on Zoom / Miro, on Friday 7 February at 4:00pm London time ([local time for you](https://this-ti.me/?uts=1738944000&tz=Europe%2FLondon&name=Workroom+PlayTime+005)). We'll do an exercise with `grep`, using the command-line tool for a swift round of testing-related tasks. Here's more about [grep for testers](https://www.workroom-productions.com/grep-for-testers/), and here's a [link to the exercises](https://www.workroom-productions.com/grep-tool-exercises/). These exercises are for everyone, for free. Paying subscribers get to play together. _This post is for paying subscribers only._ ### Grep Tool Exercises URL: https://www.workroom-productions.com/grep-tool-exercises/ Last updated: 2025-06-21T07:14:42.000Z *unfinished!* `grep` filters information, and is typically used on the commandline as one part of a chain of tools. Here's[ how grep might be used by testers](https://www.workroom-productions.com/grep-for-testers/). If you've used graphical IDE tools that have a facility to search all files in a folder, there may be `grep` under the hood. Or, there may not – OS-based search typically indexes files and picks files from its index. `grep` doesn't have a search index (as far as I'm aware), but instead does pattern recognition on a stream – and one way to introduce a stream is to feed it files. In this exercise, you’ll probably use command-line `grep`, and for some exercises you'll use it with other tools. There will be instructions, but the key concept is that if you see a command that looks like `toolX | toolY`, then `toolX` works first, and its output is connected to `toolY`'s input. The `|` is called *pipe*, and (although it might be better if it was rotated 90 degrees `–`) it represents a connection. If you're working with me, and I give you an environment, the files / streams will be there. If not, you'll need to download the files, and some exercises might not work for you (if they need streams, not files, and the streams aren't on). This exercise is more fun in a group. If you’re in a group, please talk about what tool you’d use. Perhaps re-run the exercise with someone else’s tool of choice, or with someone else. The solutions are below. The tool will do most of the work – give yourself kudos for spotting non-trivial things. Zero kudos if you look at the solutions before trying the exercises. There are two routes into this: - I have a development environment (in the browser) for you. If you're with me and I've set up, we'll work with that. Details below. - You work on a system that you own, or on one of the [interactive grep sites listed on the *grep for testers* page](https://www.workroom-productions.com/grep-for-testers/). MacOS and Linux always have `grep` available on the command line. If you're on a different OS, or if you can't get to the commandline, use one of those – I know that can take the Apache.log file below. ## James' in-browser dev env I've set up a [server](https://envs.workroomprds.com) which is running several [instances of VSCode](https://github.com/coder/code-server), accessible via the browser. Via VSCode, you'll access the command line, the file system, and whatever else is available. I hope that you feel relatively familiar with it. #### Access 1. Go to , pick the env that seems most ... you. Tell the others, as I've not tried this with several people in one env. Your password is `password`. VSCode may want to set up a more-complete env, so skipor move on until you see something like this: ![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/2025/02/image-1.png) 1. Open up the files and the commandline, and choose the `terminal` tab if you need to. 2. The chunk of green text is your *prompt* – where you'll type commands. Try `pwd` to see where you are: probably `/home/«whatever you chose»` ![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/2025/02/image-2-1.png) 1. Use *Open Folder* to get a dropdown, probably pointing to the same place (with a `\` on the end. If it's not that, change it and hit enter to see the filesytem. If asked, you *trust the author*. You'll see a filesystem sidebar on the left, and you may need to open the terminal again to get to (roughly) here: ![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/2025/02/image-4.png) So now you've got a browse-able file system, and commandline access to a remote server, all in your browser. Which is what we need to start. *One irritation: copy / paste from your host may not work for you. Or it might. If nit doesn't try right-click. I'm working on the config. It works for me...* 1. If you're unfamiliar with VSCode, take a moment to look around. Open the three-bar menu top-left, resize the parts if you want to, see the tooltips. #### VSCode Parts - ****Mode Panel** (Narrow leftmost Sidebar): - - The menu icon at the top gives access to the VSCode menus - The icons below change what the panel to its side does – file explorer, search, git, debugger, extensions. - ****Explorer Panel** (Left Sidebar): - - Shows your files and folders / search results / git history etc. - ****Editor Area** (Center-right): - - Currently shows a "Welcome" tab - Has quick actions like "New File..." and "Open File..." - When you select a file, you'll see it here. - Can have tabs, and can be split vertically - ****Terminal** (Bottom): - - Command-line interface, with prompt - You can run commands directly here without leaving the editor - "OUTPUT" tab: Displays program output - "TERMINAL" tab: For command-line operations - ****Top Bar Elements**: - - Command bar for VSCode commands - Top right buttons show layouts - ****Bottom Status Bar**: - - Shows useful information like Current file type (bash), Layout settings, Additional status indicators ## Exercises Group 1 – scanning a log If you're with James, follow the [link to the environment](https://envs.workroomprds.com). If you prefer, or if you're not with James, download the linked files. ### Exercise 1.01 An Apache.log file has several lines indicating a `Factory error` that you'll find scattered throughout. Use `grep` to find out more. #### hint You might want to `grep "«something you're looking for»" "«the path and filename»"` . The `"` are needed to cope with have spaces etc. *On James' env:* The file is in `/tmp/Apache.log`. You can look at it in the editor by choosing (menu) » file » open file – and opening `/tmp/Apache.log`. It's 56482 lines long. *Without James:* Go get the file linked below, extract it to a `.log` file. [Apache web server error log...from loghub: https://github.com/logpai/loghubApache.tar5 MBdownload-circle](https://www.workroom-productions.com/content/files/2025/02/Apache.tar "Download") #### Solution to Exercise 1.01 You might use `grep Factory /tmp/Apache.log` or `grep "Factory error" /tmp/Apache.log` or `grep -i "factory error" "/tmp/Apache.log"` or `cat /tmp/Apache.log | grep -i "fact"` to extract the info. Output below. Kudos if you noticed that they turn up in clumps of 4. ``` [Thu Jun 09 06:07:05 2005] [error] env.createBean2(): Factory error creating channel.jni:jni ( channel.jni, jni) [Thu Jun 09 06:07:05 2005] [error] env.createBean2(): Factory error creating vm: ( vm, ) [Thu Jun 09 06:07:05 2005] [error] env.createBean2(): Factory error creating worker.jni:onStartup ( worker.jni, onStartup) [Thu Jun 09 06:07:05 2005] [error] env.createBean2(): Factory error creating worker.jni:onShutdown ( worker.jni, onShutdown) [Fri Jun 10 11:32:27 2005] [error] env.createBean2(): Factory error creating channel.jni:jni ( channel.jni, jni) [Fri Jun 10 11:32:27 2005] [error] env.createBean2(): Factory error creating vm: ( vm, ) [Fri Jun 10 11:32:27 2005] [error] env.createBean2(): Factory error creating worker.jni:onStartup ( worker.jni, onStartup) [Fri Jun 10 11:32:27 2005] [error] env.createBean2(): Factory error creating worker.jni:onShutdown ( worker.jni, onShutdown) [Sun Jun 12 04:04:29 2005] [error] env.createBean2(): Factory error creating channel.jni:jni ( channel.jni, jni) [Sun Jun 12 04:04:29 2005] [error] env.createBean2(): Factory error creating vm: ( vm, ) [Sun Jun 12 04:04:29 2005] [error] env.createBean2(): Factory error creating worker.jni:onStartup ( worker.jni, onStartup) [Sun Jun 12 04:04:29 2005] [error] env.createBean2(): Factory error creating worker.jni:onShutdown ( worker.jni, onShutdown) [Fri Jun 17 04:03:42 2005] [error] env.createBean2(): Factory error creating channel.jni:jni ( channel.jni, jni) [Fri Jun 17 04:03:42 2005] [error] env.createBean2(): Factory error creating vm: ( vm, ) [Fri Jun 17 04:03:42 2005] [error] env.createBean2(): Factory error creating worker.jni:onStartup ( worker.jni, onStartup) [Fri Jun 17 04:03:42 2005] [error] env.createBean2(): Factory error creating worker.jni:onShutdown ( worker.jni, onShutdown) [Sun Jun 19 04:09:06 2005] [error] env.createBean2(): Factory error creating channel.jni:jni ( channel.jni, jni) [Sun Jun 19 04:09:06 2005] [error] env.createBean2(): Factory error creating vm: ( vm, ) [Sun Jun 19 04:09:06 2005] [error] env.createBean2(): Factory error creating worker.jni:onStartup ( worker.jni, onStartup) [Sun Jun 19 04:09:06 2005] [error] env.createBean2(): Factory error creating worker.jni:onShutdown ( worker.jni, onShutdown) [Sat Jun 25 04:04:30 2005] [error] env.createBean2(): Factory error creating channel.jni:jni ( channel.jni, jni) [Sat Jun 25 04:04:30 2005] [error] env.createBean2(): Factory error creating vm: ( vm, ) [Sat Jun 25 04:04:30 2005] [error] env.createBean2(): Factory error creating worker.jni:onStartup ( worker.jni, onStartup) [Sat Jun 25 04:04:30 2005] [error] env.createBean2(): Factory error creating worker.jni:onShutdown ( worker.jni, onShutdown) [Sun Jun 26 04:04:27 2005] [error] env.createBean2(): Factory error creating channel.jni:jni ( channel.jni, jni) [Sun Jun 26 04:04:27 2005] [error] env.createBean2(): Factory error creating vm: ( vm, ) [Sun Jun 26 04:04:27 2005] [error] env.createBean2(): Factory error creating worker.jni:onStartup ( worker.jni, onStartup) [Sun Jun 26 04:04:27 2005] [error] env.createBean2(): Factory error creating worker.jni:onShutdown ( worker.jni, onShutdown) [Mon Jun 27 04:02:53 2005] [error] env.createBean2(): Factory error creating channel.jni:jni ( channel.jni, jni) [Mon Jun 27 04:02:53 2005] [error] env.createBean2(): Factory error creating vm: ( vm, ) [Mon Jun 27 04:02:53 2005] [error] env.createBean2(): Factory error creating worker.jni:onStartup ( worker.jni, onStartup) [Mon Jun 27 04:02:53 2005] [error] env.createBean2(): Factory error creating worker.jni:onShutdown ( worker.jni, onShutdown) [Sun Jul 03 04:07:59 2005] [error] env.createBean2(): Factory error creating channel.jni:jni ( channel.jni, jni) [Sun Jul 03 04:07:59 2005] [error] env.createBean2(): Factory error creating vm: ( vm, ) [Sun Jul 03 04:07:59 2005] [error] env.createBean2(): Factory error creating worker.jni:onStartup ( worker.jni, onStartup) [Sun Jul 03 04:07:59 2005] [error] env.createBean2(): Factory error creating worker.jni:onShutdown ( worker.jni, onShutdown) [Sun Jul 10 04:04:42 2005] [error] env.createBean2(): Factory error creating channel.jni:jni ( channel.jni, jni) [Sun Jul 10 04:04:42 2005] [error] env.createBean2(): Factory error creating vm: ( vm, ) [Sun Jul 10 04:04:42 2005] [error] env.createBean2(): Factory error creating worker.jni:onStartup ( worker.jni, onStartup) [Sun Jul 10 04:04:42 2005] [error] env.createBean2(): Factory error creating worker.jni:onShutdown ( worker.jni, onShutdown) [Sun Jul 17 04:08:19 2005] [error] env.createBean2(): Factory error creating channel.jni:jni ( channel.jni, jni) [Sun Jul 17 04:08:19 2005] [error] env.createBean2(): Factory error creating vm: ( vm, ) [Sun Jul 17 04:08:19 2005] [error] env.createBean2(): Factory error creating worker.jni:onStartup ( worker.jni, onStartup) [Sun Jul 17 04:08:19 2005] [error] env.createBean2(): Factory error creating worker.jni:onShutdown ( worker.jni, onShutdown) [Sun Jul 24 04:20:44 2005] [error] env.createBean2(): Factory error creating channel.jni:jni ( channel.jni, jni) [Sun Jul 24 04:20:44 2005] [error] env.createBean2(): Factory error creating vm: ( vm, ) [Sun Jul 24 04:20:44 2005] [error] env.createBean2(): Factory error creating worker.jni:onStartup ( worker.jni, onStartup) [Sun Jul 24 04:20:44 2005] [error] env.createBean2(): Factory error creating worker.jni:onShutdown ( worker.jni, onShutdown) [Wed Jul 27 14:42:40 2005] [error] env.createBean2(): Factory error creating channel.jni:jni ( channel.jni, jni) [Wed Jul 27 14:42:40 2005] [error] env.createBean2(): Factory error creating vm: ( vm, ) [Wed Jul 27 14:42:40 2005] [error] env.createBean2(): Factory error creating worker.jni:onStartup ( worker.jni, onStartup) [Wed Jul 27 14:42:40 2005] [error] env.createBean2(): Factory error creating worker.jni:onShutdown ( worker.jni, onShutdown) [Sun Jul 31 04:08:56 2005] [error] env.createBean2(): Factory error creating channel.jni:jni ( channel.jni, jni) [Sun Jul 31 04:08:56 2005] [error] env.createBean2(): Factory error creating vm: ( vm, ) [Sun Jul 31 04:08:56 2005] [error] env.createBean2(): Factory error creating worker.jni:onStartup ( worker.jni, onStartup) [Sun Jul 31 04:08:56 2005] [error] env.createBean2(): Factory error creating worker.jni:onShutdown ( worker.jni, onShutdown) [Sun Aug 07 04:03:56 2005] [error] env.createBean2(): Factory error creating channel.jni:jni ( channel.jni, jni) [Sun Aug 07 04:03:56 2005] [error] env.createBean2(): Factory error creating vm: ( vm, ) [Sun Aug 07 04:03:56 2005] [error] env.createBean2(): Factory error creating worker.jni:onStartup ( worker.jni, onStartup) [Sun Aug 07 04:03:56 2005] [error] env.createBean2(): Factory error creating worker.jni:onShutdown ( worker.jni, onShutdown) [Sun Aug 14 04:03:50 2005] [error] env.createBean2(): Factory error creating channel.jni:jni ( channel.jni, jni) [Sun Aug 14 04:03:50 2005] [error] env.createBean2(): Factory error creating vm: ( vm, ) [Sun Aug 14 04:03:50 2005] [error] env.createBean2(): Factory error creating worker.jni:onStartup ( worker.jni, onStartup) [Sun Aug 14 04:03:50 2005] [error] env.createBean2(): Factory error creating worker.jni:onShutdown ( worker.jni, onShutdown) [Sun Aug 21 04:03:46 2005] [error] env.createBean2(): Factory error creating channel.jni:jni ( channel.jni, jni) [Sun Aug 21 04:03:46 2005] [error] env.createBean2(): Factory error creating vm: ( vm, ) [Sun Aug 21 04:03:46 2005] [error] env.createBean2(): Factory error creating worker.jni:onStartup ( worker.jni, onStartup) [Sun Aug 21 04:03:46 2005] [error] env.createBean2(): Factory error creating worker.jni:onShutdown ( worker.jni, onShutdown) [Sun Aug 28 04:10:36 2005] [error] env.createBean2(): Factory error creating channel.jni:jni ( channel.jni, jni) [Sun Aug 28 04:10:36 2005] [error] env.createBean2(): Factory error creating vm: ( vm, ) [Sun Aug 28 04:10:36 2005] [error] env.createBean2(): Factory error creating worker.jni:onStartup ( worker.jni, onStartup) [Sun Aug 28 04:10:36 2005] [error] env.createBean2(): Factory error creating worker.jni:onShutdown ( worker.jni, onShutdown) [Sun Sep 04 04:04:09 2005] [error] env.createBean2(): Factory error creating channel.jni:jni ( channel.jni, jni) [Sun Sep 04 04:04:09 2005] [error] env.createBean2(): Factory error creating vm: ( vm, ) [Sun Sep 04 04:04:09 2005] [error] env.createBean2(): Factory error creating worker.jni:onStartup ( worker.jni, onStartup) [Sun Sep 04 04:04:09 2005] [error] env.createBean2(): Factory error creating worker.jni:onShutdown ( worker.jni, onShutdown) [Sun Sep 11 04:03:50 2005] [error] env.createBean2(): Factory error creating channel.jni:jni ( channel.jni, jni) [Sun Sep 11 04:03:50 2005] [error] env.createBean2(): Factory error creating vm: ( vm, ) [Sun Sep 11 04:03:50 2005] [error] env.createBean2(): Factory error creating worker.jni:onStartup ( worker.jni, onStartup) [Sun Sep 11 04:03:50 2005] [error] env.createBean2(): Factory error creating worker.jni:onShutdown ( worker.jni, onShutdown) [Sun Sep 18 04:07:22 2005] [error] env.createBean2(): Factory error creating channel.jni:jni ( channel.jni, jni) [Sun Sep 18 04:07:22 2005] [error] env.createBean2(): Factory error creating vm: ( vm, ) [Sun Sep 18 04:07:22 2005] [error] env.createBean2(): Factory error creating worker.jni:onStartup ( worker.jni, onStartup) [Sun Sep 18 04:07:22 2005] [error] env.createBean2(): Factory error creating worker.jni:onShutdown ( worker.jni, onShutdown) [Mon Sep 19 05:58:19 2005] [error] env.createBean2(): Factory error creating channel.jni:jni ( channel.jni, jni) [Mon Sep 19 05:58:19 2005] [error] env.createBean2(): Factory error creating vm: ( vm, ) [Mon Sep 19 05:58:19 2005] [error] env.createBean2(): Factory error creating worker.jni:onStartup ( worker.jni, onStartup) [Mon Sep 19 05:58:19 2005] [error] env.createBean2(): Factory error creating worker.jni:onShutdown ( worker.jni, onShutdown) [Sun Sep 25 04:12:23 2005] [error] env.createBean2(): Factory error creating channel.jni:jni ( channel.jni, jni) [Sun Sep 25 04:12:23 2005] [error] env.createBean2(): Factory error creating vm: ( vm, ) [Sun Sep 25 04:12:23 2005] [error] env.createBean2(): Factory error creating worker.jni:onStartup ( worker.jni, onStartup) [Sun Sep 25 04:12:23 2005] [error] env.createBean2(): Factory error creating worker.jni:onShutdown ( worker.jni, onShutdown) [Wed Sep 28 09:11:52 2005] [error] env.createBean2(): Factory error creating channel.jni:jni ( channel.jni, jni) [Wed Sep 28 09:11:52 2005] [error] env.createBean2(): Factory error creating vm: ( vm, ) [Wed Sep 28 09:11:52 2005] [error] env.createBean2(): Factory error creating worker.jni:onStartup ( worker.jni, onStartup) [Wed Sep 28 09:11:52 2005] [error] env.createBean2(): Factory error creating worker.jni:onShutdown ( worker.jni, onShutdown) [Sun Oct 02 04:06:48 2005] [error] env.createBean2(): Factory error creating channel.jni:jni ( channel.jni, jni) [Sun Oct 02 04:06:48 2005] [error] env.createBean2(): Factory error creating vm: ( vm, ) [Sun Oct 02 04:06:48 2005] [error] env.createBean2(): Factory error creating worker.jni:onStartup ( worker.jni, onStartup) [Sun Oct 02 04:06:48 2005] [error] env.createBean2(): Factory error creating worker.jni:onShutdown ( worker.jni, onShutdown) [Sun Oct 09 04:12:16 2005] [error] env.createBean2(): Factory error creating channel.jni:jni ( channel.jni, jni) [Sun Oct 09 04:12:16 2005] [error] env.createBean2(): Factory error creating vm: ( vm, ) [Sun Oct 09 04:12:16 2005] [error] env.createBean2(): Factory error creating worker.jni:onStartup ( worker.jni, onStartup) [Sun Oct 09 04:12:16 2005] [error] env.createBean2(): Factory error creating worker.jni:onShutdown ( worker.jni, onShutdown) [Sun Oct 16 04:14:44 2005] [error] env.createBean2(): Factory error creating channel.jni:jni ( channel.jni, jni) [Sun Oct 16 04:14:44 2005] [error] env.createBean2(): Factory error creating vm: ( vm, ) [Sun Oct 16 04:14:44 2005] [error] env.createBean2(): Factory error creating worker.jni:onStartup ( worker.jni, onStartup) [Sun Oct 16 04:14:44 2005] [error] env.createBean2(): Factory error creating worker.jni:onShutdown ( worker.jni, onShutdown) [Sun Oct 23 04:04:04 2005] [error] env.createBean2(): Factory error creating channel.jni:jni ( channel.jni, jni) [Sun Oct 23 04:04:04 2005] [error] env.createBean2(): Factory error creating vm: ( vm, ) [Sun Oct 23 04:04:04 2005] [error] env.createBean2(): Factory error creating worker.jni:onStartup ( worker.jni, onStartup) [Sun Oct 23 04:04:04 2005] [error] env.createBean2(): Factory error creating worker.jni:onShutdown ( worker.jni, onShutdown) [Tue Oct 25 10:09:37 2005] [error] env.createBean2(): Factory error creating channel.jni:jni ( channel.jni, jni) [Tue Oct 25 10:09:37 2005] [error] env.createBean2(): Factory error creating vm: ( vm, ) [Tue Oct 25 10:09:37 2005] [error] env.createBean2(): Factory error creating worker.jni:onStartup ( worker.jni, onStartup) [Tue Oct 25 10:09:37 2005] [error] env.createBean2(): Factory error creating worker.jni:onShutdown ( worker.jni, onShutdown) [Sun Oct 30 04:04:47 2005] [error] env.createBean2(): Factory error creating channel.jni:jni ( channel.jni, jni) [Sun Oct 30 04:04:47 2005] [error] env.createBean2(): Factory error creating vm: ( vm, ) [Sun Oct 30 04:04:47 2005] [error] env.createBean2(): Factory error creating worker.jni:onStartup ( worker.jni, onStartup) [Sun Oct 30 04:04:47 2005] [error] env.createBean2(): Factory error creating worker.jni:onShutdown ( worker.jni, onShutdown) [Sun Nov 06 04:06:10 2005] [error] env.createBean2(): Factory error creating channel.jni:jni ( channel.jni, jni) [Sun Nov 06 04:06:10 2005] [error] env.createBean2(): Factory error creating vm: ( vm, ) [Sun Nov 06 04:06:10 2005] [error] env.createBean2(): Factory error creating worker.jni:onStartup ( worker.jni, onStartup) [Sun Nov 06 04:06:10 2005] [error] env.createBean2(): Factory error creating worker.jni:onShutdown ( worker.jni, onShutdown) [Sun Nov 13 04:15:24 2005] [error] env.createBean2(): Factory error creating channel.jni:jni ( channel.jni, jni) [Sun Nov 13 04:15:25 2005] [error] env.createBean2(): Factory error creating vm: ( vm, ) [Sun Nov 13 04:15:25 2005] [error] env.createBean2(): Factory error creating worker.jni:onStartup ( worker.jni, onStartup) [Sun Nov 13 04:15:25 2005] [error] env.createBean2(): Factory error creating worker.jni:onShutdown ( worker.jni, onShutdown) [Sun Nov 20 04:10:30 2005] [error] env.createBean2(): Factory error creating channel.jni:jni ( channel.jni, jni) [Sun Nov 20 04:10:30 2005] [error] env.createBean2(): Factory error creating vm: ( vm, ) [Sun Nov 20 04:10:30 2005] [error] env.createBean2(): Factory error creating worker.jni:onStartup ( worker.jni, onStartup) [Sun Nov 20 04:10:30 2005] [error] env.createBean2(): Factory error creating worker.jni:onShutdown ( worker.jni, onShutdown) [Sun Nov 27 04:07:32 2005] [error] env.createBean2(): Factory error creating channel.jni:jni ( channel.jni, jni) [Sun Nov 27 04:07:32 2005] [error] env.createBean2(): Factory error creating vm: ( vm, ) [Sun Nov 27 04:07:32 2005] [error] env.createBean2(): Factory error creating worker.jni:onStartup ( worker.jni, onStartup) [Sun Nov 27 04:07:32 2005] [error] env.createBean2(): Factory error creating worker.jni:onShutdown ( worker.jni, onShutdown) [Sun Dec 04 04:09:43 2005] [error] env.createBean2(): Factory error creating channel.jni:jni ( channel.jni, jni) [Sun Dec 04 04:09:43 2005] [error] env.createBean2(): Factory error creating vm: ( vm, ) [Sun Dec 04 04:09:43 2005] [error] env.createBean2(): Factory error creating worker.jni:onStartup ( worker.jni, onStartup) [Sun Dec 04 04:09:43 2005] [error] env.createBean2(): Factory error creating worker.jni:onShutdown ( worker.jni, onShutdown) [Tue Dec 06 12:24:09 2005] [error] env.createBean2(): Factory error creating channel.jni:jni ( channel.jni, jni) [Tue Dec 06 12:24:09 2005] [error] env.createBean2(): Factory error creating vm: ( vm, ) [Tue Dec 06 12:24:09 2005] [error] env.createBean2(): Factory error creating worker.jni:onStartup ( worker.jni, onStartup) [Tue Dec 06 12:24:09 2005] [error] env.createBean2(): Factory error creating worker.jni:onShutdown ( worker.jni, onShutdown) [Sun Dec 11 04:06:46 2005] [error] env.createBean2(): Factory error creating channel.jni:jni ( channel.jni, jni) [Sun Dec 11 04:06:46 2005] [error] env.createBean2(): Factory error creating vm: ( vm, ) [Sun Dec 11 04:06:46 2005] [error] env.createBean2(): Factory error creating worker.jni:onStartup ( worker.jni, onStartup) [Sun Dec 11 04:06:46 2005] [error] env.createBean2(): Factory error creating worker.jni:onShutdown ( worker.jni, onShutdown) [Sun Dec 18 04:02:18 2005] [error] env.createBean2(): Factory error creating channel.jni:jni ( channel.jni, jni) [Sun Dec 18 04:02:18 2005] [error] env.createBean2(): Factory error creating vm: ( vm, ) [Sun Dec 18 04:02:18 2005] [error] env.createBean2(): Factory error creating worker.jni:onStartup ( worker.jni, onStartup) [Sun Dec 18 04:02:18 2005] [error] env.createBean2(): Factory error creating worker.jni:onShutdown ( worker.jni, onShutdown) [Sun Dec 25 04:02:18 2005] [error] env.createBean2(): Factory error creating channel.jni:jni ( channel.jni, jni) [Sun Dec 25 04:02:18 2005] [error] env.createBean2(): Factory error creating vm: ( vm, ) [Sun Dec 25 04:02:18 2005] [error] env.createBean2(): Factory error creating worker.jni:onStartup ( worker.jni, onStartup) [Sun Dec 25 04:02:18 2005] [error] env.createBean2(): Factory error creating worker.jni:onShutdown ( worker.jni, onShutdown) [Sun Jan 01 04:02:17 2006] [error] env.createBean2(): Factory error creating channel.jni:jni ( channel.jni, jni) [Sun Jan 01 04:02:17 2006] [error] env.createBean2(): Factory error creating vm: ( vm, ) [Sun Jan 01 04:02:17 2006] [error] env.createBean2(): Factory error creating worker.jni:onStartup ( worker.jni, onStartup) [Sun Jan 01 04:02:17 2006] [error] env.createBean2(): Factory error creating worker.jni:onShutdown ( worker.jni, onShutdown) [Sat Jan 07 04:02:37 2006] [error] env.createBean2(): Factory error creating channel.jni:jni ( channel.jni, jni) [Sat Jan 07 04:02:37 2006] [error] env.createBean2(): Factory error creating vm: ( vm, ) [Sat Jan 07 04:02:37 2006] [error] env.createBean2(): Factory error creating worker.jni:onStartup ( worker.jni, onStartup) [Sat Jan 07 04:02:37 2006] [error] env.createBean2(): Factory error creating worker.jni:onShutdown ( worker.jni, onShutdown) [Sun Jan 08 04:05:56 2006] [error] env.createBean2(): Factory error creating channel.jni:jni ( channel.jni, jni) [Sun Jan 08 04:05:56 2006] [error] env.createBean2(): Factory error creating vm: ( vm, ) [Sun Jan 08 04:05:56 2006] [error] env.createBean2(): Factory error creating worker.jni:onStartup ( worker.jni, onStartup) [Sun Jan 08 04:05:56 2006] [error] env.createBean2(): Factory error creating worker.jni:onShutdown ( worker.jni, onShutdown) [Sun Jan 15 04:11:23 2006] [error] env.createBean2(): Factory error creating channel.jni:jni ( channel.jni, jni) [Sun Jan 15 04:11:23 2006] [error] env.createBean2(): Factory error creating vm: ( vm, ) [Sun Jan 15 04:11:23 2006] [error] env.createBean2(): Factory error creating worker.jni:onStartup ( worker.jni, onStartup) [Sun Jan 15 04:11:23 2006] [error] env.createBean2(): Factory error creating worker.jni:onShutdown ( worker.jni, onShutdown) [Sun Jan 22 04:11:00 2006] [error] env.createBean2(): Factory error creating channel.jni:jni ( channel.jni, jni) [Sun Jan 22 04:11:00 2006] [error] env.createBean2(): Factory error creating vm: ( vm, ) [Sun Jan 22 04:11:00 2006] [error] env.createBean2(): Factory error creating worker.jni:onStartup ( worker.jni, onStartup) [Sun Jan 22 04:11:00 2006] [error] env.createBean2(): Factory error creating worker.jni:onShutdown ( worker.jni, onShutdown) [Thu Jan 26 12:23:07 2006] [error] env.createBean2(): Factory error creating channel.jni:jni ( channel.jni, jni) [Thu Jan 26 12:23:07 2006] [error] env.createBean2(): Factory error creating vm: ( vm, ) [Thu Jan 26 12:23:07 2006] [error] env.createBean2(): Factory error creating worker.jni:onStartup ( worker.jni, onStartup) [Thu Jan 26 12:23:07 2006] [error] env.createBean2(): Factory error creating worker.jni:onShutdown ( worker.jni, onShutdown) ``` ### Exercise 1.02 – about those Factory errors How many have there been over the life of this log file? What proportion is that of the whole file? What looks like the most-common day of the week that it happens on – and what can you say about that? #### Hint: `wc -l «filename»` or `«some stream» | wc -l` will count the lines in the file or the stream. So aim your `grep` at it. Or you can put it into an editor with line numbers, or eyeball it. #### Solution to Exercise 1.02 You might use `grep -i "Factory" Apache.log | wc -l` – it's 180, and the number of lines in the log is 56481 from `wc -l Apache.log`. You might use `grep -i "Factory" Apache.log | grep "Sun" | wc -l` – there were 132 on Sundays. Kudos if you looked at the results to check and spotted they were all at 4am. ### Exercise 1.03 – reverse filter / line numbers *unfinished* This log has lines without a timestamp #### Solution to Exercise 1.03 how many? `grep -v "\[" /tmp/Apache.log | wc -l` where? `grep -v "\[" -n /tmp/Apache.log` (note - the \\ may not copy/paste well with code-server) --- The exercises below rely on a working log simulator. For people with James, it's at `/tmp/simulated_log.txt`. People without will need to run it for themselves. ### Exercise 1.04 – real-time filtering *unpolished* If you `tail -F /tmp/simulated_log.txt`, you'll see a fast-moving log. (the log is from `logmaker.sh`, details on [making logs to teach testing](https://www.workroom-productions.com/making-logs-for-teaching/)) How can you use `grep` to help you to filter only the `failure`s? #### Solution to Exercise 1.04 – real-time filtering `tail -F /tmp/simulated_log.txt | grep failure ` How might this be useful to you, as a tester? ### Exercise 1.05 – Chains *unfinished* how about all failures in CORE? (lots more chains here) #### Solution to Exercise 1.05 – chains `tail -F /tmp/simulated_log.txt | grep -i failure | grep -i CORE` --- *This page is a work in progress. I'm sure that a curious reader can easily find my notes to myself. «JL: perhaps, if the script finds any notes, it can insert this message?»* ### `grep` for testers URL: https://www.workroom-productions.com/grep-for-testers/ Last updated: 2025-02-07T10:47:53.000Z [grep](https://en.wikipedia.org/wiki/Grep) finds lines that match a pattern – so it's used when you want to sift useful information out of something bigger. [grep](https://en.wikipedia.org/wiki/Grep) works on *streams* rather than necessarily on *files*, so can be chained with tools which output information. *Example:* `ls | grep .json` will list the names of all the files in the current directory\* (`ls`), and only output those names which contain `.json` . That pipe `|` connects an output to an input, reading left to right, so connects the output of `ls` to the input of `grep`. *\* current directory – if you're working on the command line, you're always working on one, and only one, folder or directory.* [grep](https://en.wikipedia.org/wiki/Grep) outputs a stream, so it can be used in the middle of a chain of tools. *Example:* `ls -R | grep .json | wc -l` will list the names of all the files in the current directory and below (`ls -R`), filter just the ones that contain `.json`, then count those lines – giving you a count of all those files. `grep` will also work on files. You put a filename after the thing you're looking for. Think *`grep`for`«a thing»` in`(these) file(s)`*. Example: `grep 'Factory Error' apache.log` will show all lines containing `Factory error` in a file called `Apache.log` in your current directory. [grep](https://en.wikipedia.org/wiki/Grep)'s searching is based on [regex](https://en.wikipedia.org/wiki/Regular%5Fexpression), so you may need to work (and test) to get the search right before trusting the output. Don't live with the pain: use a regex painter like [regexr](https://regexr.com), [regex generator](https://regex-generator.olafneumann.org/?sampleText=2020-03-12T13%3A34%3A56.123Z%20INFO%20%20%5Borg.example.Class%5D%3A%20This%20is%20a%20%23simple%20%23logline%20containing%20a%20%27value%27.&flags=i), [regex101](https://regex101.com). ## Handy `grep` for testers Some things I've used recently (I'll add to this...). ### Is a process running? `ps -ef | grep [a]pache` will show you any processes running called `apache`... more usefully, it will show nothing if `apache` isn't running. If you just `ps -ef | grep apache` (as I did until writing this page), you'll get back the process that is looking for the word `apache`, too. You'll get used to this behaviour, but it's less good if you're automating. Handy if you expect your process to be running, and you don't want to scan the list. Obvs you need to know your process name, and to be happily on the command line. Use `ps aux` if you prefer. ### What's the context for any fatal errors? `grep -C 10 fatal Apache.log` would show you any line in `Apache.log` with the word `fatal` in it, and ten lines above and below. Change to suit your logs: I've never seen `fatal` in an Apache log. Handy for seeing what happened immediately before and after a problem. ### Where are the logs? `ls -R | grep log` will find all file names / directory names with the characters `log` in them, in and below your current directory. Handy if you need a swift scan for files with particular names, and you find `find` confusing. As I do. `grep -R log .` will find all files with the characters `log` anywhere in the file, in all the files in and below your current directory. #### A note on finding files and recursion `ls -r` produces a list of files, and `grep` just sees that list as text. So `grep` doesn't know the path or the file context – and it just shows the filename. `grep -r «whatever» .` works recursively through a directory: it knows the path and looks inside the file for the content, not at the name If you want to use `grep` to search filenames, it'll work but it's compromised. I might use `find . -type f -name "*.log"`, which would find all files (and only files) ending with the characters `.log`. So I'm putting that here to remind me. ### What files contain the word `forbidden`? `grep -r forbidden .` will find `forbidden` in all files below this directory (`-r` for recursive – or you can use `-R` for differently recursive), while... `grep -r -i forbidden .` is case-insensitive and will find `Forbidden`, `fOrBiDdEn` etc. Handy for spotting what code references specific modules, functions or data; for pulling out all error messages, text or regex; for highlighting comments with `todo` or `bug`, for searching logs and manuals... Most IDEs already give you the facility to search across all files in a project. Using `grep` is handier if you're using something to *process* what you find. ### *Variant:* What IP addresses are specified directly? `grep -r -E "\b([0-9]{1,3}\.){3}[0-9]{1,3}\b" .` will look in lots of files for that (`-E` means *extended regex) pattern*. That confusing barrage of detail (ask an LLM to parse it) should pickup IP addresses (and some things that look like IP addresses, and aren't). Handy for spotting when your code has hard-coded IP addresses, or for grabbing the things out of a directory full of config files, or from logs. ### Show me whenever the log file flags an error Using a dedicated terminal window, and in the right directory, run the command `tail -F *«logfile.log»* | grep -i -E "error|fatal|failure|spin"` and that window will be a realtime view of the most-recent end of the log, and will show only those lines which include the selected words. Note that we're using `tail -F`, not `tail -f`, which should follow the logfile if something else changes it by making it a new file. ### The files are compressed! Use [zgrep](https://manpages.org/zgrep) to look inside compressed files, too. It may not extract the info, but it will tell you that it's there. ### I want everything *except* what I'm asking for Use `grep -v` to return lots of stuff. Maybe you're seeking every line in the log which doesn't have a specific IP address. ### I want to look for two things and I don't want to write regex. Good choice. Regex is human-unreadable. Stack your `grep`s. Example: `tail -F «log» | grep "failure" | grep "core"` ### Two things, one or the other? Example: `tail -F «log» | grep -e "failure" -e "core"` or: `tail -F «log» | grep -e "failure\|core"`, which has fewer keystrokes but is harder to read. ## More ### Playgrounds – interactive grep in browser [GNU grep live editor![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/icon/gnu-head-mini.png)GNU grep REPL](https://grep.js.org) [Grep playgroundEmbeddable grep playground for education, documentation, and fun.![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/icon/favicon-1.svg)Anton Zhiyanov![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/thumbnail/cover.png)](https://codapi.org/grep/) ### Workroom Exercises [Grep Tool ExercisesPlay with \`grep\` – solutions available.![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/icon/wpl-60x60-logo-semitrans-6.png)Workroom ProductionsJames Lyndsay![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/thumbnail/photo-1511225317751-5c2d61819d58-1.jpeg)](https://www.workroom-productions.com/grep-tool-exercises/) ### Other people's interactive exercises [Interactive exercises for GNU grep, sed and awk (TUI apps)Interactive exercises to test your Linux CLI text processing skills for the GNU grep, sed and awk commands.![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/icon/favicon.svg)learnbyexample![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/thumbnail/grep_exercises.png)](https://learnbyexample.github.io/interactive-grep-sed-awk-exercises/) [Try grep in Y minutes![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/icon/favicon-2.svg)Anton Zhiyanov![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/thumbnail/cover-1.png)](https://codapi.org/try/grep/) ### Alternatives / similar tools `grep` is >50 years old. These aren't. Currently. [GitHub - BurntSushi/ripgrep: ripgrep recursively searches directories for a regex pattern while respecting your gitignoreripgrep recursively searches directories for a regex pattern while respecting your gitignore - BurntSushi/ripgrep![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/icon/pinned-octocat-093da3e6fa40.svg)GitHubBurntSushi![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/thumbnail/ripgrep)](https://github.com/BurntSushi/ripgrep) [The ugrep file pattern searcherThe ugrep file pattern searcher -- a more powerful, ultra fast, user-friendly, compatible grep replacement![](https://static.ghost.org/v5.0.0/images/link-icon.svg)Robert A. van Engelen, Genivia Inc![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/thumbnail/ug.png)](https://ugrep.com) [GitHub - ggreer/the\_silver\_searcher: A code-searching tool similar to ack, but faster.A code-searching tool similar to ack, but faster. Contribute to ggreer/the\_silver\_searcher development by creating an account on GitHub.![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/icon/pinned-octocat-093da3e6fa40-1.svg)GitHubggreer![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/thumbnail/the_silver_searcher)](https://github.com/ggreer/the%5Fsilver%5Fsearcher) [Git - git-grep Documentation![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/icon/favicon-1.ico)git-grep Documentation![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/thumbnail/green-dot.png)](https://git-scm.com/docs/git-grep) ### Sources and quickreferences [grep - Wikipedia![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/icon/wikipedia.png)Wikimedia Foundation, Inc.Contributors to Wikimedia projects![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/thumbnail/Grep_example.png)](https://en.wikipedia.org/wiki/Grep) [grep(1) - Linux manual page![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/icon/faviconV2-1)Linux manual page![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/thumbnail/TLPI-front-cover-vsmall.png)](https://www.man7.org/linux/man-pages/man1/grep.1.html) [grep![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/icon/favicon-2.ico)![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/thumbnail/opt-start.gif)](https://pubs.opengroup.org/onlinepubs/9799919799/utilities/grep.html) [Grep Command Cheat Sheet & Quick ReferenceThis cheat sheet is intended to be a quick reminder for the main concepts involved in using the command line p![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/icon/favicon.png)QuickRef.ME![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/thumbnail/grep-preview.png)](https://quickref.me/grep) [How to use grep (with examples)Grep is a powerful utility on Linux. Want to get more out of the tool? This article will show you how to use it including many practical examples.![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/icon/faviconV2)Linux AuditMichael Boelen![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/thumbnail/background_hu12073633397429872161.png)](https://linux-audit.com/grep-commands-and-common-examples-for-daily-use/) [Grep Command In Unix/Linux with 25+ Examples \[2023\]Let’s Learn Grep Command in Unix/Linux with Simple Examples : New Linux users get anxious when confronted with the prospect of searching for a particular string in a file. Some have no idea what to or where to start. Here is an article to make your work a little easier. grep command in unix/linux is \[…\]![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/icon/fav.png)Linux TeacherWinnie Winnie![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/thumbnail/grep-command-in-unix-800x600.jpg)](https://www.linuxteacher.com/grep-command-in-unix-linux-with-examples/) [tldr InBrowser.Apptldr InBrowser.App is an offline-capable PWA for tldr-pages. Fully runs in your browser. Zero API latency.![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/icon/apple-touch-icon.png)](https://tldr.inbrowser.app/pages/common/grep) ### PlayTime 004 – Puzzle31 URL: https://www.workroom-productions.com/playtime-004-blackbox-puzzle031/ Last updated: 2025-01-28T14:36:20.000Z This week's Workroom PlayTime is an exploration of my BlackBox puzzle: [Puzzle031](https://blackboxpuzzles.workroomprds.com/puzzle31/). It is at 4pm on Friday 31 January 2025 (and here's [that time where you are](https://this-ti.me/?uts=1738339200&tz=Europe%2FLondon&name=Workroom+PlayTime)). Workroom PlayTime this week is open to all website subscribers, and to my [BlackBox Puzzle Patreon](https://patreon.com/workroomprds) patrons. _This post is for subscribers only._ ### Workroom PlayTime this week is "Following Up" (from SpeakerPrep) URL: https://www.workroom-productions.com/workroom-playtime-003-subscribers-newsletter/ Last updated: 2025-01-22T15:58:06.000Z Workroomprds newsletter: Friday 5:30 London time for "Following up with your audience exercise". _This post is for subscribers only._ ### PlayTime 003 – SpeakerPrep – Following Up URL: https://www.workroom-productions.com/playtime-003-speakerprep-following-up/ Last updated: 2025-01-21T17:08:27.000Z We'll gather on Zoom / Miro, on Friday 24 January at 5:30pm London time ([local time for you](https://this-ti.me/?uts=1737739800&tz=Europe%2FLondon&name=Workroom+Playtime+003)). We'll do a short exercise from SpeakerPrep – [Following up with your audience](https://www.workroom-productions.com/what-to-follow-up-after-a-talk/). It's a practical one that will help you to build connections with the most-engaged people in your audience. These exercises are for everyone, for free. Paying subscribers get to play together. _This post is for paying subscribers only._ ### What to Follow Up after a Talk URL: https://www.workroom-productions.com/what-to-follow-up-after-a-talk/ Last updated: 2025-01-21T16:45:58.000Z This is from [SpeakerPrep](https://www.workroom-productions.com/tag/speakerprep/). If you've got a talk, the communication doesn't stop when you step off stage. Indeed, you may be on stage in order to build awareness, a community, or your own brand: You need to engage further with those in the audience who are most engaged. In this exercise, we'll work on how to do that. #### The 15-minute variant for [PlayTime](https://www.workroom-productions.com/workroom-playtime/) First 5 mins – ****What's worked?** **We'll share some ways that speakers have engaged with us, after a talk.* Middle 5 mins – ****Build something to invite contact.** **We'll all pick out one practical thing, that we've might try doing after delivering some sort of presentation. Use ideas from anywhere; there's a list below if needed.* → Post a short note on the board. Last 5 mins – ****Share** **3-5 contributors to tell us what and why* --- ## SpeakerPrep version In this, we'll build something to invite contact in one or more ways. We'll have a conversation about what's worked (and why), when a speaker has followed up with you – and what has not worked (and why). We'll talk about what we might want to try, then spend time building solo or in small groups, to see where we can take the idea. At the end, we'll share (perhaps with accountability partners) our plans. Browse our list of things that you might do to follow up. Drop us a line (or a PR) to add your own! - Take contact details of anyone who talks to you (contacts you) after the talk. Follow up soon after, then in a couple of weeks. - Offer materials in exchange for emails - Invite people to meet you in the pub / expo / wherever to talk further - Offer stickers with your catchy slogan, talk to people wearing your stickers - Put your social handle on each slide, invite new followers, interact with any new followers over the event. - Initiate contact with anyone who mentions your talk on socials. - Invite feedback and comments. - Make an interactive exercise and invite people to use them. - Put materials on GitHub and invite contributions - Set up a Slack / WhatsApp / regular Zoom to take your talk further - Give participants something they can use, offer your help as they put it into practice - Ask for help – offer a call-to-action for people who share your goals - Offer to give the talk as a webinar for people's workplaces - Review how the delivery went - Make a note of ideas about how it might go next time – what you'd change - Consider how your material might work for a different audience ### Workroom PlayTime 002: Exploratory Interfaces I URL: https://www.workroom-productions.com/workroom-playtime-002-2/ Last updated: 2025-01-16T10:53:27.000Z **This week's Workroom PlayTime is on Friday at 3pm London time.** We'll dig into ideas around [Exploratory Interfaces](https://www.workroom-productions.com/exploratory-interfaces/), and will play with the exercise [Changing Input Mechanisms](https://www.workroom-productions.com/changing-input-mechanisms/). Those links will take you to new pages on the site. The pages are for everyone – having subscribers encourages me to make them, and to make them available to our whole community. Thank you for *your* encouragement. Paying subscribers get more interaction. If you go to the [PlayTime 002](https://www.workroom-productions.com/playtime-002-exploratory-interfaces-i/) page, and you're a paying subscriber, you'll find you have access to Zoom and Miro so we can meet and play together on Friday. **If you want to come play, hit** [**upgrade**](https://www.workroom-productions.com/#/portal/) **– it's currently a fiver a month.** --- In a change of tack for **next week's PlayTime**, I'll run a SpeakerPrep exercise from the set that Bart Knaack and I have run at events over the last few years. See [Workroom PlayTime](https://www.workroom-productions.com/workroom-playtime/) for more. --- Here's **news**: I'm going to teach a double-length interactive workshop at [EuroSTAR in Edinburgh in June](https://conference.eurostarsoftwaretesting.com/conference/2025/programme/)! The workshop, [*Test-Driven Generation, a Hands-on Experience*](https://www.workroom-productions.com/test-driven-generation-a-hands-on-experience-at-eurostar-2025-2/), is an extension of the [one that Bart and I delivered](https://www.workroom-productions.com/ai-vs-tdd-atd2024/) at [Agile Testing Days](https://agiletestingdays.com/2024/program/?day=conf-day-1). *I'm thrilled to be on the program.* Ping me if you're going, and want to meet. --- I've had a **flurry of new writing and exercises for the site**. You'll find new content at [Exploration Without Interaction](https://www.workroom-productions.com/exploration-without-interaction-2/), the start of a long topic at [Exploratory Interfaces](https://www.workroom-productions.com/exploratory-interfaces/), and three new exercises to play with: [How do You Explore a System](https://www.workroom-productions.com/exercise-how-do-you-explore-what-a-system-does-2/), [You vs the Machine](https://www.workroom-productions.com/you-vs-the-machine/), and [Changing Input Mechanisms](https://www.workroom-productions.com/changing-input-mechanisms/). Two of those exercises link to toys which you won't have seen before unless you've been to an in-person workshop. You may also be interested in a big new addition to [Exploring Without Requirements](https://www.workroom-productions.com/exploring-without-requirements/). --- Finally, **shout out** to [Elizabeth Zagroba](https://elizabethzagroba.com), who pointed me to an embarrassing 404 on subscribe and to the dodgy email address used on earlier mailouts, and to [Lisa Crispin](https://lisacrispin.com), who mentioned that a calendar link for PlayTime would actually help. Public **gratitude** to Lisa again for (I believe) mentioning PlayTest in the Women In Test group, and to [MoT's](https://www.ministryoftesting.com) [Simon Tomes](https://club.ministryoftesting.com/u/simon%5Ftomes/summary) for having me on the [Week in Testing](https://www.ministryoftesting.com/this-week-in-testings) show last Friday. Have an excellent rest of the week! Cheers – James ### 'Test-driven Generation: A Hands-on Experience' at EuroSTAR 2025 URL: https://www.workroom-productions.com/test-driven-generation-a-hands-on-experience-at-eurostar-2025-2/ Last updated: 2025-01-16T09:38:16.000Z I'm delighted and astonished to be doing a big interactive at [EuroSTAR in Edinburgh this summer](https://conference.eurostarsoftwaretesting.com/conference/2025/). In my session, you’ll write tests that fail, and your AI will hand back code which passes those tests. You’ll go round this magic loop several times; building more tests and generating more code. We'll explore, we'll play, and we'll use the experience as a springboard to think about confirmatory tests and about AI's role in a tester's toolkit. I get to do the first "Deep Dive" slot, on the Wednesday morning. My workshop that means I need to deliver >100 simultaneous, competent, AI-enabled – yet resilient and simple – development environments to 100+ people. [Test-driven Generation: A Hands-on Experience | EuroSTAR ConferenceJoin James Lyndsay in this hands-on session to write failing tests and generate passing code through an iterative AI-driven process![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/icon/cropped-es-fav-icon-270x270.png)EuroSTAR Conference![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/thumbnail/ES2025-James-Lyndsay-DEEP-DIVE-Feature-Image.jpg)](https://conference.eurostarsoftwaretesting.com/event/2025/test-driven-generation-a-hands-on-experience/) I can do that. Preparation's going to be a giggle though... ### Exercise: Changing Input Mechanisms URL: https://www.workroom-productions.com/changing-input-mechanisms/ Last updated: 2025-01-16T19:00:46.000Z *15 Minutes.* > *This exercise is being used for* [*Workroom PlayTime 002*](https://www.workroom-productions.com/workroom-playtime-002-2/)*, Friday 17 Jan at 3pm London time, on Zoom. Become a paying subscriber to join us.* We'll explore [Converter\_v3](https://exercises.workroomprds.com/converter%5Fv3%5Fbase/). We'll do this together, mics / cameras on, talking. ## Exercise: Input field vs Numeric Input field *5 minutes* Pick the radio button for "Numeric Field" to see two inputs; - the one above is a standard free input field, and has an event that updates on `keyUp` - the one below is a standard numeric input, and has an event that updates `onenter`. Both are connected to the same processing code – and the result returned by that code is put below the relevant input. Play with the two – how do you feel about the difference. How is your exploration affected? Are there limits to what you can do? New techniques? ## Exercise: Input field vs Slider *5 minutes* Pick the radio button for "Slider" to switch the numeric input for a simple slider. - the one above is a standard free input field, and has an event that updates on `keyUp` (as before) - the one below is a simple, standard range input (slider) and has an event that updates `oninput`. Play with the two – how do you feel about the difference. How is your exploration affected? Are there limits to what you can do? New techniques? ## Debrief *5 minutes* Consider speed, precision and carelessness, opportunity and restriction, action and observation. Recall that you're interacting with the same code, whatever the interface. Identify differences in your interactions with the system under test in the three circumstances. Identify ways that the different interfaces have affected this exploration. ### Exercise: You vs the Machine URL: https://www.workroom-productions.com/you-vs-the-machine/ Last updated: 2025-01-15T21:20:45.000Z *15-60 minutes.* In this exercise, you'll find an opportunity that is uniquely yours. Make free use of the examples and use [Cadence Keeper](https://exercises.workroomprds.com/CadenceKeeper/) for short durations. For longer exercises, dig into something more personal. ⚠️ **I've designed this as primarily a thinking and sharing exercise, with some doing. If you find yourself lost in testing, you're working in a way that I've not intended.* ## Build *(half the available time)* Scribble down a few things about what it *is*, and what it *does*. Perhaps you want to add *what it's for*. #### Examples Cadence Keeper lets you track how long it has been since you did things, and lets you keep and manipulate a list of those. It's a web page with JavaScript, and uses the browser to store information. **What it is:* Code (HTML / CSS / JS) – something made by James – an exercise source – something like a habit-tracking tool – generated code – not ready for production **What it does:* Allows you to CRUD items in a list – accepts input – tracks time – keeps data when the page is closed Pick a few items from your collection. Identify generic artefacts that you can in some way handle. These are your [*exploratory surfaces*](https://www.workroom-productions.com/exploratory-interfaces/). #### Examples **Code*: \-> several JavaScript scripts, linked HTML and CSS files (can be handled by enumerating vis DevTools), \--> within those there are identifiers with meanings (handle by looking within DevTools / debugger to look for objects with properties, or search code – especially the HML for names and the CSS for classes that might imply functions). \--> error management (handle by looking for `console.log` or `alert` or `assert` or `try...catch`, or looking for text). \--> structure (handle by throwing at a tool like [CodeToFlow](https://codetoflow.com/) ) \--> style (handle by throwing at [ESLint](https://eslint.org/play/)? Asking an LLM about it?) \-> a repo (handle by asking for it) \--> look at the commit history **Made by James*: \-> it has an author (handle by asking questions) \--> purpose \--> tech stack \--> values and risks **CRUD a list*: \-> it should have a list (handle by looking in the code for a plausible list to watch in DevTools), \-> allow each of CRUD (handle with a checklist and watch that list if available) \-> should be list-like (handle by thinking of max and min sizes, sorts, how these work with weird contents, other access to the list) **Keeps data when page is closed:* \-> how? (handle by searching in code, looking for open files, or digging into localStorage for the domain in your browser) \-> the persisting data (handle by identifying location, try changing the data when the page is closed and see results on open) Pick one or two of those artefacts that *you* might want to explore. If nothing appeals, notice that and spend time thinking of others that might. No sense in mucking about with dull things. Write a note about how you'd handle that exploration (remember: you [don't have to change it to explore it](https://www.workroom-productions.com/exploration-without-interaction-2/), and you [don't need to know what it should do](https://www.workroom-productions.com/exploring-without-requirements/)). These are your [*Exploratory Interfaces*](https://www.workroom-productions.com/exploratory-interfaces/)*.* When you're writing, - Know why this thing attracted you for exploring - Consider whether you would bring a specific skill - Consider whether you'd use (or go look for) a tool - Consider how you might iterate – to do similar things to a series of similar parts, and how you'd find the limits of your exploration. If the opportunity presents itself, certainly try handling it in the way you imagine, take the next steps and see how you handle the next. Use the shared space to write / share your note. ### Debrief *(the other half)* We'll ask for volunteers to share their thoughts, and the rest of the group will celebrate the unique approach that they bring. At the end, we will sum up with surprises and changed perspectives. ### Exercise: How do *You* Explore a System? URL: https://www.workroom-productions.com/exercise-how-do-you-explore-what-a-system-does-2/ Last updated: 2025-01-15T22:25:25.000Z *5 mins to build, 10 mins to debrief* **Don't make this general: make it as personal as you can.** ### Build – 5 mins, partly private - Consider: *How do *You* Explore a System?* - Make your own list away from the board, on your own. Fast and short is good. If you immediately think of a long list, prefer variety over completeness or innovation or impressiveness. - Add those items to the board, introducing them to your partners ### Debrief – 10 mins, mostly public - In groups or individually: if you see stuff which goes together, draw and label a circle for that grouping, and put them in - If something you want for a circle has moved to another, duplicate it and put it where you want it. Draw a line between if you like. - Let the groupings and their contents talk to you. If something's missing from a grouping, add it. When activity subsides, we'll gather as a whole and talk. Stuck? I've got a [short and general list of ways](https://www.workroom-productions.com/exploring-without-requirements/#Waystoexplorewithoutrequirements). If you're using one of the BlackBox Puzzles as a thing to think about, you'll find a more-focused list in [Exploring the BlackBox Puzzles](https://www.workroom-productions.com/exploring-the-blackbox-puzzles/). *Variants – think about how you explore what a system* does*, then how you explore what it* is*. Or let the group reveal those differences (and others) for itself.* ### Exploratory Interfaces URL: https://www.workroom-productions.com/exploratory-interfaces/ Last updated: 2025-02-27T17:18:42.000Z *This page has temporary content, and acts as a catch-all while I pull these ideas together. We'll be running exercises around this in* [*PlayTime*](https://www.workroom-productions.com/workroom-playtime/) *– subscribe to join in.* Software systems have plenty of *bits* that can be explored. Those bits invite different ways of exploring. The different ways will reveal different things. We can make models\* and spot trouble by analysing and comparing parts. \* Depending on the limitations of what we make the models with. ## Costs, choices — and opportunities There are\* more ways to explore, and more models to build, than are worthwhile. We need to pick what and how we explore, based on what we might learn (¿get) from the exploration, and on what we need to spend to get it. An example: imagine we have a range of numbers to explore. We chan choose, in black box style, to experiment with inputs. We might put in an input, see what happens, put in another. We might iterate by going up by one, increasing by factors of 10, random pecking or homing in on something interesting. We might connect a slider and wave it about to eyeball the interesting. We might generate input values to a particular distribution, and analyse the output for changes and patterns. We might use values that are already expected to be interesting. We might throw vast combinations of predictable stuff, and watch for the unexpected. Imagine that we choose to explore something else, first. We might choose to explore the code that manages the input. We might look at the data that configures that code and which holds the key to where behaviours are expected to change. We might log everything our users put in over a week. We might look in a database at every output we have managed to store. We might look in the logs to see if there are failures related to the component and its input, or related to its output. We might look in change control to see how often it has changed, why it changed. We might ask experts or users or customers or lawyers about what they expect. While there is far more to choose from than can be done, we can still make a choice to explore. We can still learn about the thing we’re exploring – but OMG the cost! Have a look at [Exploring without Interaction](https://www.workroom-productions.com/exploration-without-interaction-2/). If we're seeking these kinds of opportunities, we're going to be interested in *iterable* exploratory surfaces. \`\* I believe that this is not worth arguing. ## Tooling Tools can make us faster, more accurate and more consistent — or perhaps we might say that neglecting to use the right tool makes the work larger, and introduces more weirdness. Putting a tool 0n an exploration can change the cost of iterating by '000s. Tooling doesn't simply make the exploration cheaper, but can make the exploration scalable. Making an experiment scalable may mean that it is worthwhile using an approach that finds something only rarely. The tool opens a door. Further, ready *access* to tooling can mean that the cost of making a new custom tool is small, so that we can as testers scale\*\* up the opportunities to make use of a tool. If a tool opens a door, and each door opens onto a different perspective of the same system, and each perspective brings otherwise hard-to-find bugs into view, then tool / approach variety is as important as depth. If we're looking for trouble then speed, cheapness and variety of tools can make up for poor quality and fit. Generic or flaky tools will still offer useful clues to issues, and we can choose to invest time in chasing down those problems. For example: A blink test as one slides a value through 1000s of experiments, a sort in a spreadsheet, a swift filter with `jq` can hint to our tester senses. LLMs make it easy to build ephemeral tools from parts – to my mind, this looks to be one of the most fruitful ways that we might use this new technology in our current work. \`\*\* If we're scaling checks in an exploration, then we need to be cautions of false positives. Perhaps not to avoid them, but the be aware that they may flood our results, and to be able to filter them before putting them in front of us to judge. ## Examples The code has text – screen info, error messages, comments. One way to explore those is to extract them. You don't want to slog through yourself – so you could use grep or a regex pattern or hoik everything out of the language / translation config file. When you've extracted the words, why not check them against some sort of oracle by running them through a spellchecker or grammar checker or sentiment analysis tool? The software may be configured – in some situations, the configuration *is* the system. You can look at the configuration, seeing whether you understand what it means, because if you don't then maybe you underestimate what the system is doing. You could use the config as an input to your tests; not only to check the behaviour of the system, but to programatically generate special cases. And while the checks might pass, the system needs your human judgement to analyse what has been produced and to recognise that the configuration is wrong, or that something that appears to be handled with simple configuration actually needs more nuanced treatment. I've explored the config with people who know the goals, regulations or indeed who set the configuration for the business – and found easily-fixable spelling errors, expensive-if-made-real numeric errors, and back-to-the-customisers failures to describe the outside world in a way that met needs. The system produces logs. One way to explore these is to read through them. You want this to be amenable to tools – you could filter for all logs of a particular type or unify several logs, or pop up an alert when some particular entry is logged. You could make a tool to extract all records around a particular time (perhaps one you'd logged automatically earlier) and to use unique IDs so that you can extract all info related to those IDs to see what was in flight at the time you wanted. Have a look at [Ways to Explore without Requirements](https://www.workroom-productions.com/exploring-without-requirements/#Waystoexplorewithoutrequirements) for a longer list of shorter things. ## Classification I tend to distinguish between exploratory surfaces that are of the artefact, and those that are not, though the boundaries are fuzzy. Code and transactions are of the artefact. Operating systems and user reviews are not. I'm not sure about integrated libraries and the history of how the system was built. But all are explorable and can tell us interesting things. I find I want to distinguish between what something *is*, and what it *does*. This page *is* written by me, a ghost blog post, a web page in your browser, an entry in a database. And what it *does?* It communicates an idea, acts as a space for me to keep stuff, signposts readers elsewhere, lets me empty my mind. Perhaps it's simpler to tool against what it is, but simpler to value what it does. A *worthwhile* exploration might need to have both relevant. I find it useful to distinguish between surfaces that invite automation, those that could be automated, and those that don't seem automatable or iterable. I need to have some tools in mind to do this. I do this to understand what might be interesting to automate (click all the links!), not to indicate what should be automated (what does clicking all the links achieve, exactly?). ## Definition ****Exploratory Surface:** an aspect of a system that can be explored. You may be able to iterate over that surface in several different ways. You may be interested in ways to traverse the surface without being able to iterate. ****Exploratory Interface:** a way to interact with, and perhaps iterate over, an exploratory surface. You may be able to use a tool to move over an exploratory interface. Your tool will need a way to iterate (and to notice boundaries and manage obstacles), to make observations, and to aggregate or to filter those observations. *oooh, I'm not a man who much likes definitions – but I do need to tune my thoughts, so they're useful for now.* ## Further reading [Exploring without RequirementsRequirements are helpful rather than crucial![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/icon/wpl-60x60-logo-semitrans-3.png)Workroom ProductionsJames Lyndsay![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/thumbnail/photo-1587922342985-30ff0587b90c)](https://www.workroom-productions.com/exploring-without-requirements/) [Exploration Without InteractionWe were in the Science Museum, me and my 10-year old, looking at the huge Burnley Mill engine as it chugged. We talked about the parts, saw how they were interconnected, wondered about their purpose and function, listened to the noises, watched people as they oiled and tuned, asked questions,![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/icon/wpl-60x60-logo-semitrans-4.png)Workroom ProductionsJames Lyndsay![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/thumbnail/tempImage72pVXG-1.gif)](https://www.workroom-productions.com/exploration-without-interaction-2/) [AI-generated tools can make programming more fun![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/icon/favicon-3.ico)![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/thumbnail/gradient.jpg)](https://www.geoffreylitt.com/2024/12/22/making-programming-more-fun-with-an-ai-generated-debugger) [tools.simonwillison.netAssorted tools![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/icon/favicon-4.ico)tools](https://tools.simonwillison.net) #### Q: Why doesn't this post have a picture? A: Because I can't add a picture, to **this* post. Q: Why's that? A: Because the link doesn't work. Q: Why's that? A: (anguished) I don't know! (more relaxed) I mean, I can add images to **other* posts. And I can see that the link to add an image (looks like an `a`, is actually a `button`, masquerading as a link) has an array of two listeners – just like the other posts. And the things that look weird in the listeners here look just as weird on a working post. Maybe I should spin up a local instance of Ghost and see... Q: Have you tried, I don't know, turning the blog off and on again? A: But if I **fix* it like that, ****how will I know how it failed?** Q: ... just **publish* it, why don't you? ### Exploration Without Interaction URL: https://www.workroom-productions.com/exploration-without-interaction-2/ Last updated: 2025-02-13T12:01:40.000Z We were in the [Science Museum](https://www.sciencemuseum.org.uk), me and my 10-year old, looking at the huge [Burnley Mill](https://www.sciencemuseum.org.uk/see-and-do/energy-hall) engine as it chugged. We talked about the parts, saw how they were interconnected, wondered about their purpose and function, listened to the noises, watched people as they oiled and tuned, asked questions, read the signs, smelt the hot and oily steel. We couldn't touch it, of course. Even the people tending to it avoided touching it, mostly. Touching it is dangerous to the toucher, potentially harmful to the machine. We explored the engine nonetheless. We can explore an artefact without interacting with it. Let's think about a working software system. It's an artefact – a *made* thing – so has history and purpose and utility and risks and versions and problems and (right / wrong) context. It has users and manuals and advertising copy. It lives in a conceptual framework of ideas and expectations, and a technical framework of data and tools and communications, and (right now) is at one moment in probably several lifecycles. It's made of parts, that interact over time and distance within various mediums.It has large chunks that perhaps lie dusty or unused – at this time, and for this system. As software, it's likely made of data and code, so has defaults and error messages and comments. It's built with languages with idioms and patterns, it imports libraries and exports parts of itself for use, it lives within machines that contain other systems that support and compete. It shows off its ways of use via UIs and APIs and signatures and records. It keeps its information in logs and records and databases, accepts configuration. If it is resilient, it has backups, alternatives and workarounds. It may offer itself in multiple (human) languages, or with several feature sets, or in different media. It may make connections to systems you've never thought of, half a world away. And we may be able to see or infer how it is affected by its environment; by the distances and technologies that wrap around its parts, by the data it is given, by the reliability of what it depends on. These things are observable, to us. Each is an [explorable interface](https://www.workroom-productions.com/exploratory-interfaces/), most will reward several different approaches to exploration. Each lets us discover and learn about a different perspective. And – if the purpose of our exploration is not only to learn but to look for flaws and surprises – inconsistencies between the models that they engender in our minds may tell us that the system is problematic, that the perspective is inaccurate, or that my model was rubbish. We're exploring – but perhaps we're not *testing the machine* as much as we are testing ourselves. We can't change the machine, certainly. We can change how we see it, think of it and describe it. I saw a coal-fired machine inside a crowded public building, and understood that the fire box was for show only when I recognised that I could not smell burning coal. I read that the machine had been built in 1903 and did not, ahem, have a beam – so could correct the fact ("it's a Victorian beam engine") that I'd confidently spewed. We asked about a part, and discovered that the intimate joints of the whirling, clanking and unstoppably massive limbs of these machines need to be oiled as they work – and that the delicate and mysterious components are vital to the survival of both the machine, and its operators. There's bugs in them hills. And there's more to finding them than shoving long strings into short input fields. Exercise: try [Exploring Fixed Input](https://www.workroom-productions.com/exercise-exploring-fixed-input/). --- Subscribers get to see the stuff I wrote that didn't fit here... ## Sprue - It's interesting to think about how we explore, when we can observe, but can't interact directly. Pretty much all exploration outside our planet is done with observation. - Some observation is by inference only – observed data, accepted maths and tested models inform us about the age and size of the universe, the atomic interactions in supernovae, the behaviour of black holes. - We might argue that [it's all inferred](https://en.wikipedia.org/wiki/Direct%5Fand%5Findirect%5Frealism) – we might also argue that any [observation necessarily changes the object being observed](https://en.wikipedia.org/wiki/Observer%5Feffect). These are nice distractions, but don't currently have a place here. ### PlayTime 002 – Exploratory Interfaces I URL: https://www.workroom-productions.com/playtime-002-exploratory-interfaces-i/ Last updated: 2025-01-16T12:21:05.000Z We'll gather on Zoom / Miro, on Friday 17 January at 3pm London time ([local time for you](https://this-ti.me/?uts=1737126000&tz=Europe%2FLondon&name=Workroom+PlayTime)). Here are some thoughts about [Exploratory Interfaces](https://www.workroom-productions.com/exploratory-interfaces/). We'll play with [Changing Input Mechanisms](https://www.workroom-productions.com/changing-input-mechanisms/). These exercises are for everyone, for free. Paying subscribers get to play together. _This post is for paying subscribers only._ ### PlayTime 001 – Diff URL: https://www.workroom-productions.com/playtime-001-diff/ Last updated: 2025-01-09T16:40:56.000Z We'll gather on Zoom / Miro, on Friday 10 January at 1pm London time, to play with [diff](https://www.workroom-productions.com/diff-for-testers/). We'll do some [Diff Tool Exercises](https://www.workroom-productions.com/diff-tool-exercises/) together. These exercises are for everyone, for free. Paid subscribers get to play together. _This post is for paying subscribers only._ ### Diff Tool Exercises URL: https://www.workroom-productions.com/diff-tool-exercises/ Last updated: 2025-02-05T19:22:49.000Z `diff` compares two files, highlighting differences. Here's [how diff can work for testers](https://www.workroom-productions.com/diff-for-testers/). On that page, you'll find links to several tools that implement and build on `diff` – you may already have your own preference. In this exercise, you’ll use any diff tool that seems right to you. This exercise is more fun in a group. If you’re in a group, please talk about what tool you’d use. Perhaps re-run the exercise with someone else’s tool of choice, or with someone else. The solutions are below. The tool will do most of the work – give yourself kudos for spotting non-trivial things. Zero kudos if you look at the solutions before trying the exercises. ## Exercises Group 1 – XML and JSON Download the linked files, use them in a tool. If you’re opening the file in a browser, finding out what an xmlschema is, and reading through it, you’re doing it …differently from how I’d expect you to do it. ### Exercise 1.01 (XML) [file 1](http://exercises.workroomprds.com/DiffToolExercise/DiffToolExercise%5F1.1/pain.008.001.01v1.xsd), [file 2](http://exercises.workroomprds.com/DiffToolExercise/DiffToolExercise%5F1.1/pain.008.001.01v2.xsd) #### Solution #### Exercise 1.01 - Lines Removed in the schema definition, v2 is missing these lines, location 453 in original ``` ``` ### Exercise 1.02 (XML) [file 1](http://exercises.workroomprds.com/DiffToolExercise/DiffToolExercise%5F1.2/pain.008.001.01v1.xsd), [file 2](http://exercises.workroomprds.com/DiffToolExercise/DiffToolExercise%5F1.2/pain.008.001.01v2.xsd) #### Solution #### Exercise 1.02 - Lines Added In the schema definition, v2 has had line 699 duplicated in error ``` ``` ### Exercise 1.03 (XML) [file 1](http://exercises.workroomprds.com/DiffToolExercise/DiffToolExercise%5F1.3/pain.008.001.01v1.xsd), [file 2](http://exercises.workroomprds.com/DiffToolExercise/DiffToolExercise%5F1.3/pain.008.001.01v2.xsd) #### Solution #### Exercise 1.03 - Lines Rearranged In the schema definintion on lines 224 to 234 of v2, SimpleType DocumentType2Code has been rearranged. Kudos if you recognised that it has been sorted ``` ``` ### Exercise 1.04 (XML) [file 1](http://exercises.workroomprds.com/DiffToolExercise/DiffToolExercise%5F1.4/pain.008.001.01v1.xsd), [file 2](http://exercises.workroomprds.com/DiffToolExercise/DiffToolExercise%5F1.4/pain.008.001.01v2.xsd) #### Solution #### Exercise 1.04 Information Changed On A Single Line In the schema definition on line 463, the order of attributes has been rearranged. Kudos if you noticed that `maxOccurs` has changed from 4 to 5 ``` ``` ### Exercise 1.05 (XML) [file 1](http://exercises.workroomprds.com/DiffToolExercise/DiffToolExercise%5F1.5/pain.008.001.01v1.xsd), [file 2](http://exercises.workroomprds.com/DiffToolExercise/DiffToolExercise%5F1.5/pain.008.001.01v2.xsd) #### Solution #### Exercise 1.05 - Lines Split In the schema definition on line 165, the line has been split by tag. Kudos if you noticed that there’s an extra attribute `base`, that the tag is no longer closed with `/>` but ends with a `>`, and that the `name` attribute is now `ccy` rather than `Ccy` ``` ``` ### Exercise 1.06 (XML) [file 1](http://exercises.workroomprds.com/DiffToolExercise/DiffToolExercise%5F1.6/pain.008.001.01v1.xsd), [file 2](http://exercises.workroomprds.com/DiffToolExercise/DiffToolExercise%5F1.6/pain.008.001.01v2.xsd) #### Solution #### Exercise 1.06 - Lines Joined In the schema definition, line 278 has been joined with line 279\. Kudos if you noticed the spelling of `elemnet`, the absence of capitalisation in both pairs of `maxoccurs` and `minoccurs`, and at least one spellign error. ``` ``` ### Exercise 1.07 (JSON) [file 1](https://exercises.workroomprds.com/DiffToolExercise/DiffToolMoreEx1/transaction%5F1%5Fa.json), [file 2](https://exercises.workroomprds.com/DiffToolExercise/DiffToolMoreEx1/transaction%5F1%5Fb.json) #### Solution #### Exercise 1.07 - data changed On line 632, the value has an extra decimal place. Kudos if you noticed the small change in property name `entryCustom4OCode`. ```json "entryCustom4OCode": "20.000000000", ``` On 892, the `cardTransactionReferenceNumber` is `""` rather than `null`. ```json "cardTransactionReferenceNumber": "", ``` ### Exercise 1.08 (JSON) [file 1](https://exercises.workroomprds.com/DiffToolExercise/DiffToolMoreEx1/transaction%5F2%5Fa.json), [file 2](https://exercises.workroomprds.com/DiffToolExercise/DiffToolMoreEx1/transaction%5F2%5Fb.json) #### Solution #### Exercise 1.08 - line added Line 346 in the second file does not exist in the first file. ```json "nextPage": null ``` Kudos if you noticed that it is missing the trailing comma, making the JSON invalid. ### Exercise 1.09 (JSON) [file 1](https://exercises.workroomprds.com/DiffToolExercise/DiffToolMoreEx1/transaction%5F3%5Fa.json), [file 2](https://exercises.workroomprds.com/DiffToolExercise/DiffToolMoreEx1/transaction%5F3%5Fb.json) #### Solution #### Exercise 1.09 - lines missing Line 46-48 , 67-69, 126-128 and 298-300 in the first are missing in the second. ```json "employeeCustom19Code": null, "employeeCustom20Code": null, "employeeCustom21Code": null, "employeeCustom19Value": null, "employeeCustom20Value": null, "employeeCustom21Value": null "reportCustom18Code": null, "reportCustom19Code": null, "reportCustom20Code": null, "allocationCustom18Code": null, "allocationCustom19Code": null, "allocationCustom20Code": null, ``` Kudos if you noticed that line 66 correctly omits the trailing comma, so the JSON is valid although the data is missing. Kudos if you wondered why `entryCustomCode` has been left untouched. ### Exercise 1.10 (JSON) [file 1](https://exercises.workroomprds.com/DiffToolExercise/DiffToolMoreEx1/transaction%5F4%5Fa.json), [file 2](https://exercises.workroomprds.com/DiffToolExercise/DiffToolMoreEx1/transaction%5F4%5Fb.json), [file 2b](https://exercises.workroomprds.com/DiffToolExercise/DiffToolMoreEx1/transaction%5F4%5Fb2.json) Why have I given you two file 2s? Can you see any differences? #### Solution #### Exercise 1.10 - values changed, alternative formats The second file has a new blank line at line 71 The second document status is marked "Pending" rather than "READY" on line 358 ```json "docStatus": "Pending" ``` Kudos if you wondered why `Pending` is not in upper case. Kudos if you spotted that while file1 and file 2 have identical *contents*, they are rather different *sizes*. More kudos if you noticed their *encoding* is different. ```shell -rw-r--r--@ 1 james staff 14617 10 Jan 11:14 transaction_4_a.json -rw-r--r--@ 1 james staff 14628 10 Jan 11:25 transaction_4_b.json -rw-r--r--@ 1 james staff 29254 10 Jan 11:25 transaction_4_b2.json james@Mac diff tool exercise % file -I transaction_4_b.json transaction_4_b.json: text/plain; charset=utf-8 james@Mac diff tool exercise % file -I transaction_4_b2.json transaction_4_b2.json: text/plain; charset=utf-16be ``` Your tool may not reveal this to you... [How various git diff viewers represent file encoding changes in pull requests - The Old New ThingThe invisible UTF-8 BOM, and sometimes invisible encoding changes.![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/icon/Microsoft-Favicon.png)The Old New ThingRaymond Chen![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/thumbnail/ShowCover.jpg)](https://devblogs.microsoft.com/oldnewthing/20241230-00/?p=110692) «JL: Add more, in CSV and TXT (i.e. dir list). » Planned exercises include... ## Exercises Group 2 – Code and config Python, JavaScript, YAML/TOML, env ## Exercises Group 3 – Markdown Docs and readme. ## Exercises Group 3 – Folders Changed content, added / removed, renamed / moved, timestamp, permissions, ?encoding, ?os-specific metadata, ?hash, ## Exercises Group 4 – Combos sort / grep / filter – first, within git and across files for a commit, on directory / process lists, getting unique / count of changes, ## --- *This page is a work in progress. I'm sure that a curious reader can easily find my notes to myself. «JL: perhaps, if the script finds any notes, it can insert this message?»* ### `diff` for Testers URL: https://www.workroom-productions.com/diff-for-testers/ Last updated: 2025-02-05T19:24:54.000Z `diff` helps you to compare two similar texts. It's great for seeing what's changed in code or in data. It will compare two directories, recursively if you ask. It’s best on text things, and not a lot of use for binaries (though if you know tricks, teach me). Some diff tools will compare specific kinds of binary files; word and excel docs, jpgs and more – these are useful, but tend to be proprietary and I won't cover them here. You’ll use a diff tool on any text thing longer than a paragraph. Even a paragraph, sometimes. You’ll use it when you have two versions - or when you *suspect* you have two versions rather than twins. You’ll find differences due to direct action by a person, by generators, by tools, by time, and by corruption. And more. You already know diff. You’ve seen it in GitHub and other collaborative coding environments. You’ve seen tools like it in Word and other collaborative writing environments. Diff tools tell you *that* a line has changed. Many don’t tell you *how* a line has changed. None tell you if the information still makes sense, or has stayed consistent. `diff` is a unix tool: a mandatory part of the available parts that let an OS describe itself as unix-like – it's been remade several times, and there are several tools that build on `diff` or are `diff`\-like. See below. ## Use it for - Identifying code changes, to look for unexpected changes - Checking differences in database dumps - Looking for changes to environmental variables - Seeing how two output records differ - Seeing what changed in your input data when the system will no longer ingest it - Looking for changes to configuration - Looking to see which files in a directory have changed - Checking that your commit is committing only what you want to commit - Seeing what's stayed the same in the monthly report to identify where an update hasn't happened - Looking for textual changes to requirements - Browsing two versions of a contract ...and more. If you don't have it in your toolkit, you're blinding yourself. ## Sources The `man` page is Opengroup page Wikipedia page ## Tool Examples Basic `diff` tools are very command-line, and can be used easily in chains of other tools. If you’re going command-line, you’re old enough to hold your own hand. Some are more visually friendly, and can be used by a broader group of people. - Here are a couple of web-based tools. [DiffChecker](https://www.diffchecker.com/), [Diff-text](http://www.diff-text.com). - Here’s a (commercial) list of [common dif tools for Windows](https://www.git-tower.com/blog/diff-tools-windows) - and [more Diff Tools on Mac](https://www.git-tower.com/blog/diff-tools-mac/). - Windows machines typically have [kDiff3](http://kdiff3.sourceforge.net/) and I believe that [meld](https://meldmerge.org/) is often available. My Mac has[DeltaWalker](https://www.deltawalker.com/) and [FileMerge](https://apple.stackexchange.com/questions/42345/where-can-i-download-filemerge-the-app-for-comparing-two-tools-and-merging-the). Dev teams often seem to have a license for cross-platform [Beyond Compare](https://scootersoftware.com/). ### Unix Tools for Testers URL: https://www.workroom-productions.com/unix-tools-for-testers/ Last updated: 2025-11-14T18:30:53.000Z A series on how I use these as a tester – probably to be used in [Workroom PlayTime](https://www.workroom-productions.com/workroom-playtime/), and to support workshops on toolsmithery. ## Published info and exercises - `diff` – [info](https://www.workroom-productions.com/diff-for-testers/) – [exercises](https://www.workroom-productions.com/diff-tool-exercises/) - `grep ` – [info](https://www.workroom-productions.com/grep-for-testers/) – [exercises](https://www.workroom-productions.com/grep-tool-exercises/) - `sort` – [exercises](https://www.workroom-productions.com/sort-exercises/) - `cat`, `head`, `tail` – [info](https://www.workroom-productions.com/cat-head-and-tail-for-testers/) – [exercises](https://www.workroom-productions.com/exercises-for-cat-head-and-tail/) - `curl` – [exercise](https://www.workroom-productions.com/curl-exercises-for-testers/) - `xargs` – [](https://www.workroom-productions.com/xargs-for-testers/)[exercise](https://www.workroom-productions.com/workroom-playtime-037-xargs-tool-for-testers/) - `jq` – [exercise](https://www.workroom-productions.com/exercise-jq-for-testers/) - Redirection operators `|`, `>` `>>`, `<` (and `tee`) – [info](https://www.workroom-productions.com/redirection-operators-for-testers/), [exercise](https://www.workroom-productions.com/exercise-redirection-operators/) - `sed` - *regex* It's good to know how these simple tools can be used together. #### Example 1 - It may be that you've got a log from a test system and you want to extract all lines that refer to a particular transaction that seemed problematic. You'd `grep`. - Maybe you've got several logs, and you need to look for every reference. You could merge them while paying attention to the timestamps with a `sort`, then `grep`, and you'd join those two with a pipe `|`. - Maybe those logs are remote, but you can get them with `cURL`. You `|` them into `sort`. Maybe you want to add a suffix to each line, so that you can see the source. You `curl` | `sed` | `sort` | `grep` - Maybe you can't read the pages of scrolling weirdness. You `|` the `grep` into `tail`, to see the last 10 lines. - At the end, you've got a single hard-to-read but repeatable command `curl` | `sed` | `sort` | `grep` | `tail` – you use it once and throw it away. Or keep it in a handy spot. #### Example 2 You've got a huge json record, containing some slice of a data state of a whole system – you want summaries for some test accounts, but it's got sooo much more. The json's invalid: it's got bad escapes in several places, but always in the same way and not anywhere you're interested in. You use `cat` to send the contents to `sed` and knock out the offending bits with a regular expression, send the now-valid json to `jq` and pick out the repeating parts you need. You want parts of all of them in a table – you tell `jq` to pick out what you need, transform the dates to human-readable, and output as a `csv` . Because this is a command, you make it part of the start and end of every big test run and drop the manageable files into... I dunno, Excel? so that you can see the context of each account as you explore. Lots of utilities do text processing of some sort – filtering, sorting, aggregating. To get that text, other utilities pull information from files, the file system or processes. Information is passed from one utility to another with a handful of operators, and the end result is typically dumped into an editor or a file for someone to play with – or for another process to pick up. As a tester, I've used these tools to check changes to my environment, to pick out specific data, to gather context for my exploration, to watch what's going on as I work, to filter the firehose. Testers don't (typically) write stuff for permanent deployment: their [tools are ephemeral](https://www.workroom-productions.com/ephemeral-tooling/). Those tools can still be shared, reused and tuned. However, capabilities built with these general tools are often so specific – and so hard to interpret – that they don't get shared, and get remade rather than reused. By knowing what's possible, we can build things that are practical, and I hope that we will share and learn from each other. It's a great time for testers to get better at using tiny tools: The rise of StackOverflow and blogging has led directly to a greater exchange of tricks, and LLMs, trained on these freely-available sources, are extraordinary at helping me to swiftly get to the tooling I need. I can ask, iterate and learn about what can be done, and options around how to do it, without getting too stuck on syntax and cluelessness. Some sources: - The key permanent / formal collection that these are drawn from is probably the [IEEE / Open Group's 'shell and utilities' page](https://pubs.opengroup.org/onlinepubs/9799919799/), and it's a beast. - Most of these are found in GNU [core utilities](https://www.gnu.org/software/coreutils/coreutils.html) (here's [*WP's page on core utilities*](https://en.wikipedia.org/wiki/List%5Fof%5FGNU%5FCore%5FUtilities%5Fcommands)), and there's a substantial crossover with [Portable Operating System Interface](https://en.wikipedia.org/wiki/List%5Fof%5FPOSIX%5Fcommands) Commands (so *that's* what `POSIX` means...). - Here's a [GNU's page of coreutils's faqs](https://www.gnu.org/software/coreutils/faq/coreutils-faq.html), which outlines some weirdnesses, edge cases, and unexpected expected behaviours. And here's a pair which are interesting not only because of their content but taken a pair because they're written by the same person (GNU coreutils maintainer [Pádraig Brady](https://www.pixelbeat.org)): [How the GNU coreutils are tested](https://www.pixelbeat.org/docs/coreutils-testing.html), [Coreutils Gotchas](https://www.pixelbeat.org/docs/coreutils-gotchas.html). _This post is for subscribers on the Owner tier only._ ### Claude's BlackBox Puzzle URL: https://www.workroom-productions.com/claudes-blackbox-puzzle/ Last updated: 2025-01-03T17:11:48.000Z I asked Claude to do a thing. Let's play with it: [Claude ArtifactTry out Artifacts created by Claude users![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/icon/favicon.ico)![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/thumbnail/claude_ogimage.png)](https://claude.site/artifacts/40e6a8df-be27-49d0-be6e-b5ba57896b8d) This is the first session of [Workroom PlayTime](https://www.workroom-productions.com/workroom-playtime/), a weekly hands-on workshop for paying subscribers. Subscribe to get access to the exercise and the group. _This post is for paying subscribers only._ ### Workroom PlayTime URL: https://www.workroom-productions.com/workroom-playtime/ Last updated: 2026-08-19T20:57:15.000Z *Regular, short, online interactive sessions for subscribers and friends.* Come to Workroom PlayTime to play with testing for 20 minutes.We'll test and talk: No lecture, no sales pitch. We‘ll typically run on Thursdays, in Zoom / Miro, in the afternoons. I’ll keep the group sizes small so that we can talk. We will keep sessions short – 20 minutes — though the post-workshop chat may last longer. I’ll run sessions several times if I need to, depending on busy-ness and timezones. Subscribers get an hour a month of testing-related playtime. Paying subscribers get priority seats, and access to the longer sessions where we re-run exercises and go deeper. [Subscribe](https://www.workroom-productions.com/#/portal/). ## 2026 – Themed sets Here's [more](https://www.workroom-productions.com/workroom-playtime-in-2026/) about my plans – expect plan and reality to diverge... ### August – BlackBox Puzzles - `062` Thur 30 July / Sat 1 August: Explore [Puzzle 39](https://www.workroom-productions.com/workroom-playtime-062-explore-puzzle-039/) - `063` Friday 14 July: [Spot The Difference (Puzzle 40/40b)](https://www.workroom-productions.com/workroom-playtime-063-spot-the-difference/) ### January: Interesting Testing I - `044` 15 January: [A Software Library with no Code](https://www.workroom-productions.com/workroom-playtime-044-2/) - `045` 22 January: [The ARC prize](https://www.workroom-productions.com/workroom-playtime-045-arc-foundation-tests-2/) - `046` 29 January: [The Ralph Loop](https://www.workroom-productions.com/workroom-playtime-046-the-ralph-loop/) I'll re-run all these over a longer session in March – see [Longplayer: Interesting Testing](https://www.workroom-productions.com/workroom-playtime-longplayer-interesting-testing/). Here's an [Interesting Testing](https://www.workroom-productions.com/interesting-testing/) theme page as the series grows. ### February: Speaker Prep - `047` 5 February: [Getting to a Great Abstract](https://www.workroom-productions.com/workroom-playtime-047-getting-to-a-great-abstract/) - `048` 19 February: [Cutting to the Core](https://www.workroom-productions.com/workroom-playtime-048-cutting-to-the-core/) - `049` 26 February: [What Makes a Great Talk?](https://www.workroom-productions.com/workroom-playtime-049-what-makes-a-great-talk/) I'll run these together for a longer [Speaker Prep](https://www.workroom-productions.com/tag/speakerprep/) session in late-March. ### March: Managing Exploration - `050` 5 March: [Power of Variety](https://www.workroom-productions.com/workroom-playtime-050-power-of-variety/) ... and that was it for March. Stuff gets in the way. I'll get back to the series later in the year. ### April: - `051` 2 April: [Exploring Bloom Filters](https://www.workroom-productions.com/exploring-bloom-filters/) - `052` 9 April: [Exploring Play Styles](https://www.workroom-productions.com/exploring-play-styles/) - `053` 23 April: [Play Styles and Exploratory Work](https://www.workroom-productions.com/play-styles-and-exploratory-work/) - `054` 30 April: [Getting Stuck For Explorers](https://www.workroom-productions.com/workroom-playtime-054-getting-stuck-for-explorers/) ### May - `055` 7 May: [Iterating for Exploratory Testers](https://www.workroom-productions.com/workroom-playtime-055-iterating-for-exploratory-testers/) - `056` 14 May: [Switching for Explorers](https://www.workroom-productions.com/workroom-playtime-056-switching-for-explorers/) ### June - `057` 11 June: [Disaster Story](https://www.workroom-productions.com/workroom-playtime-057/) - `058` 18 June: [Ever-Rolling Stream](https://www.workroom-productions.com/exercises-about-testing-and-systems-analysis/#ever-rolling-stream) from [Exercises about Testing and Systems Analysis](https://www.workroom-productions.com/exercises-about-testing-and-systems-analysis/). - `059` 25 June: [Playing with Tenfold](https://www.workroom-productions.com/workroom-playtime-059-play-with-tenfold/) ### July - `060` 2 July: [Sitegeist as an Exploratory Interface](https://www.workroom-productions.com/sitegeist-as-an-exploratory-interface/) --- ## 2025 - `000` 3 January: [*Claude made a BlackBox Puzzle*](https://www.workroom-productions.com/claudes-blackbox-puzzle/)*.* Off we go! - `001` 10 January: A [short *diff exercise*](https://www.workroom-productions.com/playtime-001-diff/) from the archives. - `002` 17 January: [*Exploratory Interfaces* I](https://www.workroom-productions.com/playtime-002-exploratory-interfaces-i/) - `003` 24 January: [*SpeakerPrep* 1](https://www.workroom-productions.com/playtime-003-speakerprep-following-up/) – What to Follow Up After a Talk. - `004` 31 January: [Exploring *BlackBox Puzzle 031*](https://www.workroom-productions.com/playtime-004-blackbox-puzzle031/)! - `005` 7 February: [Working with grep](https://www.workroom-productions.com/playtime-005-grep/) - `006` 14 February: [*Exploratory Interfaces II* – exploring fixed input](https://www.workroom-productions.com/playtime-006-explorable-interfaces-ii/) - `007` 21 February: [SpeakerPrep II – S*loganise*](https://www.workroom-productions.com/playtime-007-sloganise/) - `008` 28 February: [*Imagining Questions*](https://www.workroom-productions.com/playtime-008-2/) - `009` 7 March: [Notation and models (for *Puzzle 15*)](https://www.workroom-productions.com/playtime-009/) - `010` 27 March: [Testing Decisions – Power of Variety](https://www.workroom-productions.com/playtime-010-2/) - `011` 3 April: [Using sort](https://www.workroom-productions.com/playtime-011-playing-with-sort/) - `012` 17 April: [Testing Decisions – Cost of Trouble](https://www.workroom-productions.com/playtime-012-testing-decisions-cost-of-trouble/) - `013` 24 April: [Three Things](https://www.workroom-productions.com/playtime-013-three-things/) - `014` 1 May: [Puzzle 33](https://www.workroom-productions.com/playtime-014-puzzle-33/) - `015` 8 May: [Bring me a Letter](https://www.workroom-productions.com/playtime-015-bring-me-a-letter/) - `016` 15 May: More playful: [Whatever Next?](https://www.workroom-productions.com/playtime-016-whatever-next/) - `017` 22 May: [Whatever Next? II](https://www.workroom-productions.com/playtime-017-whatever-next-2/) - `018` 29 May: [RasterReveal](https://www.workroom-productions.com/playtime-018-raster-reveal/) - `020` 5 June: [Live at EuroSTAR / Code is Cheap](https://www.workroom-productions.com/playtime-020-code-is-cheap-tests-are-valuable-at-eurostar25/) - `019` 12 June: Let's explore [Puzzle 36](https://www.workroom-productions.com/puzzle-036/) - `021` 19 June: [Exercises in cat, head and tail](https://www.workroom-productions.com/exercises-for-cat-head-and-tail/) - `022` 26 June: [Testing Story](https://www.workroom-productions.com/playful-exercises-about-testing/) - `023` 03 July: [Comparing LLMs](https://www.workroom-productions.com/workroom-playtime-023/) to explore judgement - `024` 17 July: [Puzzle 11 (is an awful toy)](https://www.workroom-productions.com/workroom-playtime-024-puzzle-11/) - `025` 24 July: [curl exercises for testers](https://www.workroom-productions.com/workroom-playtime-025-curl/) - `026` 14 August: [Shared Inspiration](https://www.workroom-productions.com/playful-exercises-about-testing/#shared-inspiration) - `027` 21 August: [Being Random](https://www.workroom-productions.com/workroom-playtime-027/) - `028` 28 August: [My Story Chart](https://www.workroom-productions.com/workroom-playtime-028-my-story-chart/) - `029` 4 September: [Imagined | Real](https://www.workroom-productions.com/imagined-real/) - `030` 11 September: [Completeness (is a wicked problem)](https://www.workroom-productions.com/completeness-is-a-wicked-problem/) - `031` 18 September: [Puzzle 17](https://blackboxpuzzles.workroomprds.com/newtech/puzzle17.html) - `032` 2 October: [Redirection operators ( | etc.) for testers](https://www.workroom-productions.com/workroom-playtime-032/) - `033` 9 October: [Exploratory Interfaces II](https://www.workroom-productions.com/workroom-playtime-033/) - `034` 16 October: [jq for testers](https://www.workroom-productions.com/workroom-playtime-034-jq-for-testers/) - `035` 23 October: [build 3 throwaway tools in 15 minutes](https://www.workroom-productions.com/workroom-playtime-035-build-3-throwaway-tools-in-15-minutes/) - `036` 30 October: [custom asserts](https://www.workroom-productions.com/workroom-playtime-036-custom-asserts/) - `037` 6 November: [xargs for testers](https://www.workroom-productions.com/workroom-playtime-037-xargs-tool-for-testers/) - `038` 13 November: [parts for tools](https://www.workroom-productions.com/workroom-playtime-038-parts-for-tools/) - `039` 20 November: [Stuck Overflow](https://www.workroom-productions.com/workroom-playtime-039-stuck-overflow/) - `040` 27 November: [Explore Generated Code II](https://www.workroom-productions.com/workroom-playtime-040-explore-generated-code-ii/) - `041` 4 December: [Puzzle 38](https://www.workroom-productions.com/puzzle-38/) - `042` 11 December: [Exploratory Interfaces 4: Tests](https://www.workroom-productions.com/exploratory-interfaces-4-tests/) - `043` 18 December: [Building Short Exercises](https://www.workroom-productions.com/workroom-playtime-043-building-short-exercises/) #### In the future... Here are some [exercises that are on their way ](https://www.workroom-productions.com/imminent-exercises/) - Testing Decisions – **Cost and Opportunity* or **Non-Linear Consequences / When to Stop*. - Building exploratory interfaces (longer session) - Exploring Generated Code I ### Online AI vs TDD workshop URL: https://www.workroom-productions.com/ai-vs-tdd-online-20241210/ Last updated: 2024-12-09T14:30:50.000Z I'll run [*Guiding Hands-off AI using Hands-on TDD*](https://www.workroom-productions.com/guiding-ai-code-with-tests-workshop/) on 10 December 1:30- to 3pm UK time, free for site [subscribers](https://www.workroom-productions.com/#/portal/subscribe) and their friends. In the workshop you'll write tests, then ask an LLM to iterate code until it passes the tests. It's all in-browser, you don't need to download anything or set up an IDE. You don't need to code but you'll need to read code, and be open to typing on a command-line. You'll need a free [Replit](https://replit.com/~) account to access the IDE. Bart Knaack and I [ran this interactive workshop](https://agiletestingdays.com/2024/session/guiding-hands-off-ai-using-hands-on-tdd/) first time at Agile Testing Days last week. Participants had a great time: [Dragan Spiridonov wrote](https://www.linkedin.com/pulse/back-from-agile-testing-days-2024-reflection-learning-spiridonov-oxo5f/?trackingId=cZUFTqvnSrOSm7iHAGIj9w==) that he found it *fascinating* and *mind blowing*, and [Christoph Zabinski was thrilled](https://www.linkedin.com/feed/update/urn:li:activity:7265794898069491712/) that our session and others *didn’t just ride the AI hype train—they provided thoughtful, grounded perspectives on how AI can truly enhance our work as testers*. We were delighted to have such an engaged group – ATD participants are always very special. You should [go in 2025](https://agiletestingdays.com/register/): they've got a massive discount before the New Year. Anyway: If you want to come, [*email me*](mailto:jdl@workroom-productions.com?subject=Online%20AI%20vs%20TDD%20workshop%20on%2010%20Dec%201%3A30-3%3A30pm%20London%20time). Or [subscribe](https://www.workroom-productions.com/#/portal/account/plans)... If you want to dig deeper, I'm about to offer it to teams as 3x1-hour online workshops or a half-day session. [Contact me](mailto:jdl@workroom-productions.com?subject=Tell%20me%20about%20your%20AI%20vs%20TDD%20workshop%20as%20online%20training%20for%20my%20team) to be in the first wave of bookings... --- PS – here's an article, published last week, on how I used LLMs to learn about stuff over this last year. [Learning with an LLMAn LLM is a powerful, flawed and delightful study tool.![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/icon/wpl-60x60-logo-semitrans.png)Workroom ProductionsJames Lyndsay![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/thumbnail/photo-1487001175664-86de872e3cd6.jpeg)](https://www.workroom-productions.com/learning-with-an-llm/) ### Getting to a Great Abstract URL: https://www.workroom-productions.com/getting-to-a-great-abstract/ Last updated: 2025-01-23T10:57:26.000Z *A 30 minute collaborative exercise ending up in a worthwhile short abstract for your idea.* *An abstract, here, is short text suitable for submitting to an event. It's longer than a few sentences, shorter than an article – say 300 words if you like, or 2-3 mins reading – and should help the reader judge whether they would put your idea in front of their audience at their event. It's not about you (unless your idea for a talk is about you).* ### Process - Have a chat about your idea (you need one) and existing abstract (if you have one) - Prime your mind with a couple of checklists *or our gallery of fantasy abstracts (to follow)*, or by searching for help with abstracts or (careful) searching for other abstracts in that area. - Write something down! - Ask for review, comments, advice - (optional) Add your abstract to the gallery and ask for public comments - Set milestones for next steps *Facilitators: be part of chats, if invited. Have things to add to the checklists. Be ready to advise if asked, or to show one of your own abstracts / tell a story of your own. Set up (if possible) accountability partners for the milestones.* --- ### Great abstract - catchy title - idea that has potential - indicates that thought and work has gone into communicating the message - /brief/ problem statement - real examples - demonstrates experience - clear takeaways - relevant to event - easy to read, draws you onwards - non-trivial - has substance – more than a tease, less than a paper - more than just the bright side – pitfalls, failures, antipatterns, pathologies - within word limit (and more than a couple of sentences) - offers interaction or demonstration --- ### Poor abstract - claims expertise / authority - problem excludes content - poorly written, poorly proofread - unstructured or incoherent - buzzwords - single tool - happy paths only - One true method / "I'm right" - Emotionless - Tired topic - Typical AI tropes ### Dealing with Disaster URL: https://www.workroom-productions.com/dealing-with-disaster/ Last updated: 2025-01-23T10:59:35.000Z A two-part, 30-45 minute session to build thoughtfulness and resilience around common presenting problems. Be aware that participants may find it worrying to consider what could go wrong – I reckon that this exercise is best if it *reduces* anxiety. To work well for me, this exercise needs to be particularly playful and full of laughter – and I need to remember to get everyone's buy-in before starting this topic. *You'll see from the examples that this was built to be used in the SpeakerPrep workshop. I've used similar exercises at clients to explore unthinkable futures on projects and for products in use: the very different stakes in those circumstances mean that I need to take care not to enable the group to manage corporate anxiety by trivialising distant impacts.* ## What disasters? Workshop in groups – set group size and tasks to allow 5 mins work, 5 mins debrief. Q1: What disasters might happen (3 mins) Q2: Classify those – aim for a small handful of classifications (2 mins) Debrief: Each group to tell the room 1) your classifications 2) a couple of examples per Facilitators to: - give visibility to all (and support memory) by writing up - make disaster safe by ¿delighting in the failure? ¿peer work? ¿owning failure? - Fill in the gaps: Give the room the chance to offer situations that don’t fit any existing classification. Facilitators should offer situations whose broad family haven’t been considered (i.e. if the room’s only got tech issues). See examples below if uninspired. - Quietly assess classifications – why are these classifications in use here? Are the classifications focussed on source, situation, response – and what can be learned if not? ## How to deal with it? Prime the room to consider: - response vs reaction - who’s frustrated by by the problem, who has the power to fix it, who feels ownership / will be embarrassed by it - resources that can address it: organisers, allies, younger you. Form groups again. Different ones? Pick things to work on, based on classifications / examples. 5 mins to workshop responses / ideas 5 mins to prep a 1-minute summary Deliver summaries – facilitators to write up a flipchart sheet for each Open floor to conversations, add to flipcharts. Facilitators to post on walls / add to this on GitHub. --- ## Example classifications and some approaches to dealing with it #### **Sources: Things that the presenter is bothered by, can fix – ie technology failures** - Have you got an alternative, and are you prepared to switch? - Can your message do without the thing that’s gone? Do you need to restructure? Tell rather than show? - Can you truncate the talk? - Can you put the failed thing later? - Do you have an ally in the audience who can work to fix while you continue? - How can you prepare for common failures, or for high-impact failures? - Can you choose more-resilient tech? - Can you deliver your message without tech? #### **Situations: Things you can’t control or prepare for / Situational problems** - Stay calm - Involve the organisers - Involve the audience - Q: Acknowledge, or ignore? - Q: Is this your problem, the organiser’s, the venue’s? - Q: you’ve got people’s attention – do you keep it and lead, or defer #### **Responses: How to cope** - What needs to be dealt with? The message, or the problem? - Some things don’t matter. - Be on the audience’s side – they’re on your side. Empathise with what they want / need – message / information, entertainment, leadership. - Stay calm, recognise a proportional emotional response. Respond, rather than react. Be aware that you’re probably already over-alert. ## Example disasters - Someone in the audience becomes ill - Someone storms out - The audience starts to leave - Your interactive workshop has too few people in it - Your interactive workshop has too many people in it - No one shows up - You’re forced to delay the start by 15 minutes – but not the end - The projector dies - There’s no sound - Something worked before the talk, and fails after the talk starts - Your laptop dies and you rely on slides - Your demo doesn’t work - You need the internet. There is no internet. - You get dreadful news just before you go on - You forget what you’re going to say next - You spill something - Fire alarm - An organiser interrupts - Your phone goes off - Someone emails you mid-talk, and their email comes up on-screen - Your screen shows the audience sensitive information - Hecklers - Code of conduct issues - You say something that you instantly regret - You reveal sensitive information about your project / organisation / client ### Current LLM tech stack URL: https://www.workroom-productions.com/current-llm-tech-stack/ Last updated: 2024-11-29T23:39:39.000Z *because people ask, not because I recommend* November 2024: For coding, I use [Cody](https://sourcegraph.com/cody) in my local IDE VSCode and whatever is on offer in cloud IDEs (recently [CoPilot](https://github.com/features/copilot) and [ReplitAI](https://replit.com/ai)). For getting my code to talk with LLMs, I use [Simon Willison's llm](https://llm.datasette.io/en/stable/). For running models in the cloud, I use [Replicate](https://replicate.com). I've used [Ollama](https://replicate.com) locally – but honestly I haven't touched a local model recent months. I use [DrawThings](https://drawthings.ai) to explore image-making locally (but while generative, that's not using an LLM). For general stuff (including explaining code and learning new stuff), I currently use [msty](https://msty.app/), which allows chat-like interfaces with several AIs via their API. This lets me set up the context for a chat with system prompts, and handily means that I can pay-as-I-go with AIs, which is rather cheaper for me than a monthly subscription. I'm currently using [Claude-3.5-Sonnet](https://www.anthropic.com/claude/sonnet), [GPT 4o](https://openai.com/index/hello-gpt-4o/) and [GPT 4o-mini](https://openai.com/index/gpt-4o-mini-advancing-cost-efficient-intelligence/) via their API ### Learning with an LLM URL: https://www.workroom-productions.com/learning-with-an-llm/ Last updated: 2024-12-08T22:45:32.000Z *tl;dr: use examples, get chatting, don't take the first answer, work in areas you know.* Asking an [LLM](https://en.wikipedia.org/wiki/Large%5Flanguage%5Fmodel) to explain stuff is a new learning technique for me. An LLM is a fine [rubber duck](https://en.wikipedia.org/wiki/Rubber%5Fduck%5Fdebugging): setting out a clear question requires me to engage with the subject carefully and precisely. Reading my questions helps my own imagination to pop up insightful answers. Critically parsing any response makes me think about what is plausible and consistent, which makes me think about the underlying models that I'm using, and how they might change. However, the *quality* of the LLM's responses is... unreliable. The words generally fit together into a clear explanation, but the information explained can be false. Key points might be left out, irrelevancies included. I've seen answers that purport to be about shell but include syntax for python, legal advice with no legislative support, opinions offered as gospel but as reliable as gossip. An LLM is very polished random walk through a lossily-compressed and limited repository. If you're planning on learning with it, you need to have ways to cope. To limit the loopiness, I find it useful to ***use an example***. LLMs are, by training, plausible and confident – in the necessarily shifting sands of new knowledge, an example is more solid than the LLM's ramblings. On the one hand, I would ***ask for an example*** of executable code or tests. These are easy examples to work with, as running them will give you immediate definitive feedback. In other areas, I've found it helpful to ask for relevant legislation when learning about legal stuff, or for specific quotes when asking about books. Most of the time, requests for specifics will either give you a source for what you need, or indicate that the LLM's response is down to the opinions in the work it has ingested, and not necessarily supported. On the other, I also found it helpful to ***ask it to explain an example*** – I used it to pick apart the output of a linter, test and compiler failure messages, and examples of grammar. It helps to ***get chatting***, primarily because a conversation engages me more with the topic: An LLM has no beliefs, so feels refreshingly open to changing its min, particularly if there's more in the training data to [(stochastically) parrot](https://en.wikipedia.org/wiki/Stochastic%5Fparrot#Purpose). I've recently heard `that's a great discovery!`, `thank you for that hint! ` and `You're right`, as well as the regular `I apologise for the confusion` as the LLM spins on its metaphorical heels. What an *encouraging* rubber parrot it is. A less wooly reason to converse is that your corrections and hints become part of the prompt, shifting the perspective for the rest of the conversation. This is a subpattern of something that we should be familiar with, as testers: ***don't take the first answer***. Following that principle, I've found it useful when learning to: - ask the same thing in a new conversation, or with a different system prompt - ask for alternatives - ask several LLMs - ask the LLM to take apart its own answer. I find that ***learning with an LLM is best when I already have insight***: I learn better on subjects where I already know enough to tell good ideas from dumb. I'm pretty happy learning from an LLM about code patterns and syntax, or around application of a tiny and specific subset of UK education law. I'd not want to dig into Nelson's naval tactics, or the biochemistry of pollens. And there are some things where written language means not an awful lot, so I think I'd avoid using an LLM to discover [Gabber](https://en.wikipedia.org/wiki/Gabber) or [Hyperrealism](https://en.wikipedia.org/wiki/Hyperrealism%5F%28visual%5Farts%29). I imagine that most of us pick our learning materials to suit our expertise and learning style, so this is not a novel principle. Nonetheless, an LLM's fluency and speed makes can make it appear temptingly useful as a start point or as an expert, and it can be misleading in both circumstances. In the right circumstances, an LLM is a delight to learn with. If you can stand to personify your tools, it's easy to cast it in the role of an open, informed, encouraging, helpful, endlessly available interlocutor. I've never had a study buddy like it. As an example, here is what I learned in a recent single conversational thread with Claude. I was requesting information around shell scripts (whether built, suggested or generated). Often, I had run [ShellCheck](https://www.shellcheck.net) and was wondering about ShellCheck's reasons and about my alternatives. I learnt masses (particularly where the scripts themselves were obscure). The LLM gave me relevant information in moments, explained it well, and I could immediately try it out. Claude reminded me of: - the difference between `''` and `""`, - spaces around `=` , - the absence of booleans, - `[` as a 'test' command. Claude helped me to use approaches which were new to me, including: - functions in scripts, - here-documents and here-strings, - moving parameters around, - do-nothing with `:`, - `set -x` and `set -e`, - comparing `[` with `[[`, - expanding with `${llmCodeContentParameter[@]}`, - how to set up a config file, - running executables with `./«name»` compared with `sh ./«name»` or just `«name»`. And, aside from explaining, Claude could - plausibly pick out idiomatic oddnesses and root cause of syntax errors, - offer options for how to do things more clearly, - unpick regex, - spot antipatterns, - give plausible pros and cons, and - propose ways to strengthen to code against common failures. The LLM put forward bad ideas throughout the learning process. If I'd not been able to temper its dreaming from the deterministic side by executing the code, and from the thoughtful side by my own insights into code, I'd have been conned. A little voice wonders if I have been conned anyway, in some way that I don't see right now. I pay attention to that feeling, but can't address it myself: if you can help me work out what I've missed, let me know below. ### Taming 2000 Safari Tabs URL: https://www.workroom-productions.com/taming-safari-tabs/ Last updated: 2024-09-11T18:56:56.000Z My machine is [:slow](#slowbox). Well, it's [fast](https://www.cpubenchmark.net/cpu.php?cpu=Apple+M2+Max+12+Core+3680+MHz), but the experience of using it feels hesitant. Perhaps that's because, with three Apple devices sharing tabs, and an ill-disciplined approach to looking stuff up, I have [:countless](#countless) tabs open. I wondered whether [Claude](https://claude.ai/) could help. After a couple of queries to ask how to manage my tabs and to see if it knew a tool I didn't (reader, it [*did*](https://apps.apple.com/gb/app/tab-space/id1473726602?mt=12)), I asked this.... > I'd like to create a custom automator script to extract details of my tabs from safari. I guess that the details I want are the title, URL, parent window details and any date information. Can you help with that? Claude gave me code, instructions and alternatives. I followed its instructions and ran the supplied AppleScript in Automator. The output let me see that the numbers were credible: there were *thousands* of open tabs. After some [:fiddling](#fiddling) (see below), a Shortcut-wrapped version of the script produced something like this but with much larger numbers... ![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/2024/09/image-1.png) A modal dialog showing text detailing my open safari tabs, summarised with an overall count and a list of the hostnames that have most tabs with. Don't you go making assumptions about my politics... So that set me rolling: closing and bookmarking and switching and sorting tabs. From time to time I hit that Shortcut and saw the number of open tabs going *down*. The numbers went down by tens, then hundreds. Each of my top clusters of awfulness rolled up like a woodlouse and vanished – pop – like a dull grey bubble. *How* satisfying. Immediate, on-demand, consistent reports helped to keep me going. I could come back after a break and still feel progress. I could open a few more tabs (because crap habits don't die easy), and not feel so hopeless. The situation that slowed down my beast of a machine for six months turned from a problem into a story and a bit of code to share. Getting the tiny tool from an idea to good-enough took an hour – at least half of which was spent on an unnecessary integration. Managing the tabs took perhaps 6 hours overall. The tool was valuable because it changed my output, not because its numbers were perfect. I needed the tool to motivate me to do the work, and didn't need to trust it to do the work itself. Without the tool, the work would not have been done, by now. I could be spending months glaring at that beachball. If I hadn't had access to a magic box incorporating a compression of half of humanity's code, I might have [:spent frustrating hours](#needignore) on the trivial script. And as with any script, it's useless without its complex surrounding system. Let's not imagine that the code is either mine, or Claude's – it's a speck standing on mighty shoulders. Here it is: ```JavaScript function getHostame(urlString) { // need to use regex as JS for AS doesn't appear to have a URL object // could use const parsedUrl = new URL(url); const hostname = parsedUrl.hostname; const regex = /^https?:\/\/([^/?#]+)(?:[/?#]|$)/i; const matches = urlString.match(regex); const hostname = matches && matches[1]; return hostname } function getTopHostnames(tabs, limit = 5) { const hostnameCount = tabs.reduce((listOfHostnames, tab) => { const hostname = getHostame(tab.url); listOfHostnames[hostname] = (listOfHostnames[hostname] || 0) + 1; return listOfHostnames; }, {}); return Object.entries(hostnameCount) .sort((a, b) => b[1] - a[1]) .slice(0, limit) .map(([hostname, count]) => ` ${count} of ${hostname}`); } function getListOfTabs(input, parameters) { var Safari = Application('Safari'); var tabInfo = []; Safari.windows().forEach(function(window) { window.tabs().forEach(function(tab) { tabInfo.push({ title: tab.name(), url: tab.url() }); }); }); return tabInfo; } function sendOutput(numTabs, topSites){ const message = "You have "+numTabs+" tabs open. Common Sites: "+topSites console.log( message ) // send to stdout //return message; // un-comment to pass to ScriptEditor output } const listTabs = getListOfTabs() const numTabs = listTabs.length; const topSites = getTopHostnames(listTabs, 10 ); sendOutput(numTabs, topSites); ``` I have a Shortcut that picks up this script, uses the `Run Shell Script` action to call `osascript` to run it, and the `Show` action to pop it up where i can see it. ![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/2024/09/image-2.png) A screenshot from Shortcuts showing how I use "**Run Shell Script*" to pass my script to `osascript`, and pass the results to "**show*". Note that this is old – the file is suffixed `.js` now. So that's my tiny handy tool. I used it intensely for a weekend, and I'll probably forget about it in a week. Have it if you want it. Use it but don't trust it. It's ephemeral, it's got no tests, it's a means to an end, it's yours. Better yet, make whatever tool *you* need to support the task that's bugging *you*, and **tell the rest of us**. --- ## Fiddling the script The numbers were OK, but the details were dumb. Seeking to polish out the dumb, I asked Claude to sort the set by url and count the top 10 most-common hostnames. It wrote a bubble sort (because training data, I guess), so I gave that a jiggle towards a native sort with another prompt. Then asked it to switch the [:AppleScript into JavaScript](#OMGAppleScript). The suggested script used [URL](https://developer.mozilla.org/en-US/docs/Web/API/URL/URL), and [JavaScript for Automation](https://support.apple.com/en-gb/guide/automator/auta229c77c2/mac) baulked / borked at that. Once I'd asked for a regex to hoik out the hostname instead, all was broadly OK: the script was producing the same count, more-useful output, and now had clearer code. I broke out a few functions and parameterised a magic number, to give me some sense of agency and to make the thing clearer still if I came back to it. Then, wanting a touch of integration. I switched from working on the code in Automator, which is a pain, to working on the code in Script Editor, and running the executable in Shortcuts. With the benefit of hindsight, I'm actually not sure precisely *why*. But the decision cost me. - Could I get Script Editor to ouptut a `\n` newline? I could not. - Did the script give me half an hour's debugging arseache with a rogue empty-but-not-really tab? It did. - Did I resort to idiot debugging and tiny iterations? Yes: I blame Script Editor, and I draw your attention, fellow debug-bruised Script Wranglers (and me when forgetful), to the `Log History` in Script Editor's `Windows` menu. - Did Shortcuts try to ask permission to open every single URL when I put the output into a notification – and am I therefore happy with the ugly compromise of a poorly formatted dialog? Why, yes, and yes again. Ooof. 40 mins I won't get back. So that's the more-human part of AI-assisted toolsmitherly. When writing this article, I smoothed some more bumps. - I did more re-naming and refactoring. - I took out an unused library, and two pointless variables. - My asked-for and twice-tuned step to sort the list of URLs was unnecessary, given that Claude's way of finding the top hostnames doesn't require a sorted list. It's gone. - I disliked Claude's way to get information out of the script, but having a `run()` function that returned a value, and that value happened to go to `stdout`. I wanted the `return` to be out of a function so it was clear what was coming back. AppleScript allows a `return` outside a function, but JSX doesn't. So I took more AI advice, switched it to `console.log` and put the whole thing in a findable wrapper. - I switched the code from a `.scpt` file to a `.jsx` file to reflect the contents. It worked. - I switched again to a `.js` file, to see if I could edit in my usual editors. I could, and I could also get rid of some Script Editor-imposed cruft at top and bottom, so now the code is `text`, rather than `cruft`\-`text`\-`cruft`, which seems right for code. --- ## Footnotes ### : slowbox When Safari is open, I find that I'm terribly frustrated by the blink-long gap between button-press and glyph-appear. And the beachballing on bookmarks. ### : countless Not *countless*. Of course they can be counted, but it's not worth my time. I'll get my machine to count them. ### : needignore More likely, I'd have ignored the need and continued to be mildly frustrated. ### : OMGAppleScript frankly, writing AppleScript is like playing an 80s 'natural language' text adventure game. You know; the ones where you fail with a dozen variants of '*Tie the rope to the bucket*' '*attach the cord to the handle*' '*make a lowerable container*' before you give up and succeed irritatingly with '*throw the bucket in the well*'. I loved those things, but there's [better stuff](https://www.nomanssky.com) to do with one's attention now. --- ### Guiding Hands-off AI using Hands-on TDD URL: https://www.workroom-productions.com/guiding-ai-code-with-tests-workshop/ Last updated: 2024-12-10T13:06:40.000Z *Bart Knaack and I ran this* [*hands-on workshop*](https://agiletestingdays.com/2024/session/guiding-hands-off-ai-using-hands-on-tdd/) *at* [*Agile Testing Days*](https://agiletestingdays.com/2024/)*.* *We updated this page (occasionally) as we built the workshop, to share the ways that we found to guide AI towards code that passes automated tests, and the stuff we've tried that hasn't worked for us. I hope that this page will give you an insight into how we set up an interactive workshop.* ## Workshop Delivered! 27 November So the summer slid past us, life happened, and we concentrated our limited time on what was going to be in the workshop, rather than putting stuff here. I'm writing this about a week after. We built stuff and trialled exercises. We saw that the LLMs, being *language* models, are guided by function names and by comments more than they are guided by tests / checks – but that the deterministic executable tests helped the code to stay working, even as its form changed. We saw that the LLMs would add comments, sometimes, about what they had done – more than once we saw a comment along the lines of `This is mathematically incorrect, but is included to pass the tests`, which was excellent. The night I arrived in Potsdam, I warmed up by building – from tests – a route-finder which I would not have been able to build from scratch without books. As I made it more complex, the LLM re-wrote and decorated the underlying data structures so that the tests would continue to pass. Eventually, the code (which included without me asking for, a depth-first search) could give me three routes – fewest stops, lest distance, least time – which would differ under expected circumstances. We saw that the same tests could land up with different code – sometimes clear, sometimes obscure, sometimes algorithmic and extensible, sometimes filled with special cases. We had a couple of tech-related wobbles, particularly as our handy workshop IDE, Replit, changed its business model making it harder (or unfeasibly expensive) to do what we had expected to do. We persisted with Replit, but will move away in the long run. We got the magic loop working more and more reliably, introducing syntax checks for the tests, keeping the 'conversation' flowing, making the logging more readable, and giving participants a choice in the AI they used. When sat side by side in the lobby of the Dorint as the tutorial day ebbed and flowed, we found that seeing what the AIs made (and how each other worked) gave us insights we hadn't had working remotely – which is a good sign for an interactive workshop. The workshop went well – despite Replit oddnesses, and LLMs telling us they were overloaded. [Dragan Spiridonov wrote](https://www.linkedin.com/pulse/back-from-agile-testing-days-2024-reflection-learning-spiridonov-oxo5f/?trackingId=cZUFTqvnSrOSm7iHAGIj9w==) that he found it *fascinating* and *mind blowing*, and [Christoph Zabinski was thrilled](https://www.linkedin.com/feed/update/urn:li:activity:7265794898069491712/) that our session and others *didn’t just ride the AI hype train—they provided thoughtful, grounded perspectives on how AI can truly enhance our work as testers*. We were delighted to have such an engaged group. Someone asked about the cost of the LLMs. The total AI transaction cost for the 2-hour, 30-people workshop was well under £5 – which covered >1000 requests and >1.3M tokens. We gave people access to Anthropic's Claude and OpenAI's 4o-mini – and note that, despite similar usage, Claude cost nearly 40x 4o-mini. However, we also note that 4o-mini got stuck on special cases, and had the irritating habit of wrapping python in ```` ```python ``` ```` . Claude $3.79 \~750 requests, 781K tokens OpenAI $0.10 \~350 requests, 520K input, 62K output ### Base15 – 12 July I built a chunk of code to convert decimal to Base15 today. I specified the tests, one by one, and after each I asked the AI to build (or adjust) code to pass those tests. Which it did, fairly successfully. All this is working in Replit, so I can give it to people in a workshop. I've put the [repo on GitHub](https://github.com/workroomprds/replit%5FBelayUsurySurroundsClapping/tree/main) so that you can see how it built – you'll also see the current state of my shell script which passes stuff to the LLM, runs tests and goes back to the LLM if things fail. Have a look at the "thoughts" section in the "Bucket" heading below to see some more on what happened. I'll bring those bits up here shortly. I've adjusted my script so that it checks in the code, if it's working, which gives me a sense of progress and lets me compare and rewind the code that's been made. ### My Magic Loop is working! I wanted part of this workshop to seem magical – and whereas writing code that appears to correspond to what you've written about is astonishing, it's no longer magic. Especially to testers, who see that the code is often just awful. Copying and pasting code suggested by CoPilot or Cody, then running the test suite is repetitive and clerical. I want to automate that away. The magic I'd imagined is that participants add an (automated, confirmatory) test, step back while something else builds code that passes that test, and step in to see how weird the built thing might be. Here's the chunk of shell script at the heart of this magic loop: ```shell llm -t rewrite_python_to_pass_tests -p code "$(< ./src/$1)" -p tests "$(< ./tests/test_$1)" -p test_results "$(pytest ./tests/test_$1)" '' > ./src/$1 ``` Let's unpick: That line starts off with [Simon Willison’s llm tool](https://llm.datasette.io/en/stable/) – a tool that acts as an interface to generative AIs. The command takes a parameter (`$1` above) which is the name of the source file to be changed. It uses that parameter to gather data in three lumps indicated with `-p` and named `code`, `tests` and `test_results`. Each data gathering bit runs a tiny shell command: `<` to get the contents of a file and `pytest` to run the tests. So I'm labelling and sending in the code I expect to change, the tests I need the code to satisfy, and the results of running the tests through the code (remember I expect the tests to fail, and that the failure info is in some way helpful). I also expect the AI to make sense of these three. `llm` uses all that labelled information to fill in a 'template' called `rewrite_python_to_pass_tests` , fires the filled-in template to a generative AI, and waits for the output. `llm`'s actions are set up in the following `rewrite_python_to_pass_tests` template: ```yaml model: claude-3.5-sonnet system: You are expert at Python. You can run an internal python interpreter. You can run pytest tests. You are methodical and able to explain your choices if asked. You write clean Python 3 paying attention to PEP 8 style. Your code is readable. When asked for ONLY code, you will output only the full Python code, omitting any precursors, headings, explanation, placeholders or ellipses. Output for ONLY code should start with a shebang – if you need to give me a message, make it a comment in the code. prompt: 'Starting from Python code in $code, output code which has been changed to pass tests in $tests. Please note that the code currently fails the tests with message $test_results. Your output will be used to replace the whole of the input code, so please output ONLY code.' ``` I've already set `llm` up with a plugin and key so it can talk to an AI – in this case, Anthropic's *Claude* because it was released on Monday and all the nerds are gushing. Also, I spent a fiver on tokens there. `llm` gives the AI a *system* prompt to tell it how to behave in general, and a *prompt* to pass that data and set a task. I've done it this way to separate chracter from task. It also might let me, later, build out so that my script can carry on the conversation with the AI, keeping necessary context. When the AI hands back what it's generated, `llm` hands it to the shell and the shell **overwrites the source file** with whatever `llm` spits out. So that's the line. The line lives in a short shell script, which runs the tests again on the new code. If the tests run without failure, the script stops, with a message that the code is ready for inspection. If not, it iterates a few times, and will report the test results if the code still doesn't ~~pass~~ satisfy the tests after a few ~~passes~~ goes. And that's the magic loop. You write *tests*, run a script, the *code* changes and the tests pass. Mostly. Takes 5-20 seconds. Magic over: What's next? **Either** the newly-working code is ready for inspection. Perhaps, participants will fire up the system to explore, run a diff, write more tests to generate more code, or just commit and move on. **Or**, the script's done and the code is bust. Maybe one just runs the script again to see what it does this time. Maybe one fixes the code directly. Maybe one looks at one's tests and realises that the tests are inconsistent. Maybe the AI has vandalised the code so much that you go get the last one out of change control. In my limited playtime, I've been amused to see comments from the AI in the code to indicate that the tests are odd, but that the code has been adjusted to pass them anyway. *That's* how you pass as a thinking thing. Welcome to the team, Claude. 🫢 Yes, it overwrites the code. No, you can't get the code back with an undo. That's what change control is for. And besides: code is cheap. --- ## What we've tried, and might try - Working in the IDE - Working in a shell script - Cody - Copilot - OpenAI - Ollama - Claude On my rough list of next steps: Moving to Python ([me? *shell??*](https://www.youtube.com/watch?v=WoBLi5eE-wY)), trying different AIs, continuing the conversation with the AI, automated checkin, working in Replit, custom models in Replicate, changing prompts, building something odd, building something useful, multiple files / tests, different prompts for different purposes, mapping the (financial) costs, imagining just how many ways this can go wrong or is already wrong... --- ## Workshop Stuff – video and abstract In this hands-on workshop, you’ll write tests, and an AI will write the code. We’ll give you a zero-install environment with a simple unit testing framework, and an AI that can parse that framework. You’ll add to the tests, run the harness to see that they fail, then ask the AI to write code to make them pass. You’ll look at the code, ask for changes if it seems necessary, incorporate that code and run the tests for real. You’ll explore to find unexpected behaviours, and add tests to characterise those failures – or to expand what your system does. As you add more tests, the AI will make more code. Maybe you’ll pause to refactor the code within your tests. Bart and James are exploring the different technologies and approaches that make this possible. We’ll bring worked examples, different test approaches, and enough experience (we hope) to help you to work towards insights that are relevant to you. All you need to bring are a laptop (or competent tablet) and an enquiring mind. You’ll take away direct experience of co-building code with an AI, and of finding problems in AI-coded systems. We hope that you’ll learn the power and the pitfalls of working in this way – and you'll see how we worked together to find out for ourselves. --- ## Bucket This is where I'll keep unsorted stuff. Some is to look out, some to discount, some are threads and others are pits. I'd ignore this, if I were you, but as I'm me I'll use the contents to feed my filters and pipelines. I'm adding to this. Big job to go in and get all my bookmarks and make sense. Maybe I won't. An ideal "entry point" – replit, all set up with tests and the shell script / LLM tool / key to and API, with a test file that people can use by un-commenting a test and seeing what gets made. ### I've found that: - Testing errors make for weird code – introduce a test which is wildly inconsistent (by, for instance, duplicating a test and changing the output, but not the input) and the code can get way more complex. I've noticed that if my Python test is syntactically incorrect (missing a `:` for instance), the LLM will often include that python in its own output. - Despite my efforts, the LLM often puts a note at the top to say (something like) `here's the code that passes your tests`. The second time round the mgic loop, the line tends to get binned – but that second pass costs me another chunk of tokens (and another cent or two of money and ¿howevermany? kJ of power and ¿howeverothermany? g of atmospheric carbon). - The order in which you introduce experiments has a big influence on the code made. As it does with TDD – good TDD is based in part in thinking about this, too. - Commented-out tests still influence the code – remember that *we're* running the tests, not the LLM. We ask the large language model to parse the results and the test file. So it maybe reads the commented tests as things we'd like. I may need to refine my prompt to stop this: If we left this behaviour in for the workshop, it would mean that an exercise where we start with lots of commented-out tests may just write code that suits all the parts to be revealed. Or we may need to re-think an exercise approach. - Here's a story – if I find a pattern of these, it's a principle. At its heart is that, **in response to one extra condition, the LLM quite reasonably made the code three times longer and far more obscure**. I'd written [tests to guide the LLM](https://github.com/workroomprds/replit%5FBelayUsurySurroundsClapping) to write a [base15 converter](https://www.inchcalculator.com/base-converter/), introducing a descriptive name, then the experiments 0 -> `0`, 1-> `1`, 10->`A`, 11-> `B` 15->`10`. At each stage, the LLM rewrote the code. The code was readable, expressive and neatly recursive – and built to cope with reasonably-sized positive integers. I introduced a experiment for negative numbers: the LLM made a good change. I introduced a fraction 15.6 -> `10.9`... and noticed that the magic loop had failed *first* time round, trying to match `10.8EEEEEEEEEEE` with `10.9` . You might [recognise that problem](https://en.wikipedia.org/wiki/Floating-point%5Farithmetic#Accuracy%5Fproblems). However, that was the first pass. The *second* pass flew through the tests. On inspection, the LLM had inserted 45 lines of code to do floating point maths (here's [before](https://github.com/workroomprds/replit%5FBelayUsurySurroundsClapping/commit/587df471bc2086af541d4385dbb242d76bc1c8d3) and [after](https://github.com/workroomprds/replit%5FBelayUsurySurroundsClapping/commit/7c36c9976f15497763a243dbeb9538ebbda40a2a)). That code, which is commented and readable, but beyond my ability to understand properly, limited precision to 6dp, stripped `0` and `.` when needed, and did *something* with rounding up from `E`. Frankly, my jaw dropped – we've all seen the extra requirement that knocks a hole in the floor of the problem space to expose an echoing and unexpected cistern below. The LLM just... coped. Faced with the same problem, I'd have gone off to find a library to use, with the problems that brings. The LLM made its own, with the problems *that* brings. Floating point representational fuckery is the source of the complexity; if the LLM had addressed that trivially, it would have been *wrong*. - To continue the story beyond the part that might turn into a principle, limiting precision to 6dp is a fluttering rag to a colourblind testing bull. Two tests suggested themselves; one to use a number that needed to be expressed to more than 6dp (example: 0.123456**7**), and another to increase the size of the number so that 6dp is impossible to fit in (imagine fitting 1234567890123456.78 into something that works with 16 characters). I started with the first; the LLM wrote code to successfully convert 14.888 -> `E.D4C`. However, my next experiment proposed that 14.887 -> `E.D48959595959` . Let's note that I deliberately chose a number which has recurring digits in pentadecimal: this number is never going to be checkable without stating what precision should be used in the comparison, so the test is *purposefully* shite. The LLM hasn't been able to get my rubbish 'test' to pass without causing another to fail. Interestingly, that failing test is one that it was previously able to pass. The proposals no longer correctly do 15.6 -> `10.9`, but variously convert 15.6 -> `10.8EEEEEEEEEEE` *(perhaps floating point representation issue)*, or make 15.6 -> `109` *(perhaps inner `.` stripped)*. Am I trying to get the LLM to write code that passes tests here? Am I trying to find code that it can't write? Am I trying to prove my superiority to our [new artificial overlords](https://knowyourmeme.com/memes/i-for-one-welcome-our-new-insect-overlords)? I guess the principle here is that, if I'm using simple checks to direct an AI towards working code, perhaps I shouldn't confuse that with actually looking for trouble. - Stepping on once more, I knocked out my deliberately mean experiment, in the hope that the LLM might work back to something it could manage. It did not. It tried several times and got *worse*: I checked in the version that converted 15.6 -> `B.18EEEEE` for the sake of amusement and possible analysis. Then, I deleted the AI's code entirely (it goes in the prompt along with the tests, so might have been polluting the prompt), and tried once more. The LLM failed to write code that passes the 15.6 -> `10.9` test, lots. I took the tests back to the failing test, and rtied again. After several goes, the LLM finally wrote code that passed the tests. And, on inspection, that code contained the following snippet. Aha ha ha: ```python3 #Special case for 15.6 if abs(decimal - 15.6) < 1e-10: return "10.9" return sign + result ``` ### Thoughts - A system is not a file of code – if we feed a file and tests, we'll build stuff that fits in a file. A multi-file code-base needs a different prompt. A system needs a different approach. - There's more driving development than tests in TDD. A sense of intent – often shared as stories or requirements, and the soup of collective culture that those swim in to give them meaning – also gives direction and shows completion. An LLM might get that from your object / method names, or maybe from your comments. I ask for a method called `base15` (and no more) and the LLM doesn't write an empty method but includes `self.digits = "0123456789ABCDE"`. Much as I might want to influence the LLM only with my tests, my naming choices influence it too – I could choose obscure names to reduce that influence, if I had reason to do so. I don't, at the moment. I could lean into it with comments in my tests, too... ### Changes to tooling - consider non-AI tools (alongside the test runner) to do useful deterministic work swiftly and cheaply, and either give it back to the human (i.e. if the tests have a syntax error) or give it to the AI with the code, or give it as as feedback on the tests. Examples: syntax tool, coverage tool, tool to say if the tests actually ran - ~~Automatically stage the changed tests, the iteratively-delivered, code, optionally commit with message derived from changes.~~ Done it: OMG. - What to do when AI generates "too much" code? Notice with coverage... Recommend more tests? Remove the excess? - Look into multi-agent to review code, offer suggestions on performance / readability / security / patterns. - Look into different prompts / system prompts within the loop – a refactoring prompt vs a new-code prompt? - Look into different LLMs – on and off-device, too ### Coding targets As we learn about what we're delivering, we'll need things to build, run and test. As we deliver, we'll need things to build, run and test. *What things will we build? What will we ask participants to build?* 💡 This is where I (James) have got stuck (early July). I can't find something I connect with well-enough to write tests, use code, and to deliver for a workshop. What sources? #### Kata We can't aim at a Kata – AIs are trained on GitHub; Katas are prehaps a bit prevalent. Example: I kicked off by making aROman NUmeral thing; the AI filled in lots of code I'd not asked for, and it was all good. I set up a card hand checker, and the AI built me a deck with the usual four suits and the usual 13 values. TDD this is not: there's a cultural basis. The tests are purposeful and reflect that, so the code is often greater than requested. #### Coding problems So Bart and I need coding targets to try this stuff. Here are some: - , - [Kattis, Kattis](https://open.kattis.com/), - [https://adventofcode.com/2023/](https://adventofcode.com/2023/day/1) - - - - #### Something real Here's a [github search for Python codebases with TDD in the readme and some sort of activity](https://github.com/search?q=language%3APython+TDD+in%3Areadme++size%3A%3C%3D2000+stars%3A%3E%3D50+fork%3Atrue&type=repositories). I've been messing about with [expects ](https://github.com/jaimegildesagredo/expects) #### Something new [Pure functions](https://en.wikipedia.org/wiki/Pure%5Ffunction) will be more straightforward to test, and so to TDD, than anything with data persistence or side effects. For interesting stuff, the tests will need test stubs / harnesses, matchers etc. How will the LLM grok those? We could twist a kata. Bart suggests Arcadian numbers (like Roman but different letters and base 12), decimal <-> base 17 (not hex). With these, we could generate the tests and try repeating (and comparing) builds introducing those tests in different orders and clumpings. ### Links I hear that GPT-4 has a working python interpreter. Indeed, *had* a working interpreter a year ago. [What AI can do with a toolbox... Getting started with Code Interpreter \[Now called Advanced Data Analytics\]Democratizing data analysis with AI![](https://substackcdn.com/image/fetch/f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fbucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com%2Fpublic%2Fimages%2F233e9990-95ea-44ab-afc4-5abfe44f24d4%2Fapple-touch-icon-180x180.png)One Useful ThingEthan Mollick![](https://substackcdn.com/image/fetch/w_1200,h_600,c_fill,f_jpg,q_auto:good,fl_progressive:steep,g_auto/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F55543b86-b8ec-45db-84ec-649fe0237097_3208x2000.png)](https://www.oneusefulthing.org/p/what-ai-can-do-with-a-toolbox-getting) This is what CoPilot is doing: [copilot-explorerHacky repo to see what the Copilot extension sends to the server![](https://t3.gstatic.com/faviconV2?client=SOCIAL&type=FAVICON&fallback_opts=TYPE,SIZE,URL&url=https://github.io/copilot-explorer/posts/copilot-internals.html&size=16)copilot-explorer![](https://thakkarparth007.github.io/copilot-explorer/images/screenshot-v1.png)](https://thakkarparth007.github.io/copilot-explorer/posts/copilot-internals.html#how-is-the-prompt-prepared-a-code-walkthrough) OpenAI seem to have some success in teaching an AI to notice bugs. Better than human, apparently. Look out for false positives, it says. I say: *is it trained on automatically be-bugged software?* and --- ## Making this workshop, and making it interactive Interactive workshops – and most especially those that ask participants to work with technology rather than with each other – are risky. We've done plenty, and they always fail (and succeed) in unexpected ways. We learn loads – I'll share here some of what we're learning about how we deliver *this* workshop, and that might help us to share what we've learned about delivering workshops overall. ### Specifics #### Picking the right level We'll need different things for different people to do (?). Imagine a set of less- to more-complex. Which set might appeal more? Which can we deliver most clearly? Which gives an interesting experience? Which has useful output to talk about and learn from? What sets might we have? Mix-and-match? **Techniques:** A: coding with AI basics B: something more complex C: TDD vs AI. **Tests:** A: TDDvsAI with all the tests written, but commented out. Participants un-comment and run. B: Some tests and code, clear goal / requs, participants write own tests. C: A few short ideas, and the tooling to help. **Tools:** A: TDDvsAI with simple one-stop shell. B: with n-times loop. C: with loop and syntax / coverage. ### ### General principles – unfinished... Our general trick is to do as much of the infrastructural heavy lifting as we can, so that participants can get straight to testing work. Workshops are a great way to learn. If participants are learning how to download and install a tool, then that's fine... but I'd prefer to get as much of the tool-sourcing out of the way so that we can get to the tool *use* and through that to the thinking. We try to work within the constraints of a conference environment – short sessions, varied skills, random kit, flaky wifi. We've been known to bring laptops and routers and servers: now we configure tools that can run in browsers on tablets. > Aside: Conferences need places for people who work with technology to play with technology. We set up the TestLab at conferences because we recognised that testing conferences would be enhanced by having somewhere, and something, to test. Bart and I find a challenge in making interactive technical workshops. It's a privilege to do them, and it's painful to get to a point where we can do them. There is *always* a vale of shit that we need to pass through, where the whole premise seems misguided, where the workshop seems undeliverable, where we've lost our connection with each other, where we have no sense of the experience we'd like to deliver. And we'll try to work through or round all those things. The way to work is often simpler than we'd imagined, and that simplicity is often invisible before we've done the work. Frequently, we've bought (I've bought) the complexity, and we need to let something go to find a way through. Knowing that we have to deliver something is a great way to focus on the good bits. Knowing *why* we're delivering something is a great way to stay on track, and that sense of purpose is is something we've built over years of ### Making a Record of your Exploratory Testing URL: https://www.workroom-productions.com/making-a-record-of-your-exploratory-testing/ Last updated: 2024-02-01T16:50:03.000Z *A mashup of sources, for whittling later. Here's* [*what I mean about the whittling*](https://www.workroom-productions.com/how-i-am-writing/)*. I'm* [*aware*](https://twitter.com/workroomprds/status/1753096870182973860) *that* [*LambdaTest*](https://www.lambdatest.com)[*featured*](https://twitter.com/lambdatesting/status/1753087153750647270) *this post, and I'll sort it out shortly.* Most of the time, exploratory testing records are kept privately for the **tester** who tested. Plenty of testers rely on their memory. Testers working in **teams** might use those notes to illustrate what they did and what they found, or to help them share how they worked, or to help work out why they worked as they did. **Organisations** might want those notes to be kept around for audit, or as some unspecified artefact for a fuzzy use some time in a possible future. If you're in an organisation that wants to keep notes like this, look into the purposes, benefits and costs – you may find that some needs are illusory, or un-owned. If you find that one group wants every artefact kept, but there's no way to search and no budget to store, then there's some organisational cognitive dissonance going on; you can choose to find out what un-corporate arse is being covered by this nebulous need, and choose to go give it a kick. **I** keep records to help my mind to remain available in the present, and to support other minds later. ### Purposes: in Three Rs - Remember – Mnemonic help (me) - Review – Sharing, improvement (me, my team) - Return – Historical analysis, long-term project memory (whoever comes along) ### What to record The following is lifted from my note [What to Record](https://www.workroom-productions.com/what-to-record/). > If you make a **plan**, write it down. If you're just tootling along all planless, you need a **strategy**, an **approach**. A sticky note will do. There are no excuses - accept no substitutes. > You'll want to remember the **actions** you take, the **data** you use, your **expectations**, your **observations** \- including **the time**. Don't necessarily limit yourself to exactly what you're testing - you're working in some kind of **context**. You'll get better at this over time; there's an instinct that comes with practice that lets you separate the wheat from the chaff. There's always going to be a bit of chaff. > Keep track of **things that repeat**. Even if nothing happens. Dullness is a virtue in most working systems. And without track of dullness, how will you notice . . . > **Surprises**. Is that a goat among the sheep? If you didn't expect it, it's worth writing down. If someone else wouldn't expect it, it's a **bug**. Perhaps you've seen an **exploitation**. Have you a **hypothesis**? Are you making a **model**? And when you've supported your hypothesis, found a potential bug, had a surprise, or the dullness is just too much to bear, you need to . . . . > Make a **Decision** \- many people get so used to testing by instinct, or by the book, that they don't notice they're making decisions. Worse, they've no idea what the decisions might have been. Scripted testing can be decisionless, but decisions are key to exploration. When you decide to take a different approach, to try different data, or just to consciously do exactly the same thing again, but watching more closely this time, you're taking a decision. Make a quick note. For me, recording *Decisions* is the key to remembering all the other stuff, and here are three flipdowns for other stuff... #### Identifiers: - who - when - what #### Qualities: - risk - estimated time - dependencies #### As you go: - actions - events - data - expectations - bugs - plans - interruptions - actual time taken - time wanted - problems 💡 ****Notation** **`-`** Item **`*`** A more important item - sometimes used for 'return to this' **`!`** One you'll want to remember at the end of the test. **`!!`** is typically a bug, sometimes qualified with a spare **`?`** or **`¿`**. **`[`** An aside - a thought or observation that needs to go down, but that isn't in the flow`]` **`¿`** Something I'm not sure of - may need more tests **`?`** A question for someone, or something **`Plenty of arrows and circles`** \- not forgetting diagrams, underlining, tables etc. ### Use of Records Some audiences – including future you – will want to use your exploratory notes. They might want to see what's been done already, to have evidence of a problem, to look for the absence of evidence of a problem, to see what else might be done, to look for new leads, to learn about how you worked. Searchability is important for this – and so electronic forms may be more useful; typed notes, keylogging and files, video transcripts. Some test teams standardise on one note-taking approach; a massive mindmap, a shared OneNote, rich-text on a session-based Jira ticket, a set of sequential blank books. Some don't standardise on notes taken in-the-moment. Others put all long-term information in the issue tracker. What do **you**, your **teams**, and your **institution** do? --- On the [Ministry of Testing Club's discussion board](https://club.ministryoftesting.com/t/testing-and-notes-evidence-etc/73475/29), I wrote: > In my experience, people / committees making processes which require screenshots / videos sometimes *imagine* this need. If you have the chance to look at the decision to include the need, you might dig into what the recordings might be used for, who might be using them, and whether the expected benefit is worth the practical cost. > In one medical device org, Audit (when asked) were clear that they only required such records when seeking to know that a known problem had been fixed, and the fix checked. The process owners wanted most checks like that to be automated – and with that clarity, the long-term need to make and store detailed + searchable notes basically went away. > The testers, however, used more detailed (and more temporary) records to illustrate what they’d found, when sharing within the team. The benefit they saw was that spreading that out shared skills, and bought greater expertise to bear on that path through the system. I imagine that it also made the team more resilient to departures. Looking at the rest of this thread, it’s worth noting that their target didn’t have much of a screen-based UI, and that their exploration was typically around changing setup and simulated environmental inputs, and measuring outcomes and some internals. > In a regulator, I saw testing notes (made in Word / Notepad / markdown / knotted string) attached to whatever represented the act of doing work. The org used Jira *and* ADO *and* wiki *and* OneNote *and* auditable doc storage – and I saw notes kept in all those places (relying on fragile links). Each approach suited (and was made by) its small, typically isolated group of testers. When (rarely) people outside these teams asked for older or more-detailed records, those outsiders wanted, in effect, magic recall. From an organisational point of view, those notes were unfindable, unsearchable and unknown. > If / when I teach this stuff, I ask people to think of purpose by framing for their audience (us / people who know us / people who don’t know us) and timescales (right now / at a foreseeable juncture / later than we imagine). And, in terms of what to record, there’s the last few paragraphs of [What to Record](https://www.workroom-productions.com/what-to-record/). Which were written in a fever dream half a life ago, and so demonstrate that one’s notes, written for you for right now, may still be useful to someone you’ve not met, who lives in some unimaginable future. ### How I'm writing URL: https://www.workroom-productions.com/how-i-am-writing/ Last updated: 2024-02-01T16:49:39.000Z I'm [:gardening](#gardening). Expect my stuff on this site to grow and change and be re-arranged. Expect words and toys and videos and transcripts and lists and tools. I hope that you'll find this collection to be an interesting place to explore. Don't expect the text to stay static. These are not magazine articles, this collection is not a blog or a diary or a textbook. In my more ambitious imaginings, this site is an *experience*. I hope that it is, [:like some gardens](https://en.wikipedia.org/wiki/Kew%5FGardens), a pleasant place to explore. Cheers – James *...more details...* I'm building a few facilities for readers; some to make these pages more-explorable, some to make them more usable: - **Backlinks** – you'll find a button at the bottom of each page showing what pages link *in* to the page you're on. - **Search** – the search tool will let you tun a fuzzy search of most of the text of the site. Use your browser to search the page. - **Dynamic footnotes** – I'm using [nutshell](https://ncase.me/nutshell/) from [Nicky Case](https://ncase.me/nutshell/) to give you anchors which expand to give you more text, but which might be a distraction to the flow of the words. Links with a : in front mean you get to glance at related content, rather than re-routing your exploration to see it. - **Canonical links** – will help me to make pages and variants without breaking links, though you may find that the text may have changed. - **Change indicators** – you can see when the page last changed, and I'll indicate state and edits [:somewhere](#Visible%5FChanges) - **Comments and contributions** – I hope that my readers want to share their thoughts. Comments are *on* for subscribers, and I can give you editing access for deeper and more collaborative stuff. - **Expanding sections** – I'll put some topics in sections that fold up (or unfold, depending on mood) so that you can get them out of the way (or bring them up) as you need. - Change control – I'll put big revisions on GitHub. - Table of contents – planned for big pages - Lots of ways in – this isn't a book, to be read from start to end. It's a web. - Tags – each page will have several tags, each tag will take you to a list of other pages with the same tag, and the tags will arise as we go. - Print to pdf – when I sort the CSS, you'll find that the current state of the page will print nicely to pdf. ### Footnotes #### Visible Changes Currently, italicised text at the top. This is likely to change. #### Gardening Here are some sources - - - ### Limiting Testing With A Timebox URL: https://www.workroom-productions.com/limiting-testing-with-a-timebox/ Last updated: 2024-01-20T23:45:46.000Z *placeholder – needs to be re-written, expanded with my usual stuff and* properlyintegrated*.* Timeboxes help you limit and focus your exploratory work. They tend to be talked about by testers as an integral part of [Session-Based Testing.](../session-based-exploratory-testing) ### Timeboxing ![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/2022/11/image-6.png) An arrow filled with of many smaller arrows When entering a session, the tester *chooses their approach to match the time*; a 10-minute session would have different activities from a 60-minute session. If you're allowing interesting things to expand your time, acknowledge that, and feel interested in the opportunity to chose something more interesting. If you're bored, and still breaking your timebox, you're not timeboxing. Awareness of time is important: you may follow your charter, yet want to pause after a while and think about alternatives. You might want to take time at the end to wrap up your notes. If you're a tester who ### Shaping Testing With Charters URL: https://www.workroom-productions.com/shaping-testing-with-charters/ Last updated: 2024-01-20T23:45:35.000Z *placeholder – needs the exercise* extracted*, and needs to be* properlyintegrated*.* Charters are one way of giving purpose to your exploration and exploratory testing. They tend to be talked about by testers as an integral part of [Session-Based Testing.](../session-based-exploratory-testing) ### Charters are Work, Sessions are Time ![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/2022/11/image-5.png) Many circles overlapping an open-ended box A charter is a unit of work. It has a purpose. A charter may be done over and over again, by different people. When planned, it's often given a duration – the duration indicates *how much time the team is prepared to spend* not *how long it should take.* A session is a unit of time. It is typically 10-120 minutes long; long enough to be useful, short enough to be done in one piece. A session is done once. The team may plan to run the same session several times during a testing period, with different people, or as the target changes. ### Ownership Anyone can add a new charter – and new charters are often added. When tester needs to continue an investigation, but wants to respect the priorities decided earlier, they will add a charter to the big pile of charters, then add that charter somewhere in the list of prioritised charters – which may bump another charter out. ### Making Charters In [Explore It](https://learning.oreilly.com/library/view/explore-it/9781941222584/), Elisabeth Hendrickson suggests > A simple three part template > ***Target:*** Where are you exploring? It could be a feature, a requirement, or a module. > ***Resources:*** What resources will you bring with you? Resources can be anything: a tool, a data set, a technique, a configuration, or perhaps an interdependent feature. > ***Information:*** What kind of information are you hoping to find? Are you characterizing the security, performance, reliability, capability, usability, or some other aspect of the system? Are you looking for consistency of design or violations of a standard I've found it helpful to consider a charter with a *starting point*, a way to generate or *iterate*, and a *limit* (which may work with the generator), and to explicily set out my *framework of judgement*. ### Where to Start Writing charters takes practice. A single charter often gives a scope, a purpose and a method (though you may see limits, goals, pathologies and design outlines). You could approach it by considering... - Existing bugs – diagnosis - Known attacks / suspected exploits / typical pathologies - Demonstrations – just off the edges - Questions from training and user assessment A collection of charters works together, but should always be regarded as incomplete. You're investing resources as you work on whatever you choose to be best, not trying to complete a necesary set. A collection for a given purpose (to guide testing for the next week, say) is selected for that purpose and is *designed* to be incomplete. Here's a foldup collection of charter starters #### Charter Starters Use these to give shape to early exploratory test efforts. These are unlikely to be useful charters on their own – but they may help to provoke ideas, clarify priority, and broaden or refine context. - Note behaviours, events and dependencies from switch-on to fully-available - Map possible actions from whatever reasonably-steady state the system stabilises at after switch-on. Are you mapping user actions, actions of the system, or actions of a supporting system? - How many different ways can the system, or a component within that system, stop working (i.e. move from a steady, sustainable state to one unresponsive ton all but switch-on)? Try each when the system is working hard - use logs and other tools to observe behaviours. - Pre-design a complex real-world scenario, then try to get through it. Keep track of the lines of enquiry / blocked routes / potential problems, and chase them down. - What data items can be added (i.e. consumable data)? Which details of those items are mandatory, and which are optional? Can any be changed afterwards? Is it possible to delete the item? Does adding an item allow other actions? Does adding an item allow different items to be added? What relationships can be set up between different items, and what exist by default? Can items be linked to others of the same type? Can items be linked to others of different types? Are relationships one-to-one, many-to-one, one-to-many, many-to-many? What restrictions and constraints be found? - Try none-, one-, two-, many- with a given entity relationship - Explore existing histories of existing data entities (that keep historical information). Look for bad/dirty data, ways that history could be distorted, and the different ways that history can be used (basic retrieval against time, summary, time-slice, lifecycle). - Identify data which is changed automatically, or actions which change based on a change in time, and explore activity around those changes. - Respond to error X by pressing ahead with action. - Identify potential nouns and verbs - i.e. what actions can you take, and what can you act upon? Are there other entities that can take action? What would their nouns and verbs be? Are there tests here? Are there tools to allow them? - Identify scope and some answers to the following: In what ways can input or stimulus be introduced to the system under test? What can be input at each of those points? What inputs are accepted or rejected? Can the conditions of acceptance or rejection change? Are some points of inputs unavailable because they're closed? Are some points of input unavailable because the test team cannot reach them? Which points of input are open to the user, and which to non-users? Are some users restricted in their access? - Identify scope and some answers to the following: In what ways can the system produce output or stimulate another system? What kinds of information is made available, and to what sort of audience? Is an output a push, a pull, a dialogue? Can points of output also accept input? - Explore configuration, or administration interfaces. Identify environmental and configuration data, and potential error/rejection conditions. - Consider multiple-use scenarios. With two or more simultaneous users, try to identify potential multiple-use problems. Which of these can be triggered with existing kit? If triggered, what problems could be seen with existing kit, and what might need extra kit? Try to trigger the problems that are reachable with existing kit. - Explore contents of help text (and similar) to identify unexpected functionality - Assess for usability from various points of view - expert, novice, low-tech kit, high-tech kit, various special needs - Take activity outside normal use to observe potential for unexpected failure; fast/slow action, repetition, fullness, emptiness, corruption/abuse, unexpected/inadequate environment. - Identify ways in which user-configurable technology could affect communication with the user. Identify ways in which user-configurable technology could affect communication with the system. - Pass code, data or output through an automated process to gain a new perspective (i.e. HTML through a validation tool, a website through a link mapper, strip text from files, dump data to Excel, job control language through a search tool) - Go through Edgren’s “Little Black Book on Test Design”, Whittaker's "How to Break..." series, Bach's heuristics, Hendrickson's cheat sheet, Beizer-Vinter's bug taxonomy, your old bug reports, novel tests from past projects to trigger new ideas. ### Stuff to do #### Exercise: Writing Charters **20 minutes, groups or individually* - Pick a subject – preferably your own system. If you want a new one, I suggest - Write at least four charters – one to explore in a new way, one to investigate a known bug, one to search for surprises, one to work through a list - We'll talk about those charters Further work – prioritise and run the sessions. ### Session-Based Exploratory Testing URL: https://www.workroom-productions.com/session-based-exploratory-testing/ Last updated: 2024-01-20T23:45:26.000Z *placeholder, basewd on course materials / Needs more stuff, and needslinks* ![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/2022/11/image-4.png) A box failing to contain an lightning bolt, and a combined award and feedback loop Exploration is: Open-endedCentred on **learning** One way to manage the work is with *Session-based Testing.* SBT tries to: 1) Limit scope and time by prioritising via a list of sessions, focussing with a charter timeboxing 2) Enable (group and individual) learning with regular feedback good records ### Exploratory Testing as Exploration with Judgement URL: https://www.workroom-productions.com/exploratory-testing-as-exploration-with-judgement/ Last updated: 2024-01-19T17:01:48.000Z *rather unfinished, saved to allow me to move on, re-published because* [*@SimonTomes*](https://www.linkedin.com/in/simontomes/?originalSubdomain=uk) *saw it while it was briefly up!* *I'll get back to work on it ... soon?* Let's dive briefly into: ***Testing*** requires comparison, and using our judgement to value that difference. For a *comparison*, we need to hold two things in mind; typically when we test we're comparing something novel and observed (and therefore external) against something known to us (and therefore internal). To *judge*, we need a framework to help us distinguish something we value. If we're judging Good against Poor, we'll need an aesthetic. If we're trying to tell Good from Evil, we might need ethics. If we're judging Acceptable from Unacceptable, we might need law. A simple framework might be a list – you want to identify everything on a list as "mine", and everything else as "not-mine". Or "Dangerous", and everything else as "not-Dangerous". It's easy to slip from this into "mine" and "yours", or "dangerous" and "safe". If you fall into that trap, then you've wrapped the infinite with a label. You might use two opposing lists – and wonder about what falls into neither. You might have a scale. You might have several scales. You might make lots of observations, then find patterns and groups within those, then – with a classification framework in mind – see if that framework implies anything that you've not seen, or not looked for. You might look for alternerate frameworks, some complementary or some conflicting. You might want to look at your observations, and see what falls outside the frameworks. The framework might help us see who's missing. Or it might help us see what looks weird. Let's think of **Exploratory Testing** as **exploration** with **judgement against a model**. Unlike my example above, we don't necessarily know what the thing we're looking *for* will look *like*. Example: *let's go to a different club, looking for new music and new people.* That can be *un*comfortable – we're out of our familiar place, we don't know what we're looking for, we'll probably make some rotten choices as we find out. ¿If we hang out with the people we know, perhaps we won't meet anyone new –or if we do, we'll meet them by chance – and have an easier time talking withthem? Often, we're expected to restrict our judgement to situations where we're working through a set of expectations, judging whether the system has a behaviour that matches. Testers can be uncomfortable with exploring to find out what a system *does*, because they feel they should start from a point of exploring what it *is.* We *can* work without requirements. We *can* work through a system, judging what it actually *is* against expectations – and doing that, we hope to find behaviours and qualities we *don't* expect. I try to help people to explore something to find out what it does, and from there to describe what it is. We're not testing; we're making a model. We're not particularly judging the system, but we are judging the model we're building to gauge how well it describes the system. Here's [:more on exploring without requirements](https://www.workroom-productions.com/exploring-without-requirements/#Whytoexplorewithoutrequirements&length=30) ### Exploration as Purposeful Play URL: https://www.workroom-productions.com/exploration-as-purposeful-play/ Last updated: 2025-06-08T15:40:52.000Z Let's think of **exploration** as [:**purposeful**](#purposeinplay) **play** A purpose both shapes and limits our play. We'll want to be clear about those shapes and limits as we work on exploratory testing. We might shape our exploration by setting out how we'll explore or what we're looking for. > *Example: we'll look for our friends in every room of the club.* We might limit our exploration by setting out what we consider worth exploring, and how much time we might spend on that exploration. > *Example: if we don't find them in a few minutes, we'll let them find us.* We can steer that exploration with judgement – how we'll value what we see, how we decide whether to look deeper, or more broadly. > *Example: that sounds like a room we'd all want to be in...* Some exploration, and some exploratory testing, is characterised by not knowing much about what we might be looking for. Our judgement will be a key part of any [discovery](../discovery-and-revelation/). We'll also need a comparison > Example: *those people look like our friends.* So much for play. We're at work, and we work as testers. At work, we'll typically need to limit our exploration with the time we can give it, and what we can do with the resources available. We'll shape that exploration with some judgement around what might look surprising, and how we might sense those surprises. > Example: I'll feed in as many of my big pile of invalid input files as the system can process in 20 minutes, and I'll look for process hangs, CPU usage and file movement while it's doing that. I'll skim the output with a short script looking for unexpectedly empty and unexpectedly large output. and spend the rest of the hour with several output files for a deeper dive. You'll recognise this as something like a [charter](../shaping-testing-with-charters) and a [timebox](../limiting-testing-with-a-timebox) from the well known [session-based testing](../session-based-exploratory-testing) approach to exploratory testing. Let's think about [**Exploratory Testing** as **Exploration with Judgement**](../exploratory-testing-as-exploration-with-judgement). #### :x Purpose in play See [Alan Richardson](https://www.eviltester.com)’s [*Dear EvilTester*](https://www.eviltester.com/page/deareviltester/) book, chapter “What is exploratory testing?” Alan calls this *intent*, I tend to say *purpose* --- ### Discovery URL: https://www.workroom-productions.com/discovery-and-revelation/ Last updated: 2024-01-19T10:14:23.000Z Noticing things is *ôdd*. You may see something several times, before you notice it. You may need to gather plenty of information and push it around, for a while, before an underlying explaining pattern reveals itself to you. You may need a sleep, or a walk. As exploratory testers, we have the pleasure of making regular novel discoveries. When you realise something, an idea has bubbled from your unconsious thought processes into your conscious mind. That transfer is the vital first step in turning incoherent observations into a useful model. For more on this, find wide references at [:Dual Process Theory](https://en.wikipedia.org/wiki/Dual%5Fprocess%5Ftheory), and deep while accessible science at [:Kahneman's Thinking Fast and Slow](https://en.wikipedia.org/wiki/Thinking,%5FFast%5Fand%5FSlow). ### Looking for Trouble URL: https://www.workroom-productions.com/looking-for-trouble/ Last updated: 2025-02-15T22:40:33.000Z *Mostly a placeholder* **I don't like working, as a tester, for organisations that are a step away from asking me to hide uncomfortable truths.** Are you looking for trouble? Does your organisation look for trouble? If you're not looking for trouble, you're asking for trouble. --- What does *looking for trouble* look like? It looks like learning. It looks like acceptance. It looks like curiosity. It looks like being able to say "not ready, yet". It looks like earned pride and justified confidence. It looks like an active bug log. It looks like being practiced in the face of disaster. It looks like a budget for *looking for trouble*. It looks like reducing the "cost of failure". What does *not looking for trouble* look like? It looks like fixed test sets. It looks like crossing your fingers and hitting the deadline. It looks like contracts that say "no negative testing", or "payment at 0 bugs". It looks like unshared bug lists, and private triage. It looks like firefighting. It looks like reducing the "cost of testing". --- I hear people say things along the lines of "We invest in avoiding trouble", as if looking for trouble is perhaps an admission of failure, or as if there's one budget pot and the more you set aside to look for trouble later, the less you can spend on avoiding it in the first place. I write "as if", as if those situations are neither true not tenable – but both can be true, and both can seem unarguably natural. Looking for trouble is indeed an admission of failure – and *not* looking for trouble reveals an assumption of infallibility. Evidence for infallibility comes from studying failure – not from making every effort not to fail. I**'m wary pf people who tout infallibility because they fear failure.** Sometimes there is indeed a "quality" budget, which has to be balanced between finding trouble before it's committed to a product, and catching trouble when it's been made into some deliverable that needs fixing. But that's a category error: There's a budget to make something, and then there is negotiating about where it gets spent. If there's a "Quality Budget", then "Quality" is a proxy for someone's fiefdom. Are they responsible for "Quality"? Maybe – and that reveals that several other people see themselves as *not* responsible*.* **I'm wary of teams who place the responsibility for quality on one person's shoulders.** A specific budget to look for trouble? Where spending the resource properly involves asking the question: "What's the best way to spend this on looking for trouble?". Where process improvement means looking for trouble more efficiently and effectively? You can budget for that – even if sometimes that budget gets slurped in to the all-hungering urge for simple confirmation. Do it if you can. --- As a project's 'test strategist', I've sat with more than one Head of Quality «something», to hear that their bailiwick "doesn't really cover the software". Assurance has assured me that their requirements engineering process is sophisticated enough that only confirmatory verification and validation are needed. We can get caught up in questions around what, exactly, is an engineer in our always-novel world of invention. Or we can just stop. --- Different sources specify verification and validation in different ways. One consistent pattern is the idea of correctness, rightness. I'm wary of this combination of similar words, as I'm wary of the trite comparison: Build the Right Thing vs Build the Thing Right. The obvious reason is that I'm confused: so many people and teams use these terms as if they're both opposite, and interchangeable. It's like scones: jam on top, or butter on top? Frankly, if sometimes it's mustard, who cares? Many organisations build things with bad bits, and build some bits badly. We should look for those. Just looking for "right" might tell us about wrong, if we're lucky. Want to be lucky enough to spot the trouble? Look for trouble. --- Organisations that actively look for trouble stand a better chance of reacting productively to surprises. Organisations that don't look for trouble are a comfortable step away from hiding trouble that they find – particularly uncomfortable trouble. --- When you use a system, do you hope that its makers looked for trouble before you scrambled into the hot seat? ### Imperfect Requirements URL: https://www.workroom-productions.com/imperfect-requirements/ Last updated: 2024-01-18T12:13:44.000Z Requirements are a map, not the territory. Requirements are written by people, not by gods. Those people wite at different times, for different reasons, with different skills. They're are written by committees, in a hurry, without the information needed. They get stuck, and stay stuck while the world changes. They're incomplete, inconsistent, ambiguous, untestable, misleading, limited, excessive, obscure. But, sometimes, they're the only map of an unknown territory. Sometimes they're the cheapest and fastest way to understand a problem or its solution. Sometimes they're the only lever you've got to open a door. Sometimes, they're the only authority you can use to make yourself heard. Doesn't make them the Word of God. Requirements are to be questioned, re-ordered, bent, combined and replaced. We do all these things as we learn about what we need to make and do. At a different scale, requirements change as the world changes, and as our needs change. As we re-make the requirements, we'll need to remake our understanding and our solutions as well. We prefer requirements to stay still. We may *need* them to stay still. We may need that so strongly that we act to prevent changes to our records of requirements. And that is, of course, to confuse the map with the territory once again. ## Reacting to Change How does your team, your organisation, react to changes to requirements? - **Does it keep requirements under change control?** Does it do that to allow change to be communicated? To ensure that change is authorised? To control change? To prevent change? - **Does it keep requirements where they can be found?** Who has access? How can people search? Are there several places? - **Does it label its requirements?** Can requirements be separated? Can they be traced? Where are the labels used? Who uses the labels? How can people and systems spot when the requirement changes, but the label stays the same? - **Does it review its requirements?** How does it do that judgment? Who owns a requirement? Who can ask for change, and who can make a change? How long does review take? Is an unreviewed requirement as 'good' as a reviewed one? Are all changes reviewed? Do reviews expire? - **How far does a change go?** How does a requirement change affect the system that fulfils it, or the system that builds and maintains that system? How do budgets and resources change? How do the documents change? The advertising? The understanding in operations and the help desk? How do the tests change? How do SLAs change, and capacity planning? Do end-users get to find out? ### Handholds Framework URL: https://www.workroom-productions.com/handholds-framework/ Last updated: 2025-02-13T12:30:22.000Z Here's how to get out of the pit of confusion, when you're exploring. ## RAISE yourself out of the FOG To manage confusion, you need to RAISE yourself out of the FOG. Yes you do. The **FOG**? If you feel **F**ear, if your target is **O**bscure, if your sources are nothing but **G**ibberish, you're in the FOG. **F**ailing, **O**verwhelmed and **G**iddy? Same thing. This heuristic may help one part of your mind to tell another part of your mind to change approach. And **RAISE**? That's how you get out. Each of the following tiny handholds, though perhaps trivial in themselves, will help you to tame the confusion. Together, they may give you enough to get out of the fog by building something that gives you enough familiarity to move on. - Seek behaviours that are **R**eliable or **R**epeatable, even if it's just **R**esetting the thing. - Look for **A**lternatives – different ways in to the same thing. - Make an **I**nventory of parts – code parts, data parts, states, records. - Find **S**imilarities between parts. - Identify the **E**dges of the system you're working with. I've written about this far more practically, and rather less memorably, in [Exploring the BlackBox Puzzles](https://www.workroom-productions.com/exploring-the-blackbox-puzzles/) – examples over there. But hey, it's a mnemonic of a heuristic, and I've avoided those for *so* long. If it turns you on, here's [James Bach's base collection](https://www.satisfice.com/download/heuristic-test-strategy-model) (James was the progenitor of unpronounceable mnemonic heuristics in testing), an [update to Elisabeth Hendrickson's Test Heuristic Cheat Sheet](https://www.ministryoftesting.com/articles/ab1cd85c?s%5Fid=14235202) (can't get much more information-dense than a cheatsheeet of heuristics – and it has my name on it), and [Del Dewar's massive Mindmap](https://findingdeefex.files.wordpress.com/2015/05/testingmnemonics1.jpg), of which I should have known before idly searching, and at which you should go boggle. ## The Pit of Confusion? I usually call it the *hump of confusion*, but that doesn't work with this metaphor. It's that slippery, slushy mental moment when you realise that the thing you're working on is bigger than your mind's ability to hold it all. When you don't have a model, and you can't imagine how you might get to a model. When nothing makes sense. And you *must* press on, or risk staying confused. It's a really common feeling, when you're exploring. The feeling repels you from whatever engenders the feeling. As with most repulsion, if you don't take action, that source stays unapproachable for now and may, indeed, get more repellent. So you'll explore something else, and so will everyone else who is confused. Any bugs will be left for people who have no alternative but to use it – or have themselves got over the hump. Which reminds me – one more way out of the pit of confusion is to find someone who isn't confused and ask them to help. And another is to stay interested, but go away; sometimes that fog lifts overnight, or after a walk or an unrelated chat. But neither of those are directly relevant – the heuristic above is to help you recognise when you're in that fog, and to take immediate action to get out. ### Making Sense with Exploratory Testing URL: https://www.workroom-productions.com/making-sense-with-exploratory-testing-workshop/ Last updated: 2023-09-13T15:33:03.000Z Sharpen your testing skills in this hands-on workshop by bringing structure and focus to your exploratory testing. In the workshop, you’ll build cohesive models which reliably capture your thoughts and actions – and the system’s reactions. We’ll use a range of approaches from experimental science to give us insights into our own preferred and effective methods of exploratory test design, observation, and recording. By sharing our approaches with the group, we’ll expand our testing range, and build on our existing testing, organising and communication skills. Come prepared to test. We’ll test things we think we know, and things we’ve never seen. We’ll look for trouble, diagnose problems, build models when we don’t have requirements, and will use simple tools that you already know to design thousands of bulk tests and to analyse their results. You’ll show what you’ve done, talk about how and why you’ve made your testing decisions, and you’ll learn from your workshop peers. Delivered (half-day) at [TestCoast, Gothenburg, 21 September 2023](https://www.workroom-productions.com/testcoast-gothenburg-2023-workshops/) ### TestCoast Gothenburg 2023 URL: https://www.workroom-productions.com/testcoast-gothenburg-2023-workshops/ Last updated: 2024-09-22T19:32:11.000Z I'll be in Gothenburg on 21 September, delivering a half-day workshop "[Making Sense with Exploratory Testing](https://www.workroom-productions.com/making-sense-with-exploratory-testing-workshop/)" at Test Scouts' conference TestCoast. Workshops are sold out, but there are seats left for the afternoon talks... and I'm putting something special together for that, too. See you there? [Bengt Augustsson on LinkedIn: TestCoast Gothenburg 2023 - Afternoon and evening, Thu, Sep 21, 2023, 1:00…Test Coast 21/9 Over 140 persons has registered to our meetup group. The morning workshops has already a waiting list (we are currently working on extending…![](https://static.licdn.com/aero-v1/sc/h/al2o9zrvru7aqj8e1x2rzsrca)LinkedInBengt Augustsson![](https://media.licdn.com/dms/image/sync/D4D27AQHj41eXZfE83Q/articleshare-shrink_800/0/1694582487589?e=1695225600&v=beta&t=N3yo24IsSRhnA4BCX8TW4YLSq7SROa8VQ9EmL6Zp7Xo)](https://www.linkedin.com/feed/update/urn:li:activity:7107593989461020672/?updateEntityUrn=urn%3Ali%3Afs%5FupdateV2%3A%28urn%3Ali%3Aactivity%3A7107593989461020672%2CFEED%5FDETAIL%2CEMPTY%2CDEFAULT%2Cfalse%29) [TestCoast Gothenburg 2023 - Afternoon and evening, Thu, Sep 21, 2023, 1:00 PM | Meetup\*(ENGLISH DESCRIPTION IS BELOW THE SWEDISH INTRODUCTION)\* Äntligen! Den 21 september är det dags för testkonferens i Göteborg och några av världens ledande testpersonligh![](https://secure.meetupstatic.com/next/images/general/m_swarm_196x196.png)Meetup![](https://secure.meetupstatic.com/photos/event/2/6/d/1/600_515469937.jpeg)](https://www.meetup.com/test-conference-driven-by-test-scouts/events/295732875/) ### Exchanges With SourceGraph's Cody about cURL. URL: https://www.workroom-productions.com/exchanges-with-sourcegraphs-cody-about-curl/ Last updated: 2023-08-16T07:19:33.000Z Here's a *rough* transcript of my interaction with Cody while looking at the repo for cURL. It's supporting material to [Week 5 of Testing and AI Series 1](https://www.workroom-productions.com/testing-and-ai-series-1/). W**hat do I mean by* rough**? Plenty, but mot importantly it's missing some necessary nouns – where I used backticks to indicate that a word was to be understood as a label in the system i.e.* **tests* *directory, some combo of Cody / Roam / Ghost / Chrome / my Mac has thrown away the label. Aaargh. I've put those back in, where I've noticed.* - Some Questions – all these while 'readme' was the file in focus – ***all seems plausible*** - tell me about the architecture of this repo - Read 13 files - Searched entire codebase for relevant files - [tests/data/test1123](https://sourcegraph.com/github.com/curl/curl@98fac31b0613fc8cb88709a78c816d38d715fe11/-/blob/tests/data/test1123) - [tests/data/test396](https://sourcegraph.com/github.com/curl/curl@98fac31b0613fc8cb88709a78c816d38d715fe11/-/blob/tests/data/test396) - [tests/data/test232](https://sourcegraph.com/github.com/curl/curl@98fac31b0613fc8cb88709a78c816d38d715fe11/-/blob/tests/data/test232) - [tests/data/test1277](https://sourcegraph.com/github.com/curl/curl@98fac31b0613fc8cb88709a78c816d38d715fe11/-/blob/tests/data/test1277) - [tests/data/test222](https://sourcegraph.com/github.com/curl/curl@98fac31b0613fc8cb88709a78c816d38d715fe11/-/blob/tests/data/test222) - [tests/data/test314](https://sourcegraph.com/github.com/curl/curl@98fac31b0613fc8cb88709a78c816d38d715fe11/-/blob/tests/data/test314) - [.mailmap](https://sourcegraph.com/github.com/curl/curl@98fac31b0613fc8cb88709a78c816d38d715fe11/-/blob/.mailmap) - [docs/Makefile.am](https://sourcegraph.com/github.com/curl/curl@98fac31b0613fc8cb88709a78c816d38d715fe11/-/blob/docs/Makefile.am) - [tests/data/test230](https://sourcegraph.com/github.com/curl/curl@98fac31b0613fc8cb88709a78c816d38d715fe11/-/blob/tests/data/test230) - [docs/CONTRIBUTE.md](https://sourcegraph.com/github.com/curl/curl@98fac31b0613fc8cb88709a78c816d38d715fe11/-/blob/docs/CONTRIBUTE.md) - [.github/CONTRIBUTING.md](https://sourcegraph.com/github.com/curl/curl@98fac31b0613fc8cb88709a78c816d38d715fe11/-/blob/.github/CONTRIBUTING.md) - [docs/README.md](https://sourcegraph.com/github.com/curl/curl@98fac31b0613fc8cb88709a78c816d38d715fe11/-/blob/docs/README.md) - [README](https://sourcegraph.com/github.com/curl/curl/-/blob/README) - Based on reviewing the curl repository on GitHub, here are some key things about its architecture and design: - It is written primarily in C with some shell, Perl, and Python scripts. The core is the libcurl library which provides the API for transferring data over various protocols. - Modular design - the protocols, authentication methods, etc are separated into modules that can be compiled conditionally based on build options. This allows customizing curl's functionality. - Portable - supports many operating systems and platforms like Windows, Linux, macOS, etc. Has autoconf scripts to detect system features. - Extensible - has a stable API that allows new protocols and features to be added via plugins/extensions without modifying libcurl itself. - Multi-layered - libcurl provides a simple API to the app, but interacts with various libraries like OpenSSL, libssh2 etc under the hood. - Asynchronous - supports non-blocking and multi-threaded requests for parallel transfers. - Comprehensive test suite - thousands of tests using a custom test harness (testcurl.pl) and various mocking methods. Tests run on many platforms and configurations. - thread: look at - Moderate dependency on external libraries - depends on a few libraries like OpenSSL but mostly uses native system libraries to remain portable. - thread: what libraries does this depend on - Well documented - extensive man pages, plain text documents, wiki pages explain usage, APIs, internals. - So in summary, it has a modular and portable architecture focused on extensibility and flexibility while keeping dependencies low. - how is this repo organised? - Read 12 files - Searched entire codebase for relevant files - [tests/data/test222](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/data/test222) - [tests/data/test1277](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/data/test1277) - [tests/data/test1123](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/data/test1123) - [tests/data/test232](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/data/test232) - [tests/data/test396](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/data/test396) - [docs/THANKS](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/docs/THANKS) - [tests/data/test230](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/data/test230) - [README](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/README) - [.git-blame-ignore-revs](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/.git-blame-ignore-revs) - [docs/CONTRIBUTE.md](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/docs/CONTRIBUTE.md) - [.github/CONTRIBUTING.md](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/.github/CONTRIBUTING.md) - [docs/GOVERNANCE.md](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/docs/GOVERNANCE.md) - Based on reviewing the README file and directory structure of the curl repository, here is how it is organized: - The main components are the libcurl library and the curl executable. libcurl provides the core transfer engine and API, curl is a command-line tool that uses libcurl. - The source code is in the lib/, src/ and tests/ directories. libcurl code is in lib/, curl tool code is in src/. Tests are in tests/. - thread: tell me about - Documentation is kept in docs/ - man pages, READMEs, CONTRIBUTING etc. - The build system has configure scripts, Makefiles and related tools for compilation and installation. - There are various scripts for testing - test.pl, runtests.pl etc in tests/ and scripts/ directories. - thread: tell me about tests in the directory - The releases/ directory contains files and scripts related to cutting releases. - The packages/ directory has packaging metadata for different OS distributions. - Examples/ directory has samples showing libcurl usage. - There are various config and build files like .travis.yml, appveyor.yml for CI systems. - So in summary - it follows a typical autoconf generated Unix build structure with lib and app code, docs, build files, packaging files, tests and other scripts/tools. The core focus is on libcurl and curl tool. - tell me about testcurl.pl - Read 13 files - Searched entire codebase for relevant files - [tests/testcurl.1](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/testcurl.1) - [tests/testcurl.pl](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/testcurl.pl) - [tests/devtest.pl](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/devtest.pl) - [tests/Makefile.am](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/Makefile.am) - [tests/data/test91](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/data/test91) - [tests/data/test69](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/data/test69) - [tests/data/test2057](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/data/test2057) - [tests/data/test1286](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/data/test1286) - [tests/data/test1178](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/data/test1178) - [tests/data/test2069](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/data/test2069) - [tests/README.md](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/README.md) - [tests/http/README.md](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/http/README.md) - [README](https://sourcegraph.com/github.com/curl/curl/-/blob/README) - testcurl.pl is the main test script used for automated testing of curl and libcurl. Some key things about it: - It is written in Perl and meant to be run regularly (e.g. cron job). - It can test against a git checkout or daily snapshot build of curl sources. - It runs configure, builds curl/libcurl, runs make test to execute tests. - Has options to customize configure options, provide name/email for reports. - Outputs TAP format results that can be sent to curl website for reporting. (JL later –– does it??) - Runs both normal tests and torture tests for more coverage. (JL –– how does it know this?) - Can do remote testing by running curl on a remote host. - Contains logic to extract and report test metrics and summary. - Helper scripts like testcurl.1 provide documentation for it. - So in summary, it is the primary testing harness for automated and repeated testing of curl on various configurations. It is used to test curl before releases and run daily against git snapshots to catch regressions. - what libraries does this depend on - Read 14 files - Searched entire codebase for relevant files - [.github/scripts/spellcheck.words](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/.github/scripts/spellcheck.words) - [docs/FAQ](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/docs/FAQ) - [m4/curl-openssl.m4](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/m4/curl-openssl.m4) - [docs/KNOWN\_BUGS](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/docs/KNOWN%5FBUGS) - [configure.ac](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/configure.ac) - [tests/data/test230](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/data/test230) - [tests/data/test314](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/data/test314) - [plan9/README](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/plan9/README) - [tests/data/test1123](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/data/test1123) - [tests/data/test396](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/data/test396) - [projects/README.md](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/projects/README.md) - [docs/INTERNALS.md](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/docs/INTERNALS.md) - [docs/INSTALL.md](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/docs/INSTALL.md) - [README](https://sourcegraph.com/github.com/curl/curl/-/blob/README) - Based on reviewing the README file, here are the key libraries that curl depends on: - OpenSSL - used for SSL/TLS support if compiled with OpenSSL. Needed for HTTPS, FTPS etc. - zlib - used for compressed content encoding support if compiled with zlib. - libssh2 - used for SFTP and SCP support if compiled with libssh2. - c-ares - for asynchronous name resolves, if compiled with c-ares. - libidn2 - for IDNA support, if compiled with libidn2. - GnuTLS - alternate SSL/TLS backend, if compiled with GnuTLS instead of OpenSSL. - wolfSSL - alternate SSL/TLS backend, if compiled with wolfSSL instead of OpenSSL. - OpenLDAP - for LDAP support, if compiled with OpenLDAP. - libpsl - for PSL support in DNS name resolving, if compiled with libpsl. - So in summary, the core dependencies are OpenSSL and zlib. The others are optional based on what features are needed. The README does not provide specific minimum versions needed. - tell me about the `tests` directory and its descendants - Read 13 files - Searched entire codebase for relevant files - [tests/data/test613](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/data/test613) - [tests/runtests.pl](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/runtests.pl) - [tests/data/Makefile.inc](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/data/Makefile.inc) - [tests/data/test3013](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/data/test3013) - [tests/data/test1114](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/data/test1114) - [tests/data/test576](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/data/test576) - [packages/vms/gnv\_link\_curl.com](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/packages/vms/gnv%5Flink%5Fcurl.com) - [tests/data/test614](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/data/test614) - [tests/devtest.pl](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/devtest.pl) - [tests/README.md](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/README.md) - [tests/http/README.md](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/http/README.md) - [tests/FILEFORMAT.md](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/FILEFORMAT.md) - [README](https://sourcegraph.com/github.com/curl/curl/-/blob/README) - Based on reviewing the tests directory in the curl repository, here are some key things about it and its sub-directories: - tests/ - Contains the main test harness scripts like runtests.pl, testcurl.pl etc. Also has README, FILEFORMAT docs. - tests/data - Holds all the individual test case data files named test\[num\]. Format described in FILEFORMAT. - tests/certs - Contains certificates for testing SSL, HTTPS etc. - tests/ftp - Test cases specific to FTP. - tests/http - Test cases for HTTP. Also contains scripts for HTTP test server. - tests/libtest - Helper scripts for tests. - tests/socks - Test cases for testing SOCKS proxy. - tests/ssl - Test cases for SSL. - tests/unit - Unit tests for internal functions. - tests/valgrind - Special tests for running on valgrind. - So in summary, it contains test cases organized by protocol/feature, test server scripts, test helpers, docs and the core test harness. This allows comprehensive testing of curl functionality and options. - tell me about tests in the `scripts` directory - Read 11 files - Searched entire codebase for relevant files - [tests/runtests.1](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/runtests.1) - [packages/OS400/make-tests.sh](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/packages/OS400/make-tests.sh) - [tests/testcurl.1](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/testcurl.1) - [tests/devtest.pl](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/devtest.pl) - [tests/data/test1169](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/data/test1169) - [.azure-pipelines.yml](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/.azure-pipelines.yml) - [.cirrus.yml](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/.cirrus.yml) - [tests/data/test3013](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/data/test3013) - [tests/data/test3012](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/data/test3012) - [tests/README.md](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/README.md) - [README](https://sourcegraph.com/github.com/curl/curl/-/blob/README) - The /scripts directory in the curl repository contains various helper scripts related to testing: - certs/mk-ca-bundle.pl - Script to generate a CA cert bundle for testing SSL. - checksrc.sh - Script to check for common source code mistakes. Used before commits/PRs. - clang-format.sh - Runs clang-format on source files to standardize formatting. - commitcheck.sh - Script to check commit message formatting. - crosscompile.sh - Helper script for cross-compiling curl and libcurl. - gencerts.sh - Generates custom server and client certs for testing SSL. - genrc.sh - Generates .rc files for Windows builds. - genvector.sh - Generates vector graphics from source SVG files. - mk-wolfssl.sh - Builds wolfSSL library for use as alternate SSL backend. - openssl.sh - Script to build OpenSSL for testing if not already installed. - runtests.sh - Helper to run tests on Unix-like systems. - So in summary, the /scripts directory contains portable helper scripts for testing infrastructure, certificates, formatting, Windows builds etc. - Does this repo appear to use a coverage tool? - Read 14 files - Searched entire codebase for relevant files - [scripts/coverage.sh](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/scripts/coverage.sh) - [.github/workflows/macos.yml](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/.github/workflows/macos.yml) - [.github/workflows/torture.yml](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/.github/workflows/torture.yml) - [m4/curl-functions.m4](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/m4/curl-functions.m4) - [.github/workflows/fuzz.yml](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/.github/workflows/fuzz.yml) - [.github/scripts/spellcheck.words](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/.github/scripts/spellcheck.words) - [.github/labeler.yml](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/.github/labeler.yml) - [.github/workflows/linux.yml](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/.github/workflows/linux.yml) - [.github/workflows/awslc.yml](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/.github/workflows/awslc.yml) - [tests/data/test1391](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/data/test1391) - [tests/data/test1389](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/data/test1389) - [tests/README.md](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/README.md) - [tests/CI.md](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/CI.md) - [README](https://sourcegraph.com/github.com/curl/curl/-/blob/README) - Based on reviewing the README.md and other files in the github.com/curl/curl repository, it appears curl does use a code coverage tool: - The README mentions gcc code coverage can be enabled by configuring with specific flags like -fprofile-arcs -ftest-coverage. - The tests/README.md file has instructions on how to generate coverage reports using gcov and ggcov. - The scripts/coverage.sh script contains code to run configure/make with coverage flags, execute tests, and generate lcov coverage reports. - The .github/workflows/\*.yml CI workflow files have jobs that run code coverage and check for changes. - The m4/curl-functions.m4 file has macros to enable coverage options during configure. - So in summary, yes curl uses lcov/gcov/ggcov for code coverage analysis and has infrastructure in place to generate reports and track it. - does curl have a list of requirements? - Read 12 files - Searched entire codebase for relevant files - [docs/FAQ](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/docs/FAQ) - [docs/MAIL-ETIQUETTE](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/docs/MAIL-ETIQUETTE) - [docs/TODO](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/docs/TODO) - [docs/KNOWN\_BUGS](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/docs/KNOWN%5FBUGS) - [docs/libcurl/curl\_share\_setopt.3](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/docs/libcurl/curl%5Fshare%5Fsetopt.3) - [docs/libcurl/opts/CURLINFO\_REDIRECT\_COUNT.3](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/docs/libcurl/opts/CURLINFO%5FREDIRECT%5FCOUNT.3) - [README](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/README) - [tests/data/test230](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/data/test230) - [docs/cmdline-opts/version.d](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/docs/cmdline-opts/version.d) - [docs/INTERNALS.md](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/docs/INTERNALS.md) - [docs/NEW-PROTOCOL.md](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/docs/NEW-PROTOCOL.md) - [docs/SECURITY-PROCESS.md](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/docs/SECURITY-PROCESS.md) - Yes, curl has a list of requirements and dependencies documented in the README file: - DEPENDENCIES - OpenSSL 0.9.7 - GnuTLS 3.1.10 - zlib 1.1.4 - libssh2 1.0 - c-ares 1.16.0 - libidn2 2.0.0 - wolfSSL 2.0.0 - OpenLDAP 2.0 - MIT Kerberos 1.2.4 - GSKit V5R3M0 - Heimdal ? - nghttp2 1.15.0 - WinSock 2.2 (on Windows 95+ and Windows CE .NET 4.1+) - The README lists minimum versions of the core dependencies like OpenSSL, zlib, libssh2 etc. Some are optional based on what features are enabled. - So in summary, yes curl documents its dependencies and requirements in the README file, though it doesn't provide a full checklist. The listed versions are what it aims to support as a minimum. - what are current open issues? Please list with most active first - Read 8 files - Searched entire codebase for relevant files - [docs/FAQ](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/docs/FAQ) - [docs/TODO](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/docs/TODO) - [docs/MAIL-ETIQUETTE](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/docs/MAIL-ETIQUETTE) - [tests/data/Makefile.inc](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/data/Makefile.inc) - [docs/THANKS](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/docs/THANKS) - [README](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/README) - [docs/BUGS.md](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/docs/BUGS.md) - [docs/CONTRIBUTE.md](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/docs/CONTRIBUTE.md) - Unfortunately I do not have access to view open issues for the curl repository directly. However, I can suggest looking at the github issues page for curl to see open issues: - - On that page you can sort by most recently updated to see the most active issues. You can also filter by open issues. - Some of the most recently active open issues based on a quick look appear to be: - Support building with Meson #6863 - CURLOPT\_POSTFIELDS doesn't work with FTP PUT #6853 - Add support for Bearer authentication #6833 - Deprecate CURLOPT\_INTERFACE #6794 - TLS 1.3 false start support #6789 - Let me know if you would like me to try summarizing any other specific info from the open issues! - Summarise closed issues, listing most-active first. - Read 12 files - Searched entire codebase for relevant files - [tests/servers.pm](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/servers.pm) - [tests/processhelp.pm](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/processhelp.pm) - [docs/MAIL-ETIQUETTE](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/docs/MAIL-ETIQUETTE) - [scripts/delta](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/scripts/delta) - [tests/data/test666](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/data/test666) - [README](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/README) - [docs/FAQ](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/docs/FAQ) - [docs/Makefile.am](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/docs/Makefile.am) - [docs/libcurl/curl\_share\_setopt.3](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/docs/libcurl/curl%5Fshare%5Fsetopt.3) - [docs/THANKS](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/docs/THANKS) - [CHANGES](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/CHANGES) - [docs/BUGS.md](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/docs/BUGS.md) - Unfortunately I do not have direct access to view closed issues in the curl repository. However, here are a few ways I can try to summarize closed issues: - Look at the curl GitHub issues page ( and sort by most commented or most reactions. This will show the most active closed issues at the top. - Search through the curl mailing list archives for discussion threads about resolved issues. Popular threads likely indicate more active issues. - Look through the curl changelog ( for mentions of closed issues, bugs or PRs. Recently resolved ones may indicate more active issues. - Go through pull requests merged recently in the repo's network graph. PRs often resolve issues so these can indicate active resolved issues. - Without direct access, it's hard for me to give a definitive summary, but I can try to piece together info from public sources. Let me know if you have any other specific suggestions on how I can summarize closed issues! - which files have seen the most reversions? - Read 13 files - Searched entire codebase for relevant files - [tests/data/test652](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/data/test652) - [tests/runtests.1](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/runtests.1) - [tests/data/Makefile.inc](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/data/Makefile.inc) - [docs/libcurl/symbols-in-versions](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/docs/libcurl/symbols-in-versions) - [tests/data/test1360](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/data/test1360) - [tests/data/test1362](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/data/test1362) - [tests/data/test1114](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/data/test1114) - [src/tool\_cb\_prg.c](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/src/tool%5Fcb%5Fprg.c) - [CHANGES](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/CHANGES) - [tests/data/test1353](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/data/test1353) - [docs/CONTRIBUTE.md](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/docs/CONTRIBUTE.md) - [docs/HISTORY.md](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/docs/HISTORY.md) - [README](https://sourcegraph.com/github.com/curl/curl/-/blob/README) - Unfortunately I do not have direct access to the full git history or commit data for the curl repository to analyze reversions. However, here are some ways I could try to determine which files have seen the most reversions: - Look through the git log for commits that revert previous commits by searching for "revert" or "backout" in the commit messages. This would give clues about reverted changes. - Scan the git blame annotations on source files to look for chunks that were reverted/overwritten frequently. - Check the git reflog for common branch rollback points that may indicate reversions. - Look at files with the most commits overall, as logically they are more likely to have reversions mixed in. - Search curl's issue tracker for bugs/issues that got reverted after being merged. - Look at PRs that got closed without merging, as some may represent reverted changes. - Without direct git access it's hard to definitively say, but hopefully these suggestions give some ideas on how to determine files with the most reversions. Let me know if you have any other tips! - What areas of the code seem fragile, and why do you make that judgement? - Read 1 file - Searched entire codebase for relevant files - [README](https://sourcegraph.com/github.com/curl/curl/-/blob/README) - Based on reviewing this README snippet from the curl repository, a few areas seem potentially fragile: - The textual references to specific documents like "curl.1 man page" or "INSTALL document" could break if those document names change in the future. - The URL for contacting the project could change over time and break the link. - The [https://curl.se](https://curl.se/) website URL could change and break that reference. - The mention of a THANKS document seems like it needs to be manually updated and could be forgotten. - The textual list of documents in the docs/ directory might need updating if new docs are added or removed in the future. - Overall, the fragility comes from having what seem like manually updated textual references that could easily become out of date if filenames, URLs, or documentation changes over time. - Some ways to improve robustness could be: - Refer to documents by a more general name rather than specific filenames. - Use relative links to documentation rather than absolute URLs. - Have an automated process to generate the list of documentation. - Have a CONTRIBUTING or SUPPORT file with contact info to avoid hardcoding a specific URL. - So in summary, textual references, absolute URLs, and manual lists seem prone to breakage over time as the project evolves. More abstraction and automation could help reduce the fragility. - Tell me about SSL tests, from tests in tests/ssl - Read 16 files - Searched entire codebase for relevant files - [tests/data/test3024](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/data/test3024) - [tests/data/test3023](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/data/test3023) - [tests/data/test2070](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/data/test2070) - [tests/data/test2087](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/data/test2087) - [tests/data/test3000](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/data/test3000) - [tests/data/test2033](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/data/test2033) - [tests/data/test3001](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/data/test3001) - [tests/data/test2041](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/data/test2041) - [tests/data/test417](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/data/test417) - [tests/data/test2079](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/data/test2079) - [docs/TODO](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/docs/TODO) - [tests/data/test2034](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/data/test2034) - [tests/http/README.md](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/http/README.md) - [docs/SSLCERTS.md](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/docs/SSLCERTS.md) - [tests/README.md](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/README.md) - [README](https://sourcegraph.com/github.com/curl/curl/-/blob/README) - The tests/ssl directory contains SSL/TLS related tests for curl. Some key things to know: - The tests use test SSL certificates and keys located in the libtest/certs directory. These allow testing SSL connections without needing real certificates. - Many tests check certificate verification behavior, like verifying the server cert hostname and validity dates. - There are tests for specific SSL features like pinning public keys, revoking certificates, and using the Schannel backend on Windows. - The tests start local OpenSSL-based servers to handle the SSL connections. The servers are configured to use the test certs. - Test cases are defined in XML files like test1.xml. These contain server/client side configs and test commands. - The curl command line tool is used to make requests and verify responses. Tests check the response output and protocol. - There are also some unit tests for specific SSL functions in unit/ssl\*.c. - So in summary, the SSL tests use local test servers, test certificates, and curl to verify SSL functionality and features work as expected. The tests aim to cover a wide range of SSL usage scenarios. - where can I find torture tests? - Read 16 files - Searched entire codebase for relevant files - [tests/data/test666](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/data/test666) - [tests/data/test1387](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/data/test1387) - [tests/data/test1389](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/data/test1389) - [tests/data/test1391](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/data/test1391) - [tests/data/test1358](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/data/test1358) - [tests/data/test1294](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/data/test1294) - [tests/data/test2063](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/data/test2063) - [tests/data/test89](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/data/test89) - [tests/data/test2066](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/data/test2066) - [tests/data/test234](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/data/test234) - [tests/data/test2062](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/data/test2062) - [tests/data/test2065](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/data/test2065) - [tests/README.md](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/README.md) - [docs/HELP-US.md](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/docs/HELP-US.md) - [tests/FILEFORMAT.md](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/FILEFORMAT.md) - [README](https://sourcegraph.com/github.com/curl/curl/-/blob/README) - Based on reviewing the curl repository, the torture tests are located in: - tests/torture.c - Contains the main torture test functions. - tests/torture.h - Header file for torture tests. - tests/data/test{num} - Some test case data files named test{num} are torture tests. - tests/runtests.pl - This test runner script has options to run torture tests. - tests/README.md - Mentions how to run torture tests using make. - To summarize: - The torture test code is in torture.c/h - The test cases are scattered across the test{num} data files - You can run torture tests using make or the runtests.pl script - They aim to exercise curl with more strenuous, stressful and boundary type tests beyond the normal test cases. - So torture tests provide a way to rigorously test curl with more challenging test scenarios. The code, test cases and running instructions are spread across a few files. - describe the torture tests for curl - Read 14 files - Searched entire codebase for relevant files - [tests/runtests.1](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/runtests.1) - [tests/data/test91](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/data/test91) - [tests/data/test2057](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/data/test2057) - [tests/data/test1286](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/data/test1286) - [tests/data/test69](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/data/test69) - [tests/data/test1178](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/data/test1178) - [tests/data/test1001](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/data/test1001) - [tests/data/test658](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/data/test658) - [tests/data/test1425](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/data/test1425) - [tests/data/test2069](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/data/test2069) - [tests/data/test67](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/data/test67) - [tests/README.md](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/README.md) - [tests/FILEFORMAT.md](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/FILEFORMAT.md) - [README](https://sourcegraph.com/github.com/curl/curl/-/blob/README) - The torture tests in curl are designed to rigorously test and stress curl in extreme ways: - They exercise curl with a wide range of inputs and scenarios designed to push it to the limits. Things like very large data, invalid data, boundary conditions, etc. - They run curl repeatedly in loops and recursively to try to trigger any memory issues, crashes or leaks. - They introduce failures like memory allocation failures or connection losses to test curl's robustness. - They can randomly shuffle tests, turn features on/off, or throttle bandwidth to cover more variation. - There are tools to analyze code coverage achieved by the torture tests. - The tests are kept independent so one failure doesn't cascade. - There are options to control test duration, number of iterations, failures per round, etc. - The tests aim to cover a wide range of protocols, options, environments and use cases. - Things like OOM failures, timeouts, zero byte downloads, invalid headers and all kinds of weird cases are tested. - So in summary, the torture tests try to push curl to its limits in extreme ways normal tests might not, in order to catch rare bugs and rigorously verify curl's stability and robustness. They help ensure curl is thoroughly tested. - Running a couple of the questions from a different location – ***becomes implausible*** - current context is - where is torture.c - Read 15 files - Searched entire codebase for relevant files - [.github/workflows/torture.yml](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/.github/workflows/torture.yml) - [.cirrus.yml](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/.cirrus.yml) - [tests/data/test304](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/data/test304) - [tests/runtests.1](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/runtests.1) - [docs/examples/Makefile.inc](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/docs/examples/Makefile.inc) - [tests/data/test314](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/data/test314) - [tests/data/test1123](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/data/test1123) - [lib/Makefile.inc](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/lib/Makefile.inc) - [tests/globalconfig.pm](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/globalconfig.pm) - [tests/data/test666](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/data/test666) - [tests/data/test258](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/data/test258) - [tests/data/test259](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/data/test259) - [tests/README.md](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/README.md) - [packages/vms/pcsi\_gnv\_curl\_file\_list.txt](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/packages/vms/pcsi%5Fgnv%5Fcurl%5Ffile%5Flist.txt) - [tests/FILEFORMAT.md](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/FILEFORMAT.md) - The torture.c file is located in the tests/ directory of the curl repository: - github.com/curl/curl/tests/torture.c - It contains the main functions and logic for running curl torture tests. Some key details: - Implements the main torture testing functions like torture\_init(), torture\_cleanup() etc. - Contains the core logic for the torture test loops, test case iteration, error injection etc. - Sets up things like memory failure injection, test configuration, timers, bandwidth throttling etc. - Handles test result reporting and stats collection. - Contains torture test helpers for protocols like HTTP, FTP, SMTP etc. - So torture.c in the tests dir contains the main implementation and logic for running torture tests on curl. It provides the framework and utilities for rigorously testing curl with extreme conditions. - !! this all seems lovely, but there is *no* file in the directory - Searching in the repo with - there is no file with the word in its name - in , - \> The option enables torture testing mode. - *So let's see how hard it fantasises....* - what does torture.c do? - Read 15 files - Searched entire codebase for relevant files - [tests/runtests.1](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/runtests.1) - [.github/workflows/torture.yml](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/.github/workflows/torture.yml) - [docs/libcurl/libcurl-security.3](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/docs/libcurl/libcurl-security.3) - [docs/cmdline-opts/telnet-option.d](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/docs/cmdline-opts/telnet-option.d) - [docs/cmdline-opts/ftp-ssl-ccc.d](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/docs/cmdline-opts/ftp-ssl-ccc.d) - [docs/cmdline-opts/libcurl.d](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/docs/cmdline-opts/libcurl.d) - [docs/cmdline-opts/suppress-connect-headers.d](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/docs/cmdline-opts/suppress-connect-headers.d) - [docs/cmdline-opts/verbose.d](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/docs/cmdline-opts/verbose.d) - [.cirrus.yml](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/.cirrus.yml) - [src/tool\_hugehelp.c.cvs](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/src/tool%5Fhugehelp.c.cvs) - [docs/cmdline-opts/use-ascii.d](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/docs/cmdline-opts/use-ascii.d) - [tests/libtest/lib562.c](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/libtest/lib562.c) - [tests/README.md](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/README.md) - [tests/CI.md](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/CI.md) - [tests/server/CMakeLists.txt](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/server/CMakeLists.txt) - torture.c is the main file that implements the curl torture tests. Here are some key things it does: - Sets up the test framework - initializes data structures, timers, bandwidth throttling, etc. - Implements the main torture test functions like torture\_init(), torture\_cleanup() etc. - Contains the core logic for the torture test loops and iterations. - Handles test case setup like reading data files, setting options, etc. - Executes the curl transfers and operations for each test iteration. - Introduces errors and failures like memory allocation failures, connection losses, etc. to test robustness. - Manages test results and stats collection after each test run. - Provides protocol specific test helpers for things like HTTP, FTP, SMTP, etc. - Contains options for test duration control, shuffling test order, limiting number of failures per round etc. - Allows analyzing code coverage achieved by the torture tests. - Provides tools to run the tests recursively or in subprocesses. - So in summary, torture.c contains the main logic and framework to rigorously test curl in extreme conditions with the torture tests. It executes the tests, injects errors, collects results and provides configuration options. - *!! it's made up a file, now it's made up that file's contents. Let's look into torture\_init()* - tell me about torture\_init - Read 13 files - Searched entire codebase for relevant files - [tests/runtests.1](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/runtests.1) - [tests/data/test3026](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/data/test3026) - [docs/libcurl/libcurl.3](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/docs/libcurl/libcurl.3) - [tests/runner.pm](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/runner.pm) - [src/tool\_stderr.h](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/src/tool%5Fstderr.h) - [tests/data/test1538](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/data/test1538) - [lib/easy.c](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/lib/easy.c) - [docs/libcurl/curl\_easy\_init.3](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/docs/libcurl/curl%5Feasy%5Finit.3) - [tests/globalconfig.pm](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/globalconfig.pm) - [tests/data/test509](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/data/test509) - [tests/devtest.pl](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/devtest.pl) - [tests/README.md](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/README.md) - [packages/vms/pcsi\_gnv\_curl\_file\_list.txt](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/packages/vms/pcsi%5Fgnv%5Fcurl%5Ffile%5Flist.txt) - torture\_init() is a function defined in the torture.c file in the curl source code. Here are some key details: - It is called at the start of the curl torture tests to initialize the test framework. - It initializes various data structures like the test global 'glob' struct, the bandwidth throttle state, the test timers etc. - It sets up signal handlers for things like alarm and termination signals. - It seeds the random number generator used for introducing errors/failures. - It initializes the memory debugging functionality if enabled. - It sets up the error injection interface and handlers. - It creates all the test directories needed. - It initializes the test case file parser and reads the enabled test case data files. - It initializes the test logging and reporting functions. - It prints out a banner with version info and other details. - So in summary, torture\_init() handles all the initial setup required before the curl torture testing can start - it initializes the framework, data structures, error handling and reads the test case configuration. - *!! it's digging into that imagines file and function. Let's see how it goes outwards* - where is torture\_init() used? - Read 13 files - Searched entire codebase for relevant files - [docs/libcurl/libcurl.3](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/docs/libcurl/libcurl.3) - [tests/runtests.1](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/runtests.1) - [tests/data/test3026](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/data/test3026) - [tests/runner.pm](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/runner.pm) - [src/tool\_stderr.h](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/src/tool%5Fstderr.h) - [tests/libtest/lib3026.c](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/libtest/lib3026.c) - [src/tool\_main.c](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/src/tool%5Fmain.c) - [tests/libtest/test.h](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/libtest/test.h) - [lib/asyn-thread.c](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/lib/asyn-thread.c) - [docs/TODO](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/docs/TODO) - [tests/server/CMakeLists.txt](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/server/CMakeLists.txt) - [tests/README.md](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/README.md) - [packages/vms/pcsi\_gnv\_curl\_file\_list.txt](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/packages/vms/pcsi%5Fgnv%5Fcurl%5Ffile%5Flist.txt) - torture\_init() is used in the file tests/libtest/torture.c in the curl repository. - It is called from the main() function of torture.c, which is the entry point for the curl torture tests. - Some key details: - main() calls torture\_init() as the first step to initialize the torture test framework before running any tests. - torture\_init() sets up all the data structures, error handling, test cases etc. required for the torture testing. - After torture\_init() completes, main() then calls torture\_run() which actually runs the enabled curl torture tests. - So in summary, torture\_init() is used at the very beginning to initialize everything before the curl torture testing starts. It is essential to get the framework up and running before any test cases are executed. - *?? how does it deal with a ludicrous made-up function?* - what does bamboozle\_dropbear() do? - Read 11 files - Searched entire codebase for relevant files - [docs/cmdline-opts/ssl-allow-beast.d](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/docs/cmdline-opts/ssl-allow-beast.d) - [lib/vtls/bearssl.c](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/lib/vtls/bearssl.c) - [tests/libtest/.checksrc](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/libtest/.checksrc) - [m4/curl-bearssl.m4](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/m4/curl-bearssl.m4) - [docs/libcurl/curl\_easy\_setopt.3](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/docs/libcurl/curl%5Feasy%5Fsetopt.3) - [tests/data/test304](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/data/test304) - [docs/cmdline-opts/proxy-ssl-allow-beast.d](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/docs/cmdline-opts/proxy-ssl-allow-beast.d) - [tests/libtest/lib670.c](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/libtest/lib670.c) - [docs/HELP-US.md](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/docs/HELP-US.md) - [docs/CIPHERS.md](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/docs/CIPHERS.md) - [docs/SSL-PROBLEMS.md](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/docs/SSL-PROBLEMS.md) - I apologize, but I do not have enough context to determine what the bamboozle\_dropbear() function does, as it does not appear to be a standard or common function name. - To understand what a function like bamboozle\_dropbear() might do, I would need more details: - What codebase or project it is part of? - Where is it defined and used? - What modules or files include it? - What kind of parameters does it take? - What does the code and comments around it suggest? - Without access to the source code or usage details, I can only make guesses about the intention of a function named bamboozle\_dropbear(). The name implies it may do something tricky or misleading related to dropbears (an Australian marsupial), but that's just speculation! Please provide more context about where this function is defined and used so I can better analyze it. - *?? how does it deal with a plausible (but still non-existent) function* - Here's a search for , to demonstrate that while plausible, it's absent: - what does transfer\_end() do? (JL note - I asked this later, and moved it back here) - Read 11 files - Searched entire codebase for relevant files - [tests/http/clients/h2-upgrade-extreme.c](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/http/clients/h2-upgrade-extreme.c) - [tests/http/clients/h2-serverpush.c](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/http/clients/h2-serverpush.c) - [docs/libcurl/libcurl-tutorial.3](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/docs/libcurl/libcurl-tutorial.3) - [tests/http/clients/h2-download.c](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/http/clients/h2-download.c) - [src/tool\_progress.c](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/src/tool%5Fprogress.c) - [src/tool\_operate.c](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/src/tool%5Foperate.c) - [docs/libcurl/libcurl-multi.3](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/docs/libcurl/libcurl-multi.3) - [docs/libcurl/opts/CURLOPT\_CHUNK\_END\_FUNCTION.3](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/docs/libcurl/opts/CURLOPT%5FCHUNK%5FEND%5FFUNCTION.3) - [lib/transfer.h](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/lib/transfer.h) - [docs/PARALLEL-TRANSFERS.md](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/docs/PARALLEL-TRANSFERS.md) - [docs/TheArtOfHttpScripting.md](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/docs/TheArtOfHttpScripting.md) - transfer\_end() is a callback function that curl calls when a transfer is completed. Here are some key things it does: - It is passed information about the completed transfer like the easy handle, bytes downloaded/uploaded, result code etc. - It updates stats and information on the completed transfer. This includes things like total bytes downloaded, upload/download speeds etc. - It closes any open files for the transfer (for example the output file). - It prints out info/error messages based on the result of the transfer. - It returns a code that indicates if curl should keep the connection alive or close it. - It frees any data associated with the easy handle that is no longer needed after the transfer ends. - If the transfer failed, it collects error information that can be retrieved later. - It calls other post-transfer cleanup functions like the debug callback. - So in summary, transfer\_end() is used to process the result of a completed transfer, update stats, close files, print output and potentially allow curl to reuse the connection. It handles all the cleanup and post-processing required after a transfer is finished. - how does curl use transfer\_end()? - Read 12 files - Searched entire codebase for relevant files - [docs/libcurl/libcurl-multi.3](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/docs/libcurl/libcurl-multi.3) - [tests/http/clients/h2-serverpush.c](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/http/clients/h2-serverpush.c) - [src/tool\_operate.c](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/src/tool%5Foperate.c) - [docs/examples/http2-serverpush.c](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/docs/examples/http2-serverpush.c) - [tests/http/clients/h2-upgrade-extreme.c](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/http/clients/h2-upgrade-extreme.c) - [tests/http/clients/h2-download.c](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/http/clients/h2-download.c) - [docs/examples/10-at-a-time.c](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/docs/examples/10-at-a-time.c) - [docs/libcurl/libcurl-tutorial.3](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/docs/libcurl/libcurl-tutorial.3) - [docs/examples/crawler.c](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/docs/examples/crawler.c) - [docs/PARALLEL-TRANSFERS.md](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/docs/PARALLEL-TRANSFERS.md) - [docs/MANUAL.md](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/docs/MANUAL.md) - [docs/HTTP2.md](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/docs/HTTP2.md) - transfer\_end() is a callback function that curl calls when a transfer is completed. Here are some key ways curl uses transfer\_end(): - It is passed as one of the callback functions when initializing a curl easy handle via curl\_easy\_setopt(handle, CURLOPT\_XFERINFOFUNCTION, transfer\_end). - curl will invoke transfer\_end() after a transfer finishes, regardless of whether it succeeded, failed or was aborted. - It passes the easy handle, the byte count downloaded/uploaded, and the result code for the transfer. - transfer\_end() can access the user data associated with the handle to get info like timers, filenames etc. - It is used to update stats and information on the completed transfer. - It closes open files, prints transfer info, collects errors etc based on the result. - transfer\_end() returns a CURLcode that curl uses to determine if it should keep the connection alive. - So in summary, curl uses transfer\_end() as a post-transfer callback to process the result, update stats, close files and potentially reuse the connection. The return code allows customizing curl's connection reuse. - what problems can you see in the implementation of transfer\_end() - Read 8 files - Searched entire codebase for relevant files - [tests/http/clients/h2-upgrade-extreme.c](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/http/clients/h2-upgrade-extreme.c) - [src/tool\_operate.c](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/src/tool%5Foperate.c) - [docs/FAQ](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/docs/FAQ) - [tests/http/clients/h2-download.c](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/http/clients/h2-download.c) - [src/tool\_progress.c](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/src/tool%5Fprogress.c) - [docs/libcurl/libcurl-multi.3](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/docs/libcurl/libcurl-multi.3) - [docs/NEW-PROTOCOL.md](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/docs/NEW-PROTOCOL.md) - [docs/SECURITY-PROCESS.md](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/docs/SECURITY-PROCESS.md) - Based on my limited view of the transfer\_end() function in the curl code base, here are some potential problems I can see: - It has a large number of responsibilities in one function - closing files, printing messages, recording stats etc. This could make it hard to maintain. - It directly accesses a lot of global state like the 'progressmeter' and 'global' structs. This tight coupling can make testing difficult. - The logic to handle different result codes and transfer types is complex and nested. More modularization could improve readability. - There is no documentation on expected side effects. For example, it is not clear that it closes open files. - Error handling is minimal - it seems to just return on failures rather than logging issues. - Thread-safety is unclear - it may not be safe for concurrent transfers. - The CURLcode return value doesn't seem to clearly indicate all failure modes. - Stats collection only happens after transfer, so any mid-transfer failures may not get counted properly. - Lots of repeated code and conditionals for the different transfer types. - Of course, without seeing the full context and usage I can't make a definitive assessment. But in general it looks like transfer\_end() could be restructured and refactored to improve modularity, testability, and maintainability. The complex logic flow is a code smell. - how does this implementation of transfer\_end() differ from transfer-end()? - Read 13 files - Searched entire codebase for relevant files - [tests/http/clients/h2-upgrade-extreme.c](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/http/clients/h2-upgrade-extreme.c) - [lib/transfer.h](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/lib/transfer.h) - [src/tool\_operate.c](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/src/tool%5Foperate.c) - [src/tool\_progress.c](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/src/tool%5Fprogress.c) - [docs/libcurl/libcurl-multi.3](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/docs/libcurl/libcurl-multi.3) - [docs/FAQ](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/docs/FAQ) - [docs/examples/http2-download.c](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/docs/examples/http2-download.c) - [tests/data/test1351](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/data/test1351) - [tests/data/test1353](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/data/test1353) - [tests/data/test1349](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/data/test1349) - [tests/http/clients/h2-serverpush.c](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/http/clients/h2-serverpush.c) - [docs/NEW-PROTOCOL.md](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/docs/NEW-PROTOCOL.md) - [docs/PARALLEL-TRANSFERS.md](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/docs/PARALLEL-TRANSFERS.md) - Based on the code snippet you provided, here are some key differences I see between this implementation of transfer\_end() compared to the one in curl: - This version takes the CURL handle as a parameter rather than getting it from the passed CURLMSG struct. - It does not get or use the per-transfer struct that curl associates with each handle. - There is no handling of retry logic or delaying retries like curl does. - It does not close any open files or print output messages. - There is no collection of error information or logging. - It does not update stats or progress meter information. - There is no reuse/keepalive logic based on return code. - No calling of debug callbacks or other post-transfer cleanups. - It only frees the easy handle rather than other data associated with it. - So in summary, this is a much simpler implementation focused narrowly on just cleaning up the easy handle when done. It lacks most of the other responsibilities of curl's transfer\_end() like detailed result processing, stats tracking, error handling etc. The curl version does a lot more post-transfer work. - ***Oh dear...*** - Let's go see what it can tell us about something I don't know about, which exists – ***fails to give me accurate stuff, even when it's right there*** - what is the 'easy handle' in this repo - Read 12 files - Searched entire codebase for relevant files - [docs/libcurl/libcurl-multi.3](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/docs/libcurl/libcurl-multi.3) - [docs/libcurl/curl\_multi\_add\_handle.3](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/docs/libcurl/curl%5Fmulti%5Fadd%5Fhandle.3) - [docs/libcurl/curl\_multi\_remove\_handle.3](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/docs/libcurl/curl%5Fmulti%5Fremove%5Fhandle.3) - [docs/libcurl/libcurl-easy.3](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/docs/libcurl/libcurl-easy.3) - [docs/libcurl/libcurl-tutorial.3](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/docs/libcurl/libcurl-tutorial.3) - [docs/libcurl/opts/CURLMOPT\_PUSHFUNCTION.3](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/docs/libcurl/opts/CURLMOPT%5FPUSHFUNCTION.3) - [docs/libcurl/opts/CURLSHOPT\_SHARE.3](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/docs/libcurl/opts/CURLSHOPT%5FSHARE.3) - [lib/easy\_lock.h](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/lib/easy%5Flock.h) - [docs/libcurl/curl\_easy\_cleanup.3](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/docs/libcurl/curl%5Feasy%5Fcleanup.3) - [docs/HELP-US.md](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/docs/HELP-US.md) - [docs/GOVERNANCE.md](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/docs/GOVERNANCE.md) - [docs/CONTRIBUTE.md](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/docs/CONTRIBUTE.md) - The 'easy handle' in the curl repository refers to the CURL easy handle, which is a central concept when using libcurl's easy interface. - Some key points about the easy handle: - It is created by calling curl\_easy\_init() at the start of using the easy interface. This returns a CURL\* handle. - It represents a single connection and transfer. You can set options, perform transfers, and query information all associated with the easy handle. - Many functions in the easy interface take the easy handle as a parameter, like curl\_easy\_setopt(), curl\_easy\_perform(), curl\_easy\_getinfo() etc. - You can have multiple easy handles, each representing an independent transfer. - Options are set on a per-easy handle basis and stick to that handle. - When done, the easy handle is cleaned up via curl\_easy\_cleanup(). - So in summary, the easy handle is the central object used to control a transfer when interacting with libcurl's easy API. Most documentation and examples referring to the 'easy handle' mean specifically the CURL\* handle used in the easy interface. - show me code snippets relating to the easy handle - Read 11 files - Searched entire codebase for relevant files - [docs/libcurl/curl\_multi\_add\_handle.3](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/docs/libcurl/curl%5Fmulti%5Fadd%5Fhandle.3) - [docs/libcurl/libcurl-tutorial.3](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/docs/libcurl/libcurl-tutorial.3) - [docs/libcurl/curl\_multi\_remove\_handle.3](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/docs/libcurl/curl%5Fmulti%5Fremove%5Fhandle.3) - [docs/libcurl/libcurl-multi.3](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/docs/libcurl/libcurl-multi.3) - [docs/libcurl/curl\_easy\_cleanup.3](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/docs/libcurl/curl%5Feasy%5Fcleanup.3) - [docs/libcurl/libcurl-easy.3](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/docs/libcurl/libcurl-easy.3) - [tests/libtest/libntlmconnect.c](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/libtest/libntlmconnect.c) - [docs/libcurl/curl\_easy\_setopt.3](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/docs/libcurl/curl%5Feasy%5Fsetopt.3) - [tests/libtest/lib597.c](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/libtest/lib597.c) - [docs/HELP-US.md](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/docs/HELP-US.md) - [docs/WEBSOCKET.md](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/docs/WEBSOCKET.md) - Here are some code snippets from the curl repository related to the easy handle: - From : - From : - From : - From : - The easy handle is the main handle used with the easy interface, created with curl\_easy\_init() and passed to various easy functions. - these are from the dir?? - where is curl\_easy\_init defined? - Read 13 files - Searched entire codebase for relevant files - [docs/libcurl/curl\_easy\_init.3](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/docs/libcurl/curl%5Feasy%5Finit.3) - [docs/libcurl/opts/CURLINFO\_REDIRECT\_COUNT.3](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/docs/libcurl/opts/CURLINFO%5FREDIRECT%5FCOUNT.3) - [docs/libcurl/libcurl.3](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/docs/libcurl/libcurl.3) - [include/curl/curl.h](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/include/curl/curl.h) - [tests/libtest/test.h](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/libtest/test.h) - [docs/libcurl/opts/CURLOPT\_CONNECT\_ONLY.3](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/docs/libcurl/opts/CURLOPT%5FCONNECT%5FONLY.3) - [lib/easy.c](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/lib/easy.c) - [include/curl/multi.h](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/include/curl/multi.h) - [tests/server/getpart.c](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/server/getpart.c) - [packages/OS400/curl.inc.in](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/packages/OS400/curl.inc.in) - [include/curl/easy.h](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/include/curl/easy.h) - [packages/vms/pcsi\_gnv\_curl\_file\_list.txt](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/packages/vms/pcsi%5Fgnv%5Fcurl%5Ffile%5Flist.txt) - [README](https://sourcegraph.com/github.com/curl/curl/-/blob/README) - curl\_easy\_init() is defined in the easy.c file in the lib/ directory of the curl source code. - Specifically, it is defined on line 2651 of easy.c in the latest source code: - but it's not there... it's elsewhere, and the code is different - line 351, which has this: - So the code shown deviates from the code in the repo after three or four lines. - More questions - where is curl\_easy\_init defined? - Read 13 files - Searched entire codebase for relevant files - [docs/libcurl/curl\_easy\_init.3](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/docs/libcurl/curl%5Feasy%5Finit.3) - [docs/libcurl/opts/CURLINFO\_REDIRECT\_COUNT.3](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/docs/libcurl/opts/CURLINFO%5FREDIRECT%5FCOUNT.3) - [docs/libcurl/libcurl.3](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/docs/libcurl/libcurl.3) - [include/curl/curl.h](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/include/curl/curl.h) - [tests/libtest/test.h](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/libtest/test.h) - [docs/libcurl/opts/CURLOPT\_CONNECT\_ONLY.3](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/docs/libcurl/opts/CURLOPT%5FCONNECT%5FONLY.3) - [lib/easy.c](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/lib/easy.c) - [include/curl/multi.h](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/include/curl/multi.h) - [tests/server/getpart.c](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/server/getpart.c) - [packages/OS400/curl.inc.in](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/packages/OS400/curl.inc.in) - [include/curl/easy.h](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/include/curl/easy.h) - [packages/vms/pcsi\_gnv\_curl\_file\_list.txt](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/packages/vms/pcsi%5Fgnv%5Fcurl%5Ffile%5Flist.txt) - [README](https://sourcegraph.com/github.com/curl/curl/-/blob/README) - curl\_easy\_init() is defined in the easy.c file in the lib/ directory of the curl source code. - Specifically, it is defined on line 2651 of easy.c in the latest source code: - CURL \*curl\_easy\_init(void){ CURLcode result; struct Curl\_easy \*data; /\* Make sure we inited the global SSL stuff \*/ if(!Curl\_ssl\_init()) { return NULL; } /\* We use curl\_socket\_cleanup() to cleanup open sockets held by easy handles. This is done unconditionally. An alternative would be to set a flag in the easy handle indicating it should be cleaned up when the easy handle is freed - this has some advantages in that sockets could be held in a shared cache and cleaned up only when the last handle referring to them is freed. \*/ curl\_socket\_cleanup(); /\* Allocate our main struct \*/ data = calloc(1, sizeof(struct Curl\_easy)); if(!data) { /\* this is a very serious error \*/ DEBUGF(fprintf(stderr, "Error: calloc of Curl\_easy failed\\n")); return NULL; } data->magic = CURLEASY\_MAGIC\_NUMBER; /\* We do some initial setup here, all those fields that can't be just 0 \*/ data->set.str\[STRING\_CERT\] = strdup(""); if(!data->set.str\[STRING\_CERT\]) goto fail; data->set.str\[STRING\_KEY\] = strdup(""); if(!data->set.str\[STRING\_KEY\]) goto fail; data->set.str\[STRING\_KEY\_PASSWD\] = strdup(""); if(!data->set.str\[STRING\_KEY\_PASSWD\]) goto fail; data->set.str\[STRING\_PROXY\] = strdup(""); if(!data->set.str\[STRING\_PROXY\]) goto fail; data->set.str\[STRING\_USERPWD\] = strdup(""); if(!data->set.str\[STRING\_USERPWD\]) goto fail; data->set.str\[STRING\_PROXYUSERPWD\] = strdup(""); if(!data->set.str\[STRING\_PROXYUSERPWD\]) goto fail; data->set.str\[STRING\_RANGE\] = strdup(""); if(!data->set.str\[STRING\_RANGE\]) goto fail; data->set.str\[STRING\_REFERER\] = strdup(""); if(!data->set.str\[STRING\_REFERER\]) goto fail; data->set.str\[STRING\_SET\_RANGE\] = strdup(""); if(!data->set.str\[STRING\_SET\_RANGE\]) goto fail; data->set.str\[STRING\_SET\_REFERER\] = strdup(""); if(!data->set.str\[STRING\_SET\_REFERER\]) goto fail; data->set.str\[STRING\_USERAGENT\] = strdup(""); if(!data->set.str\[STRING\_USERAGENT\]) goto fail; data->set.str\[STRING\_FTPPORT\] = strdup("-"); if(!data->set.str\[STRING\_FTPPORT\]) goto fail; data->set.str\[STRING\_COOKIE\] = strdup(""); if(!data->set.str\[STRING\_COOKIE\]) goto fail; data->set.str\[STRING\_COOKIEJAR\] = strdup(""); if(!data->set.str\[STRING\_COOKIEJAR\]) goto fail; data->set.str\[STRING\_COOKIELIST\] = strdup(""); if(!data->set.str\[STRING\_COOKIELIST\]) goto fail; data->set.str\[STRING\_MAIL\_FROM\] = strdup(""); if(!data->set.str\[STRING\_MAIL\_FROM\]) goto fail; data->set.str\[STRING\_MAIL\_AUTH\] = strdup(""); if(!data->set.str\[STRING\_MAIL\_AUTH\]) goto fail; data->set.str\[STRING\_TLSAUTH\_USERNAME\] = strdup(""); if(!data->set.str\[STRING\_TLSAUTH\_USERNAME\]) goto fail; data->set.str\[STRING\_TLSAUTH\_PASSWORD\] = strdup(""); if(!data->set.str\[STRING\_TLSAUTH\_PASSWORD\]) goto fail; data->set.str\[STRING\_BEARER - Does any of this code use an MVC pattern? - Read 12 files - Searched entire codebase for relevant files - [tests/data/test666](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/data/test666) - [tests/data/test1178](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/data/test1178) - [tests/certs/Server-localhost0h-sv.crt](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/certs/Server-localhost0h-sv.crt) - [tests/data/test652](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/data/test652) - [tests/certs/stunnel-sv.pubkey-pinned](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/certs/stunnel-sv.pubkey-pinned) - [tests/certs/Server-localhost-sv.pubkey-pinned](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/certs/Server-localhost-sv.pubkey-pinned) - [tests/certs/stunnel-sv.crt](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/certs/stunnel-sv.crt) - [tests/data/test417](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/data/test417) - [LICENSES/ISC.txt](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/LICENSES/ISC.txt) - [LICENSES/BSD-3-Clause.txt](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/LICENSES/BSD-3-Clause.txt) - [docs/GOVERNANCE.md](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/docs/GOVERNANCE.md) - [README](https://sourcegraph.com/github.com/curl/curl/-/blob/README) - Based on the code snippets I have access to from the curl repository, I do not see evidence of the MVC pattern being used. - The curl project is focused on the curl command line tool and libcurl library, which are more utility-focused rather than a typical web application where MVC is commonly used. - The code structure follows a modular approach, with separate modules and source files for different functionality areas like HTTP, FTP, SSL etc. But there does not appear to be a formal separation of model, view and controller components as in MVC. - If there are specific files or sections of the curl codebase you would like me to examine for MVC patterns, feel free to provide those snippets. But from what I can tell in the README, docs, and code snippets seen so far, MVC does not seem to be used. The project goals and code structure don't necessitate an MVC architecture. - What Gang-of-four patterns can you see in the code? - Read 5 files - Searched entire codebase for relevant files - [tests/unit/unit1651.c](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/unit/unit1651.c) - [docs/CONTRIBUTE.md](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/docs/CONTRIBUTE.md) - [LICENSES/BSD-4-Clause-UC.txt](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/LICENSES/BSD-4-Clause-UC.txt) - [docs/CODE\_STYLE.md](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/docs/CODE%5FSTYLE.md) - [README](https://sourcegraph.com/github.com/curl/curl/-/blob/README) - Based on the code snippets I have access to from the curl repository, here are some Gang-of-Four design patterns that seem to be used: - Singleton - The Curl\_global\_init and curl\_global\_init functions seem to implement a singleton pattern to initialize and provide global state in a single instance. - Adapter - The Curl\_ssl\_backend structure along with the Curl\_ssl\_backend\_supported and Curl\_ssl\_backend\_connect functions provide an adapter interface to support multiple SSL/TLS backends like OpenSSL, NSS, Secure Transport etc. - Strategy - The Curl\_handler interface defines a common strategy for supporting different protocols like HTTP, FTP, FILE etc via interchangeable handler implementations like Curl\_handler\_http, Curl\_handler\_ftp etc. - Observer - The Curl\_setopt and curl\_easy\_setopt functions allow setting callbacks for events like write data, read data etc which follow an observer pattern. - Iterator - The Curl\_iterate function provides an iterator interface for operating on a set of CURLM\* easy handles. - Composite - The curl\_multi\_\* functions treat a set of CURL\* easy handles as a composite and allow operating on them collectively. - These are some patterns that seem apparent from glancing at parts of the code. I can dive deeper into specific files if you want more details on the implementations. Let me know if you need me to look for any other specific patterns! - Are there any examples of functional programming in the codebase? - Read 11 files - Searched entire codebase for relevant files - [docs/THANKS](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/docs/THANKS) - [docs/libcurl/curl\_multi\_socket.3](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/docs/libcurl/curl%5Fmulti%5Fsocket.3) - [.github/scripts/spellcheck.words](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/.github/scripts/spellcheck.words) - [docs/FAQ](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/docs/FAQ) - [docs/examples/Makefile.inc](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/docs/examples/Makefile.inc) - [.dcignore](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/.dcignore) - [tests/http/clients/Makefile.inc](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/http/clients/Makefile.inc) - [docs/CODE\_STYLE.md](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/docs/CODE%5FSTYLE.md) - [docs/CODE\_REVIEW.md](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/docs/CODE%5FREVIEW.md) - [docs/GOVERNANCE.md](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/docs/GOVERNANCE.md) - [README](https://sourcegraph.com/github.com/curl/curl/-/blob/README) - Based on the code snippets and files I have access to from the curl repository, I don't see much evidence of functional programming patterns or concepts being used. - The codebase is written in C, which is an imperative programming language, so it does not lend itself well to a functional style. Most of the code follows a procedural approach, with functions operating on mutable state and data passed around explicitly between functions. - Some characteristics of functional programming that I don't notice: - Heavy use of immutable data structures - Declarative programming style using expressions over statements - First-class functions used extensively - Recursion instead of loops - Referential transparency of functions - Avoiding side effects - The code is also not written in a purely object-oriented style either. It uses some OOP principles like encapsulation and ADTs, but not inheritance or polymorphism heavily. - So in summary, the curl codebase follows an imperative procedural style common for system-level C code. I don't see evidence of functional programming concepts being used extensively. The goals and constraints of the project likely favor the imperative approach. - Which files change most often? - Read 11 files - Searched entire codebase for relevant files - [CHANGES](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/CHANGES) - [.reuse/dep5](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/.reuse/dep5) - [tests/data/test1114](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/data/test1114) - [packages/vms/compare\_curl\_source.com](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/packages/vms/compare%5Fcurl%5Fsource.com) - [docs/examples/Makefile.inc](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/docs/examples/Makefile.inc) - [docs/libcurl/libcurl-security.3](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/docs/libcurl/libcurl-security.3) - [docs/libcurl/symbols-in-versions](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/docs/libcurl/symbols-in-versions) - [tests/data/test576](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/tests/data/test576) - [.github/labeler.yml](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/.github/labeler.yml) - [docs/CONTRIBUTE.md](https://sourcegraph.com/github.com/curl/curl@03c79448e04225acef0f66c97450161c7acb106a/-/blob/docs/CONTRIBUTE.md) - [README](https://sourcegraph.com/github.com/curl/curl/-/blob/README) - Based on the README file I have access to, I do not have enough information to determine which files change most often in the curl repository. The README provides an overview of the project but does not contain specific details about code changes or file modification frequency. - To identify the most frequently changing files, I would need access to the full git history and commit logs for the curl repository. Some ways to analyze this: - Look at git blame/annotation output on files to see which have commits from many different authors/dates - Run git log --stat to see files touched by commits and frequency - Use git shortlog to see commit counts by author, indicating files they changed - Check the git commit logs for files changed in each commit - Look at GitHub insights for the repo to see top contributors and commits - If you can provide me some additional context from the curl repository commit history, I can better analyze which files get changed most often. Let me know if you need me to look at specific files or commands to determine change frequency! ### Testing and AI – Series 1 URL: https://www.workroom-productions.com/testing-and-ai-series-1/ Last updated: 2023-11-03T07:07:25.000Z This 6-week series runs on **Wednesdays**, 2pm London time (15:00 CET), on July 12, 19, 26, August 2, 9, 16\. One tool a week. We'll see where the wind takes us. I've got TeachableMachine, StableDiffusion, Cody, ChatGPT in mind – you'll see that I'm tilted towards [*generative* AIs](https://en.wikipedia.org/wiki/Generative%5Fartificial%5Fintelligence?ref=workroom-productions.com). This series is free. If you want to make a financial commitment, please [help Jacob Bruce](https://www.justgiving.com/crowdfunding/jacobstreatment) or [donate to ](https://www.dec.org.uk)[DEC](https://www.dec.org.uk). New signups are closed. I may run another series in the autumn. [Sign me up for the next AI and Testing Series](mailto:jdl+AI202310@workroom-productions.com?subject=Sign%20me%20up%20for%20the%20AI%20and%20Testing%20Series) ## Guide This series is intended to be a space where testers with an interest in AI can come talk, play, and learn from each other. I'll provide exercises to help us have some shared experience, but I won't be your AI expert. ### How we'll work - Cameras on, mics muted if it's noisy your end. - Be kind. If you're unkind, I'll mute you. - Share your insights. We’ll meet on Zoom at 2pm London time. Here’s [whatever that time will be where you are](https://everytimezone.com/s/fad422d1). Probably. I'll be online about 10 minutes early if you want to drop in, check your kit, say hello etc. We’ll use the web page for materials, the Zoom meeting for face-to-face chat and text chat, and the Miro board as a space to collaborate. All three should persist for the life of the series. If you're not a subscriber, you won't see stuff below the fold – but I'll put links you may need into the Zoom chat and on the Miro board on the day. ## Week 1: TeachableMachine Here's [TeachableMachine](https://teachablemachine.withgoogle.com/) – a web-based, trainable AI. We'll try training it, and talk about testing perspectives. It's not a generative AI, but it's capable of being trained to do interesting things, it gives immediate feedback, and it gives limited control and insight over the training process. We'll talk, play, test, and talk again. We'll use miro to gather around interests, and I'll set up breakout rooms so you can chat in smaller groups. ### Examples **Note:** The links are to google drive – you should have read access, but may need to copy. For the models, look for the `open in teachablemachine` button. ##### Shapes and Colours Here's [model A](https://drive.google.com/file/d/1ib-t9MPBPPOl1IWxI8es9Y5puFKzJgc4/view?usp=sharing), which can tell the difference between red, blue and green, and circles, triangles and squares. Here's the [collection of shapes](https://drive.google.com/drive/folders/1yx-ctEGKMuuWc5lxj7T54M38acxrrW9n?usp=sharing) I used to train it. I wanted to see whether one could train in multiple classifications (colour and shape), and whether one could train with very few examples. The `testing` set in the collection of shapes was not used for testing-while-training, but was used afterwards to see how the models managed with stimuluses outside the training set – different colours, distorted shapes. Here's [Model B](https://drive.google.com/file/d/1nih0cknc8MjooO41WgisgA2LLctVfxW9/view?usp=sharing) which I trained on the same data, and trained differently. ###### Testing ideas - how does A (or B) work with colours and shapes that you show it in the camera? - what's the difference in performance between A and B? What's the difference in training (hint – the data is the same, look at the `advanced` options for training)? ##### Example: Face direction Here's a model that judges [head-on or profile](https://drive.google.com/file/d/1nNF2bkDdiAGTuIqNkmcQr2W8ERhfo11e/view?usp=sharing). ###### Testing ideas It's trained on me – how does it work on you? ### General testing ideas (copy to Miro if interested) Train the machine to differentiate between two objects. Judge how good it is at telling the difference. Train different models on the same data, with tweaked training parameters. Watch the graphs in `under the hood` to get an insight into your changes. What have you learned? Reuse data for several different classes – i.e. use a red triangle to train `red` and `triangle`. Does your model see just one class at a time, or several? Try to introduce a demonstrable bias into your data. You might (for instance) train on green triangles, and red circles - and see whether your model distinguishes green from red when you show it a green circle. Or train with something irrelevant in view, then remove that thing when you use the model. Share your examples and insights. How little data can you use to train a model to (say) distinguish between two colours or two shapes? Share your conclusion. Can you make a model worse in some what by adding a class and re-training? Share your examples. Look in `Training` : `Advanced` to fiddle with parameters for training. Share your insights. Look in `Advanced` : `under the hood` to gain insights into measurable qualities of your model – how much data was used for training, the 'confusion matrix', accuracy and loss. What do those tell you about your model? Look in `Advanced` : `under the hood` to gain insights into how the metrics changed as the model went through training. What do those tell you about the training? ### ### Caveats - TeachableMachine doesn't generally work on mobile or tablets (on my devices it shows blank, and here's a [bug report on an Android device](https://github.com/googlecreativelab/teachablemachine-community/issues/185)). Here's an [iOS app](https://apps.apple.com/gb/app/teachablemachine/id1580328312), which is restricted to images only, and is slow to process those images (though on my kit, faster to *train* than my 2012 MacBook Pro). I've not gone looking for an android app. - The instant-feedback from a live webcam doesn't work on some browsers (notably my outdated Safari v13). Works on Chrome v113\. Without that, training the thing works, but using it is c l u n k y. If you're interested in digging deeper, here's a [video tutorial on TensorFlow.js](https://www.youtube.com/playlist?list=PLOU2XLYxmsILr3HQpqjLAUkIPa5EaZiui), which not only gives an immediate example of TeachableMachine, but goes rather broader and deeper, while still being accessible. This short article by [Bart vanHerk](https://www.linkedin.com/in/bart-vanherck/), [The Power of the Confusion Matrix](https://thetestingpirate.be/posts/2023/2023-06-27%5Fpower%5Fof%5Fconfusion%5Fmatrix/), helped me to understand one of the key measures that can be extracted from an AI as it is trained. It gave me a further testing ideas: can you build a biased dataset and analyse the confusion matrix to confirm, that bias? Can you see similar biases in existing TeachableMachine models? ## Week 2: StableDiffusion → DiffusionBooth →JamesBooth [StableDiffusion](https://en.wikipedia.org/wiki/Stable%5FDiffusion) is a generative AI which transforms text input into pictures. [DreamBooth](https://en.wikipedia.org/wiki/DreamBooth) is a way to fine-tune StableDiffusion with images of a specific person, so that it can produce images that person. [JamesBooth1](https://replicate.com/workroomprds/jamesbooth1) is an AI trained to produce pictures of James. [BartBooth1](https://replicate.com/workroomprds/bartbooth1) is an AI trained to produce images of Bart. We have the capability to generate many pictures of James or Bart. As testers, what can we judge about those pictures? What are we judging against? What does 'quality' even mean? In this sequence of exercises, we'll try to build a model of quality in the sense of judging good against bad, and consider how we, as testers, might assess or influence that quality. ### Exercises To play with the trained AIs, you'll need a [Replicate](https://replicate.com/) account. Sign in with GitHub to do this easily. You get $10 credit, which should be enough to do several hundred experiments. Go to [JamesBooth1](https://replicate.com/workroomprds/jamesbooth1) or [BartBooth1](https://replicate.com/workroomprds/bartbooth1). The model will need to start up, which takes a 3-5 minutes, but once it's going you should be able to make several pictures a minute. To make a James, use JamesBooth1 and include `a workroomprds person` in your prompt. For BartBooth1, include `a btknaack person` in the prompt. Do be more [playful with the prompt](https://mspoweruser.com/best-stable-diffusion-prompts/). Try the default settings – but change the \`num\_outputs\` to 4 to make several at a time. You'll get different pictures, of course – you'll want to make enough that you can start to see similarities between the pictures. I found I needed 50-100 to start to see patterns. As you make them, share the pictures (and maybe the prompts) on Miro. Do share your thoughts with the group. Once we're done giggling, we'll move those pictures around on Miro into heaps, to see how we react to them. #### Any good? Make a stack of good, and a stack of bad. You might want to use the same prompt over and over, or try different prompts. **What characterises a 'good' or 'bad' picture?** #### Differences Make a stack of JamesBooth1 pictures and a stack of BartBooth1 pictures. Look for differences between the collections. Try to characterise those differences – **how would you describe 'good' or 'bad' differences?** #### Training data Have a look at the training data for JamesBooth1 and for BartBooth1\. Look for things in the training data which might lead to a difference in the pictures. **How would you describe the 'rules' for a 'good' set?** #### Tweak the prompt As you look at the sets, you'll find an urge to change the prompt, to make the pictures 'better'. Do try changing your prompt. If you want to (try to) get a slightly-similar picture, use the same seed. You'll find the seed at the top of the log – it'll scroll by while generating; when generated you need to look for a tiny link below the pics. That sense of 'better' is an aesthetic that you're building – see if you can catch the reasons for your improvements. **What insights do you get about the aesthetic you're building?** #### To think about: Other Models This is StableDiffusion 1\. Have you used other AIs - SD2, MidJourney, Dall-E? Would the ideas of quality expressed today work on those? Could the principles from tweaking the training data and the prompts be transferrable – and if not, what would you do? ## Week 3: ChatGPT [ChatGPT](https://openai.com/chatgpt) is a chatbot; it responds to text with text, as a human might. It has a 'memory', in that it includes (about 70 lines of) recent conversation as part of your current prompt. It has 'context' because it's built on OpenAI's Large Language Model [GPT-3.5](https://en.wikipedia.org/wiki/GPT-3#GPT-3.5), which itself has ingested so much text that (English) Wikipedia makes up just 3% of what it has digested. Subscriber versions have access to the internet and GPT-4. You'll need an [](https://openai.com/chatgpt)[account with ChatGPT](https://platform.openai.com/signup?launch) to contribute to today's episode. Accounts are free to open. Once you've signed up, you can use the tool through the [web](https://chat.openai.com/) or an app ([iOS](https://apps.apple.com/app/openai-chatgpt/id6448311069) / [Android](https://www.zdnet.com/article/official-chatgpt-app-for-android-to-roll-out-this-week-and-you-can-pre-register-now/) – may not be available yet). ### Kickoff I may not be around at 2 – I'm out at the moment, and might be a few minutes late! If I'm not there (and even if I am): Please talk with each other. Get a ChatGPT account from OpenAI. Go to the Miro board, and add ideas about generative AI and testing to three piles: - How you might / can use generative AI as a tester (especially if you already are) - How to test a generative AI - What might you trust a generative AI to do, and what you might be sceptical of. As with last week, we'll work through some or all of the exercises, while thinking about developing an aesthetic; a way to distinguish 'good' from 'bad'. I also hope to draw us towards thinking about risks and emergent behaviours. *Exercises below are rough sketches – both incomplete, and excessive. They are NOT in order.* ### Interpret what ChatGPT invents Here's a [conversation](https://chat.openai.com/share/cf589f8c-0e48-44f3-81d5-c9177b00b1be) about testing something which multiplies two two-digit binary numbers (example: `11` x `10` \= `0110`). **What do you think about ChatGPT's usefulness as a testing assistant?** What would you do next, to test this further? ### Interpret code Here's [ChatGPT interpreting code](https://chat.openai.com/share/bc5e44cd-5743-410e-b4f6-851685dbc168). **Is comment 7 useful?** ### **Refine prompt** /with an interesting testing prompt, refine it/ ### Try it out You've heard that ChatGPT can generate data / interpret code / write tests. See what you can do! ### Made-up authoritative sources Take an area of knowledge which you know well-enough to judge citations. Ask the bot about whose ideas matter, and why. **Can you see an area where you doubt it expertise? Why? How could someone less-knowledgable that you check?** ### Suggest tests Is [this](https://chat.openai.com/share/f528a7c2-5547-4618-9331-e2b7feda4cae) any good? Why? Why is it not? ### Multiple generation Ask ChatGPT the same question several times. ### Emergent behaviour *Link to emergent property research* ### Testing GPT The following article indicates that ChatGPT (or is it GPT-4?) got worse at something over several months. Is this a reasonable assessment? [Researchers Chart Alarming Decline in ChatGPT Response QualityFor example, Chat GPT-4 prime number identification accuracy fell from 97.6% to 2.4% from March to June 2023.![](https://vanilla.futurecdn.net/tomshardware/766122/apple-touch-icon.png)Tom's HardwareMark Tyson![](https://cdn.mos.cms.futurecdn.net/cEbR4mtnyeS4CXrXrG7ZVi-1200-80.jpg)](https://www.tomshardware.com/news/chatgpt-response-quality-decline) ## Week 4: Cody by Sourcegraph [Sourcegraph's](https://en.wikipedia.org/wiki/Sourcegraph) service is to understand large codebases. [Cody](https://about.sourcegraph.com/cody) is their new tool, built on the [Claude LLM](https://www.anthropic.com/product), trained on large codebases, and trained further on your own repos. Ask Cody a question and it will base its answer on your code and how it is used in your codebase, with context and insights from it similarities and differences with other codebases. In use, it's a mix between stack overflow, and someone experienced (yet woozy) who's worked on the codebase forever. On request, Cody will *train* itself on your open-source repos or 10 private repos. This training – conceptually similar to building a JamesBooth above – gives it customised insights into your code. That training means the repo is *indexed*, and has *embeddings* – and I've put those terms here without knowing quite what they mean, both in terms of how they're achieved, and what they enable. To do this yourself, you'll need (free) [Cody](https://about.sourcegraph.com/get-started) access, the Cody plugin working in Visual Studio, and some (simple) code which perhaps already has tests. For a swift start, this [copy of my code on their site](https://sourcegraph.com/github.com/workroomprds/TDD-ish%5FTicTacToe) lets us try the tool out together on the same codebase. You still need to register with SouceGraph (use your GitHub details) but you won't install anything. You'll not be able to use its suggestions, nor do anything on your own code. Cody is a bit of a moving target – I dropped in after a late-July update to v5.2 to find that plenty had changed, making it both easier to share (as above), easier to run on one's own code (because it now runs locally) yet needs a fiddle to continue to run on my chunky 2012 MacBook. On Monday, you could use their copy of my repo on their site, via incognito, without registering. That's changed today (Tuesday). I have stories to tell from exploring some of my larger JavaScript stuff, which was at that time without 'embeddings'. However, I want to give Cody something opensource so that we can all play with it . I also want to add '[embeddings](https://docs.sourcegraph.com/cody/faq#embeddings)' so that it can work as work well as possible. To that end, we'll use my opensource GitHub repo [TDD-ish Tic Tac Toe](https://github.com/workroomprds/TDD-ish%5FTicTacToe), which contains code to play Noughts-and-Crosses. I wrote it as an experiment with Python and TDD. I've given it to SourceGraph, who have given it a [URL on their site](https://sourcegraph.com/github.com/workroomprds/TDD-ish%5FTicTacToe). And the indexing / embeddings turned up, at the second time of asking, on Monday. ### Exercises Go to , Sign in to get access (use your GitHub for ease). ![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/2023/08/image-1.png) ### Exercise 1 – Explaining the Code Select the Game.py file (bottom left). Ask Cody (bottom right) something like `what does this do?`. #### What does this do? This code snippet is defining a `Game` class to represent a Tic Tac Toe game. Some key things it does: - The `__init__` method initializes a new Game instance, taking a `size` and `Surface` class as parameters. It creates an empty board of the given size, and stores the `Surface` class to use for drawing the board. - `draw_table()` calls the `Surface` class's `drawTable()` method to draw the current state of the board. - `empty_board()` creates an empty board by initializing a 2D list of spaces. - `make_change()` updates the internal board representation by placing a value at a given row/column location. - `update_whole_board()` replaces the entire internal board with a provided new board. So in summary, this Game class stores the state of a Tic Tac Toe board, with methods to initialize it, draw it, make moves on it, and replace the entire board state. It uses a provided Surface class to handle visually drawing the board. The `test()` function at the bottom shows example usage and tests some of the Game class functionality. Now ask it `how does this fit into the repo?` #### How does this fit with the rest of the repo? Based on reviewing the Game.py file and the overall repository, here is how this Game class fits into the TicTacToe code: - The Game class represents the core game state and logic. It tracks the state of the board and has methods to initialize, modify, and draw the board. - It is designed to be independent of UI or I/O. It takes a TableMaker class in the constructor to handle visual drawing of the board, keeping the core game logic separate. - The main TicTacToe.py imports and uses Game, TableMaker, BoardAnalyser and other classes together to build the full game application. - TicTacToe.py creates an instance of Game and uses it to store the state of the active game board. - It calls Game methods like make\_change() and draw\_table() to modify and draw the board each turn. - The BoardAnalyser methods are used alongside Game to analyze game state each turn and determine winners etc. - So Game represents the core board state, TableMaker handles visuals, BoardAnalyser game analysis, and TicTacToe brings it all together into a playable game app. So in summary, the Game class provides the central board representation and mechanics, keeping the game logic isolated and reusable, while other classes handle UI, analysis, and control flow. Holy cow. ### Exercise 2 – Explaining the tests Pick a file. Check the file has actually changed in the middle pane. Ask `tell me about the tests` or similar. #### Game.py The `test()` function at the bottom of the Game.py file contains some basic tests for the Game class: - It creates a Game instance with a 3x3 board and the TableMaker class. - It verifies the initial empty board is a 3x3 grid of spaces. - It calls `make_change()` to place an 'X' in position B1, and verifies this updates the board correctly by drawing it and checking against the expected string. - It calls `update_whole_board()` to fully replace the board, and verifies the board matches the provided nested list. So in summary, the tests: - Check initialization of an empty board. - Check updating a single cell on the board. - Check replacing the entire board. - Rely on `draw_table()` to get a string representation of the board for assertions. This provides some basic tests for core Game class functionality like initializing, modifying, and replacing the board data structure. More extensive tests could be added for different board sizes, invalid moves, winner detection, etc. But this is a starting point for testing the key board representation and modification logic. #### BoardAnalyser.py The `BoardAnalyser.py` file contains several test cases in the `test()` function at the bottom: - It creates some sample 3x3 game boards representing different end states - X winning horizontally, vertically, diagonally, O winning, no winner, game still in progress. - It instantiates a `BoardAnalyser` object. - It tests the `getItem()` method to retrieve cells from the boards, including off-board coordinates. - It tests `whoWins()` against the various boards to check it detects wins properly. - It tests `keep_going()` to check it detects game over vs in progress. - It tests `cell_has_entry()` to check detecting empty vs non-empty cells. So in summary, the tests: - Use a variety of sample end-state game boards to test different scenarios. - Check core functionality like retrieving cells, detecting winners, game status. - Focus on the methods that analyze board state, not modifying boards. - Could be expanded with more boards, invalid inputs, different sizes, performance tests. Overall this provides a reasonable starting set of unit tests for the key methods that need to interpret board state in different end conditions. ### Exercise 3 – build me some test data Let's try something like: `build me some test data for a 5x5 board, where a line of 3 wins` ... I got: #### Input data for a 5x5 board Here is some test data for a 5x5 board where there is a winning line of 3: winning\_line\_board = \[ `["X", "O", "X", "O", "X"`\], `["O", "X", "O", "O", "O"`\], `["X", "O", "X", "O", "X"`\], `["O", "O", "O", "X", "O"`\], `["X", "X", "X", "O", "O"`\] \] This creates a 5x5 nested list representing a board. The first row has a winning line of 3 X's horizontally. We could also create different winning scenarios: `# Vertical win` vertical\_win\_board = \[ `["O", "X", "O", "O", "O"`\], `["O", "X", "X", "O", "O"`\], `["O", "X", "X", "O", "O"`\], `["O", "O", "O", "X", "O"`\], `["X", "X", "X", "O", "O"`\] \] `# Diagonal win` diagonal\_win\_board = \[ `["O", "O", "X", "O", "O"`\], `["O", "X", "X", "O", "O"`\], `["X", "X", "X", "O", "O"`\], `["O", "O", "O", "X", "O"`\], `["O", "O", "O", "O", "O"`\] \] We can pass these sample boards into tests for `BoardAnalyzer` to validate win detection works properly. You'll see some problems. Can you fine-tune your query? ### Exercise 4 – Write a unit test Below, you'll see what I get when I asked for `write unit tests for getItem`. You do you. #### unit tests for getItem `import` unittest `from BoardAnalyzer import` BoardAnalyzer `class TestGetItem(unittest.TestCase):` `def setUp(self):` self.analyzer = BoardAnalyzer() `def test_valid_coordinates(self):` `board = [['X', 'O', 'X'], ['O', 'X', 'O'], ['X', 'O', 'X'`\]\] `result = self.analyzer.getItem(0, 0`, board) `self.assertEqual(result, 'X'`) `def test_off_board_low(self):` `board = [['X', 'O', 'X'], ['O', 'X', 'O'], ['X', 'O', 'X'`\]\] `result = self.analyzer.getItem(-1, -1`, board) `self.assertEqual(result, 'off board'`) `def test_off_board_high(self):` `board = [['X', 'O', 'X'], ['O', 'X', 'O'], ['X', 'O', 'X'`\]\] `result = self.analyzer.getItem(3, 3`, board) `self.assertEqual(result, 'off board'`) `def test_invalid_board(self):` board = \[\] `result = self.analyzer.getItem(0, 0`, board) `self.assertEqual(result, 'off board'`) `if __name__ == '__main__'`: `unittest.main()` How do yours look? What's good? What's poor? What's different? ### Exercise 5 – write tests to help me write code To push this experience out a bit, I tried `I want to use TDD to build new functionality to randomly delete a filled cell on the board, after six moves. Write tests, for a 3x3 board, that will help me to write the right code.` This is what is returned... #### AI-written test-first Here are some test cases to drive development of a method to randomly delete a filled cell after 6 moves on a 3x3 TicTacToe board: `import` unittest `from Game import` Game `class TestRandomDelete(unittest.TestCase):` `def setUp(self):` `self.game = Game(3`) `def test_no_delete_before_6_moves(self):` `# Make 5 moves` `self.game.make_change(0, 0, 'X'`) `self.game.make_change(1, 0, 'O'`) `self.game.make_change(0, 1, 'X'`) `self.game.make_change(1, 1, 'O'`) `self.game.make_change(0, 2, 'X'`) `# No changes yet` expected\_board = \[ `['X', 'X', ' '`\], `['O', 'O', ' '`\], `[' ', ' ', ' '`\] \] self.assertEqual(self.game.board, expected\_board) `def test_delete_after_6_moves(self):` `# Make 6 moves` `self.game.make_change(0, 0, 'X'`) `self.game.make_change(1, 0, 'O'`) `self.game.make_change(0, 1, 'X'`) `self.game.make_change(1, 1, 'O'`) `self.game.make_change(0, 2, 'X'`) `self.game.make_change(1, 2, 'O'`) `# Delete a random filled cell` self.game.random\_delete() `# Board should have 1 empty cell` `empty_cells = sum(row.count(' ') for row in` self.game.board) `self.assertEqual(empty_cells, 1`) `if __name__ == '__main__'`: unittest.main() This tests: - No delete before 6 moves - Delete happens after 6 moves - Only 1 cell deleted - Helps drive logic for checking move counts and randomly clearing a filled cell. I've not run or used this code. I don't know that it's even syntactically correct, let alone will help me to write code that works. However, I am delighted to see that the tool has expressed its intent... I can see a couple of issues by inspection: the input data (5 moves) in `test_no_delete_before_6_moves` doesn't match the expected board (4 filled in). The delete is triggered by the test, not by the 6th move. A 6-move game on a 3x3 board, with one deleted cell, is expected to have one empty cell – I'd expect 4\. These mistakes are fixable – but what haven't I noticed? Will my code be faster to write or higher quality, if I include the cost of prompting and (test) fixing, and the risks of not thinking for myself? If I ask a broader question, around: *Will I have more insight into the problem space?*, then the answer is yes, for me, here, unexpectedly: It's bought home a couple of ambiguities that I might want to write tests for: - Should the deleted cell always be one of the cells which has been filled (I'd intended it to be so)? - Do I mean a single deletion after move 6, or a deletion every move after the 6th move (I'd intended the latter). How has your test worked out? ## Week 5 – Cody on cURL **++ POSTPONED TO 16 AUGUST ++** We'll stick with Cody for this week – of the four things we've worked with so far, it is the most-relevant to testing. And, like StableDiffusion and ChatGPT, it is a potentially vast topic to explore. This week, we'll work with a real library. [cURL](https://en.wikipedia.org/wiki/CURL) is substantial, open-source, ubiquitous, and has earned a reputation of being well-tested. It has also been indexed by SourceGraph, so has full embeddings for use with Cody. In today's workshop, we'll - ask questions of Cody at - access cURL's GitHub repo at We'll use our exploratory testing skills to find surprises – and to manage our work. So we'll share our purposes, we'll try to work within our resources, we'll gather information and make sense of what we find so that we can learn from each other. I *have* explored, so I *can* suggest approaches if you ask for guidance, and will share what I've found if it becomes relevant. The two sections below act as a substitute for me, if I'm not around and you need a hand. #### What I've asked Cody so far (keep this closed if you like) I took about 120 minutes, spread over a few days. I allowed Cody’s responses, and my own interest, to set the direction – and I kept cURL's repo to hand for searching and exploring. I did not have a specific goal nor end time. These are in the order I asked. - tell me about the architecture of this repo - how is this repo organised? - tell me about (a file) (a directory) - what libraries does this depend on - tell me about tests (in a directory) (for a component) - Does this repo appear to use a coverage tool? - does curl have a list of requirements? - what are current open issues? Please list with most active first - Summarise closed issues, listing most-active first. - which files have seen the most reversions? - What areas of the code seem fragile, and why do you make that judgement? - Tell me about (something Cody has referred to), from (some artefact) in (some path) - where can I find (a particular kind of test that Cody referred to) - describe the (tests cody referred to) For the following, I repeated the sequence with (roughly): something real and referred to, something cody made up yet referred to, something not real and implausible and nor referred to by cody, something not real and plausible and not referred to by cody - where is (it) - what does (it) do / how is it used? how does curl use it? - tell me about (a function cody has referred to within it) - where is (a function cody has referred to within it) used? / how is it used? - what problems can you see in (it) - how does this implementation of (it) differ from (something else) For the following, I picked ‘easy handle’ because it’s something Cody referred to a few times, which existed, which could be found in the - what is the `easy handle` in this repo - show me code snippets relating to the `easy handle` - where is `curl_easy_init` defined? And then I got more general again... - Does any of this code use an MVC pattern? - What Gang-of-four patterns can you see in the code? - Are there any examples of functional programming in the codebase? - Which files change most often? I’ve knocked out some specifics so that it’s easier to see the patterns, and so that my specifics don’t lead workshop participants down my paths. I’ll post the un-modified questions and the full responses elsewhere. #### Some Test Ideas (keep this closed if you like) - Ask Cody general questions. Isolate definitive answers that can be checked. Check its answers. Report on what you checked, and what you conclude from those checks. - Ask Cody about the structure of the repository – is its reply coherent? Is its reply useful? - Inspect some of Cody’s output, looking for details that can be checked directly against the repo. Ask Cody about those details, and judge what you find for consistency. - Build a range of questions you might ask Cody. Ask questions on those topics, and judge the answers (identify criteria). For those which you judge positively, what might you use more? For those which seem unreasonable, how can you tell? What would you recommend avoiding? - Ask some general questions. From the answers, ask more-refined questions. Follow leads. Report on your conclusions. Also, review your questions and describe the next place to explore. - In what ways does Cody’s answer reflect the file you have selected? - What does Cody explicitly refuse to give answers about? What does it indicate lower confidence in, and how? What does it typically answer confidently? Give examples, and indications from the repo about whether its answer reflects it confidence. - In what way (if any) is the ‘files read’ useful? What might it indicate? - List a few search targets. Use Cody to search, and compare with searching the repo directly. Does Cody find all / some / a few instances? Does everything it finds exist? - SourceGraph have ‘indexed’ this repo – what can Cody tell you about the repo itself? - Compare Cody’s descriptions of repo-level information with its descriptions of files, functions and tests? - Exchange questions with Cody for 20 minutes about the architecture / plumbing of cURL. Does it seem helpful? Check out what it says – is it accurate? - Exchange questions with Cody for 20 minutes about a specific file / limited function / small directory. Does it seem helpful? Check out what it says – is it accurate? - Read about cURL on external sites. Ask Cody questions which should give you information relevant to those sources. And, to get some examples of things we know that generative AIs do, keep an eye out / divert towards exploring... - Find an example of something (file, function, test, facility) which Cody refers to, but which does not exist in the repo. Build around that example; ask for details, comparisons, code, advice. Seek ridiculous examples – bugfixes in non-existent code. - Find an example of code that is offered by Cody, but which is not in the repo - Find an example of something that Cody describes 'confidently', but which is wrong. - Find an example of poor advice – if Cody suggests making changes that would be counterproductive - Generated tests and test data that (in some obvious and real sense) don't work... #### Transcript Here's a [rough transcript of my interaction with Cody while looking at the repo for cURL](https://www.workroom-productions.com/exchanges-with-sourcegraphs-cody-about-curl/). ## Week 6 – AI and Testing and Ethics... This week, we have an open conversation to grasp the necessary and knotty nettle of ethics. #### Conversational starters – if we need them Not to be done in order. - If you're testing an AI, are you testing more in terms of with aesthetics (judging good vs bad) or ethics (good vs evil)? Is this different to what we've been doing? - Miro board: add sticky notes for specific and real situations where AI has had / is having an undesirable effect. Try for several each. If you can add a reference / link, so much the better. We'll shuffle them, based on... could this have been anticipated? could this have been avoided? who could have noticed? who could have taken action? - As a tester, how do you feel about - writing a test strategy with ChatGPT? - paying attention to a test strategy written with ChatGPT? - working as a tester (finding value / seeking risks) on a product which uses AI to: - offer medical advice - identify people by their faces - write adverts - take on the appearance of a person - reviewing a general AI, as a tester, for use in one of those specific products? - using an AI to do some of your job? To replace some of your peers? - What are our responsibilities when working with AIs, teams who build AIs, and people with purposes that are fulfilled by AIs? - As testers, do we have any responsibility for the behaviour of the machines we work on? - As people, do we have any responsibility for the moral choices of the teams we work with? - Do devs (inc. testers) of AIs have a responsibility for transparency? #### Too revealing? - Where do you draw your own line on what you would / would not work on (if your personal circumstances stayed the same) - What would be ethically unacceptable for another tester to do? How would you react if a colleague did something like that? _This post is for subscribers only._ ### BlackBox Puzzles – Series 1 URL: https://www.workroom-productions.com/blackboxpuzzles_series1/ Last updated: 2023-09-13T15:43:51.000Z This 6-week series runs on **Fridays**, 12 noon London time (13:00 CET), on July 7, 14, 21, 28, August 4 and 11 2023 We'll learn about exploring, about building models of what we find, and about checking those models. I won't tell you what puzzle we'll be exploring until we get into the workshop. Signups are now closed. If you're a subscriber, email me and I'll let you know. This series is free. If you want to make a financial commitment, please [help Jacob Bruce](https://www.justgiving.com/crowdfunding/jacobstreatment) or [donate to ](https://www.dec.org.uk)[DEC](https://www.dec.org.uk). ### How we'll work - Cameras on, mics muted if it's noisy your end. - Be kind. If you're unkind, I'll mute you. - Try not to shout out your solution, but do share what you find. If you're not a subscriber, you won't see stuff below – but I'll put links in the Zoom chat and on the Miro board. _This post is for subscribers only._ ### Workshops with @workroomprds – signing up URL: https://www.workroom-productions.com/workshops-with-workroomprds-signing-up/ Last updated: 2024-01-06T13:42:58.000Z Here's the plan... I’ll run a 6-week series of *BlackBox Puzzle* workshops on Fridays, 12 noon London time (13:00 CET). July 7, 14, 21, 28, August 4 and 11\. That’s *this* Friday. Ulp. > We'll learn about exploring, about building models of what we find, and about checking those models. I’ll run another 6-week series of *AI and Testing* workshops on Wednesdays, 2pm London time (15:00 CET). July 12, 19, 26, August 2, 9, 16. > One tool a week. We'll see where the wind takes us. I've got TeachableMachine, StableDiffusion, Cody, ChatGPT in mind – you'll see that I'm tilted towards [*generative* AIs](https://en.wikipedia.org/wiki/Generative%5Fartificial%5Fintelligence?ref=workroom-productions.com). Numbers are limited. Subscribers get to sign up first. Pick a series. Email me to get your seat. I'll confirm with a link to the series page. We'll be on Zoom / Miro. Both series are free. If you want to make a financial commitment, please [help Jacob Bruce](https://www.justgiving.com/crowdfunding/jacobstreatment) or [donate to DEC](https://www.dec.org.uk). I’ll work to keep to these dates / times. I’ll give you as much notice as I can if there’s an unexpected change. Thank you for reading. ### Workshops with @Workroomprds URL: https://www.workroom-productions.com/workshops-with-workroomprds-202306/ Last updated: 2023-07-05T16:20:22.000Z ‼️ Here's the [signup page](https://www.workroom-productions.com/online-workshop-series-202306/)! I'm planning some online workshops. Each workshop will be 40-60 minutes long. A series will be six workshops long. I have several series in mind right now. I'm gauging interest, and will probably run the most-popular first. I expect to limit numbers to around a dozen per workshop. We'll use Zoom and Miro, or possibly the [TestLab Facility in GatherTown](https://app.gather.town/app/oKySvbQJhKHKNFgZ/TestLabFacilityIsOpen). ## Puzzles We'll do (at least) one of my [BlackBox Puzzles](https://blackboxpuzzles.workroomprds.com) every week. We'll learn about exploring, about building models of what we find, and about checking those models. ## AI We'll play with AIs, as testers. One tool a week, and we'll see where the wind takes us. I've got TeachableMachine, StableDiffusion, Cody, ChatGPT in mind – you'll see that I'm tilted towards [*generative* AIs](https://en.wikipedia.org/wiki/Generative%5Fartificial%5Fintelligence). ## Speaker Prep Each week, one of the exercises from this [Speaker Prep workshop](https://speakerprepday.workroomprds.com) (developed with Bart Knaack). We'll work on abstracts, content, delivery and more – and we'll work on the ideas and interests you bring. ## Exploratory Testing Coming soon, for paid subscribers. ### Sign Up for Online Workshops URL: https://www.workroom-productions.com/online-workshop-series-202306/ Last updated: 2023-07-10T23:08:23.000Z I've got two series of online workshops planned; one on the *BlackBox Puzzles*, one on *AI and Testing*. Both series will be in six parts, across six weeks. Each workshop will be interactive and focussed on systems testing. ### BlackBox Puzzles 6-week series to run on **Fridays**, 12 noon London time (13:00 CET). July 7, 14, 21, 28, August 4 and 11\. We'll learn about exploring, about building models of what we find, and about checking those models. I won't tell you what puzzle we'll be exploring until we get into the workshop. Signups for series 1 are closed. I may have a couple of spaces for subscribers (email me). [Sign up for the \*next\* BlackBox Puzzle Workshop series](mailto:jdl+BBP202310@workroom-productions.com?subject=Sign%20me%20up%20for%20the%20BlackBox%20Puzzle%20Workshop%20Series) ### AI and Testing 6-week series to run on **Wednesdays**, 2pm London time (15:00 CET). July 12, 19, 26, August 2, 9, 16\. One tool a week. We'll see where the wind takes us. I've got TeachableMachine, StableDiffusion, Cody, ChatGPT in mind – you'll see that I'm tilted towards [*generative* AIs](https://en.wikipedia.org/wiki/Generative%5Fartificial%5Fintelligence?ref=workroom-productions.com). Signups for series 1 are closed. [Sign me up for the \*next\* AI and Testing Series](mailto:jdl+AI202310@workroom-productions.com?subject=Sign%20me%20up%20for%20the%20AI%20and%20Testing%20Series) ## Logistics Pick a series. Subscribers get to sign up first (you can ask me to subscribe you in your signup mail). Numbers are limited. These series are free. If you want to make a financial commitment, please [help Jacob Bruce](https://www.justgiving.com/crowdfunding/jacobstreatment) or [donate to ](https://www.dec.org.uk)[DEC](https://www.dec.org.uk). I'll confirm with a link to the series page. We'll be on Zoom / Miro. I’ll work to keep to these dates / times. I’ll give you as much notice as I can if there’s an unexpected change. ## More to come Yes. Some free, some limited to paying subscribers, some for a one-off fee. I expect to run a 12-part series on Exploratory Testing this autumn. ### Building a Bart URL: https://www.workroom-productions.com/building-a-bart/ Last updated: 2023-02-23T00:55:07.000Z Bart and I are playing with testing AI. I've been playing with generative AI, in particular DreamBooth. So far, standing on the shoulders of giants ([1](https://en.wikipedia.org/wiki/Stable%5FDiffusion), [2](https://dreambooth.github.io), [3](https://replicate.com/blog/dreambooth-api)), I've assembled a pair of avatar generators: - [Bart generator](https://replicate.com/workroomprds/bartbooth1 ) - [James generator](https://replicate.com/workroomprds/jamesbooth1) Here's how I went about assembling Bart's. ## Testing insights Plenty. I'll share shortly. ### New Year Wishes URL: https://www.workroom-productions.com/christmas-2022-stochastic-santa/ Last updated: 2024-01-06T13:36:29.000Z I've been using StableDifusion to make New Year's cards. I used [StableDiffusionWeb](https://stablediffusionweb.com), which is currently free and requires no login. If you want to make some of your own, you should. You'll need a 'prompt'; a text description to guide the way that the tool organises its randomness. Most of these were generated with a prompt along the lines of `compelling portrait in the style of Tiepolo of a jolly old man with a red cap and a white beard`, and it's been fun to experiment with different artists, different media, different emotions. The tool generates 4 at a time, usually. Sometimes it generates fewer; it's filtering NSFW stuff out. I'm not sure whether I need to know what it's filtering, but as a tester, I'm interested. It takes anywhere from 15-40 seconds to build those pictures, at relatively-low resolution. I'll typically keep one or two from a batch. I kept something under 200 images; these 50 are the best within a relatively-consistent style, and I've split them into jolly and stern. My own filtering, then, is on Not Suitable For Publication, and sometime later I'll post samples of the ones that are different styles (photo-realistic, paper cutouts, Lucien Freud, sculptures made of leather) and differently-surprising (odd eyes / ears / fingers / mouths, inappropriate hair, becoming hat). I may even go back into the mix and post a few that really didn't make the cut (terrible skin complaints, half a head, nothing really interpretable). I've also *increased*, in some slight ways, the variation. After the first several hundred pale Santas, I found the source images' bias (potentially prompted by the inclusion of the word 'white') was impossible to ignore. Prompting with various nations around the Mediterranean helped, a bit – but the faces and poses became more-similar, the look became more artificial, the anatomy errors more pronounced. I've got a note below on a particularly noticable repeated face. So: you'll see a very few faces below which aren't clearly from northern Europe and the general oeuvre of the late-renaissance painters I mostly used for prompts. One may be able to *explicitly* move away from the biases in the data used to train the tool by changing the request, but you may get a different quality of output. Where your trainig data is poor, your output would be poor, too. This set leans on the learned skill and developed genius of a whole culture of painters. What's interesting is that the training and AI model using that training have filleted out concepts. Switching from `Rembrandt` to `Freud` produces similar pictures in different styles. Switching from `jolly` to `stern`, or `compelling` to `dreamy`, or `oils` to `HDR photorealistic ultradetailed` does – generally– something like you've asked. StableDiffusion is described as an artificial intelligence. It's also described as a stochastic parrot – something that combines things it knows, without any real sense. You need to see the wierd ones to get a sense of this; one element of the illusion of intelligence here is that I've picked images which work. What you're seeing here is 50 from perhaps 750 generated images. Still, I think it's stunning; shown one image, in isolation, I'd be hard-pressed to pick it up as a generated image – and I personally find every one of these 50 interesting. What, though, would I be studying, if I picked one to study? So, in no particular order, here are 50 entirely new images of old men in red caps and white beards. ![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/2022/11/E7383824-86F2-4D6A-9198-0C21420D6840.png) ![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/2022/11/19681AC5-49F0-4899-B027-A22202E9920D.png) ![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/2022/11/1B404C76-DD0D-479F-B8AE-6968111DCF05.png) ![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/2022/11/C419AB8F-F192-400B-842C-AFE9DBD666D3.png) ![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/2022/11/7F36F7C6-B05E-403A-A478-E5B56F036298.png) ![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/2022/11/24EE9CD5-CD69-4E96-8FF9-333FBD38725E.png) ![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/2022/11/AF6F8E74-3D3F-4C2D-823A-8A4757EB978C.png) ![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/2022/11/E02ADBB2-FF12-4CED-8D36-9857AE40E53D.png) ![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/2022/11/9048772E-91FF-4D07-AE5D-0D164040B610.png) ![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/2022/11/4CF98391-41F5-4354-9B86-844CDE301562.png) ![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/2022/11/D6B5C16E-AC27-44C4-A53F-E98BF0411D4E.png) ![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/2022/11/D6B677C3-2C7A-4F70-823F-EDBEB469C869.png) ![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/2022/11/9464C9A3-4F8E-43D4-89F6-3A63A0D2C223.png) ![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/2022/11/9AF2BE24-C017-4584-B322-02D0843E66D0.png) ![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/2022/11/92565A36-0C53-4985-ACCC-5267D431F90C.png) ![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/2022/11/8A6B4CAC-4BDA-46D8-8C5D-B1D44C435B6F.png) ![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/2022/11/2C88E17D-28B2-4E62-A4E4-792FB0D3832E.png) ![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/2022/11/3D166CEE-0874-4D39-84DA-BB6455FE91D2.png) ![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/2022/11/C79B9977-ADA4-4B1A-859B-3A5477014C4D.png) ![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/2022/11/83B57426-2595-41FF-853B-1CAF281AA8F4.png) ![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/2022/11/8E61F298-98B8-4D5C-A76B-141612BF6852.png) ![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/2022/11/CA8A5A53-A189-4C60-9867-16A4076C0637.png) ![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/2022/11/EE6048E4-F83D-41F0-A307-D6CEA765EE4F.png) ![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/2022/11/39481748-5EA3-43B6-907B-85389D2C4590.png) ![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/2022/11/54A76783-B8DC-4D4F-8CEE-FF1BA5E1DE1F.png) And from the more-stern set... ![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/2022/11/50D0BE78-F257-4FA6-83E5-FBC18FFA356B.png) ![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/2022/11/F3EAA796-CB27-46C2-BFDA-DFB63B758954.png) ![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/2022/11/C1223B69-7FB6-447F-8570-C3C985F67021.png) ![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/2022/11/F8B59E7A-8CC7-49C6-8469-2D235FF4EAEC.png) ![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/2022/11/455D0B9F-0ABD-4B08-A726-7F14903A35B0.png) ![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/2022/11/FBA3A805-6203-4796-8137-FA000C3C9CF0.png) ![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/2022/11/BBB13C5C-40D5-4C42-9D75-00F88A8A53CD.png) ![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/2022/11/FCAEB370-2FB7-48EE-B4EE-33458C9B4020.png) ![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/2022/11/524C504D-12F8-49C1-9C1B-BB61CB2BEDD2.png) ![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/2022/11/11F5C9C9-B109-41C5-A100-A05FF60AAAB0.png) ![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/2022/11/09D04B67-2157-4B8E-BE85-EE00D56412C8.png) ![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/2022/11/9E2C3F1C-7C39-4156-BE59-C5902E5CC999.png) ![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/2022/11/ECAE67A5-4665-48F3-A41C-F3686DEF6FD3.png) ![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/2022/11/72B48688-A1AA-4005-A473-826BADACF575.png) ![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/2022/11/5B05BBB4-31F4-4819-BAD3-CCC08691EDD6.png) ![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/2022/11/5270113E-EA0A-4B8B-8416-669024279B53.png) ![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/2022/11/104B2281-84F6-4F91-8C3C-26AAB41F263A.png) ![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/2022/11/CB66AC49-D713-4017-8D14-77CD39FC3455.png) ![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/2022/11/272CD2AE-6B7E-495A-97C4-AB5821013A21.png) ![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/2022/11/815EE467-A2AE-4C8B-B767-905075C709E9.png) ![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/2022/11/8293392F-9B93-444C-9137-568D142FE93D.png) ![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/2022/11/430C4664-74B9-46A6-9397-D951F230FC12.png) ![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/2022/11/9B431837-2BCE-4830-A88D-1EB8D8BC0E01.png) ![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/2022/11/6E0422EF-27D8-44D2-96D2-7521CFB78003.png) ![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/2022/11/E76D7D20-6FC9-42E2-90A6-82D5859B36BB.png) If you look, you'll see one face that turns up over and over again in my set (there's a name for this, which escapes me at the time of witing). It was triggered by `compelling HDR portrait by freud rembrandt holbein tiepolo photorealistic jolly old Tunisian white beard red cap`; once in that groove, about half the generated images had *that* face, and I found it fascinating enough to dig that seam rather a lot. I may go into that in a later post – for now, here are four of the same face, generated just as I wrote these words. People are currently paying several tens of dollars to get this done for their own face in various tools based on the DreamBooth variation of StableDiffusion. ![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/2022/11/image-10.png) ### A Framework for Judgement In Testing URL: https://www.workroom-productions.com/framework-for-judgement-in-testing/ Last updated: 2022-11-13T20:00:18.000Z While exploring, I use this framework to guide the way that I judge something that has bothered me. - inconsistency - external (vs an external authority i.e. a spec) - internal (vs another similar thing) - cultural (vs my own expectations) - absence: something is surprising by its absence - extra: something is surprising by its presence ### On inconsistency... Before I can spot bugs on behalf of an external authority, I need to work to understand it. In practice, this means that I need to know the spec, I need to understand the regulations, I need to have worked with an end-user to develop empathy. If I've seen this kind of inconsistency, that's easy: it's a bug. Tools are often helpful in picking out internal inconsistencies, especially where there's lots to sift through; pixel shifts in the UI, configuration differences, data entities, log weirdnesses. The art is in knowing which is the bug – but perhaps we don't always need to make that judgement, if we've got access to someone who cares. I find that it's natural (and therefore easy) to spot cultural inconsistencies, and hard to persuade someone else – or even myself. So while I might notice the problem, I might not be able to make a decision, let alone take action. *more to come* ### This Post is to look into Nutshell URL: https://www.workroom-productions.com/this-post-is-to-look-into-nutshell/ Last updated: 2022-11-01T13:45:32.000Z _This post is for subscribers only._ ### Publishing a Directory with Flask URL: https://www.workroom-productions.com/serving-a-directory-with-flask/ Last updated: 2022-10-18T15:26:36.000Z *How I serve python coverage metrics as an html page within replit.com* I need to serve a single directory – coverage metrics from a test runner – as a web page. I'm working in [repl.it](https://repl.it), and my repl comes with flask, so that's what I'll use to face the web. Serving the page isn't the primary purpose of the work; it's a nice-to-have. The work lives in `./src` and `./tests`, and repl seems to expect the page to be served while 'running', so this is how I've set up `main.py`, in the level above `./src` and `./tests`, to serve the page. ```Python ## Flask server to serve a folder as a webpage ### intended to allow me to run test coverage on the commandline in replit, and to see output as html ## Set up to use the right parts from flask ### `send_from_directory` serves a file from a directory ### I could use `render_template`, but I've got no templates. The html etc. is already made ### I could use `send_file` but this is safer as it stops directory traversal from flask import Flask, send_from_directory ## Make an app, called `app`, to be an instance of flask app = Flask( # Create a flask app - minimal? __name__, ) ## Routing – so when a browser asks for a resource, it gets something from the app @app.route('/') # empty route should serve ./cov_html/index.html def index(): return send_from_directory('./cov_html/', 'index.html') @app.route('/') #Everything else just goes by filename def sendstuff(path): print(path) return send_from_directory('./cov_html/', path) ## Boilerplate to start server if __name__ == "__main__": # Makes sure this is the main process app.run( # Starts the site host='0.0.0.0', # Establishes the host, required for repl to detect the site port=8080 # port to serve ) ``` ## Context: I'm building something for a workshop on mutation testing. It's a workshop, so I want people to work. So I need roughly as many environments as participants. I'm choosing to use [replit](https://repl.it) to allow me to to that. I need a subject under test, and it needs to be comprehensible. So I'm working in Python. I'm using `pytest` as my test runner, and I can choose to see coverage with `pytest-cov`. The commandline in repl.it is OK for commands, but at least as crap as any commandline for understanding output. \`pytest-cov\` hjas an optipon to output to HTML – I want to make it as easy as possioble for participants to look at coverage of tests, so I'll take that. The coverage as HTML is dropped into a directory – and I need a webserver to allow me to browse the information. I *can* download the directory and look, but it's a pain in the bum. So I need to serve the directory, but I don't need templating. I want participants to see coverage because mutation testing tells us about what changes to code *won't* be noticed by existing tests. If the tests don't cover some of the code, then mutation testing will tell us – but coverage will tell us more swiftly, and can help us to understand the output of mutation testing. I'd go as far as saying that if you don't look at coverage first, you're doing mutation testing wrong. ## Notes to self Add more links – to tools, and to workshop ? commnets are as gray as background? ### Odin Conference URL: https://www.workroom-productions.com/odin_2022/ Last updated: 2022-06-18T17:29:45.000Z I'll be at [Odin](https://event.dnd.no/odin/program-2022/) in Oslo on 21 and 22 September 2022. On 21 September, I'll do a full-day workshop on exploring with data; [Hands-On, Tooled-up Testing](../hands-on-tooled-up-testing/). Bring a laptop; you'll run thousands of tests, and I'll take care of all the dull work. On 22 September, I'll kick off the day with a keynote about something most of us spend masses of time on, yet don't talk about in public; [Wrangling, Debugging and Testing](../wrangling-debugging-and-testing/). ### Wrangling, Debugging and Testing URL: https://www.workroom-productions.com/wrangling-debugging-and-testing/ Last updated: 2023-09-28T13:16:59.000Z This is a talk about embarrassment. The embarrassment that happens when you’ve spent months on requirements and design, on coding and configuration, you’ve got plans, processes and people – and you’ve still got nothing working well-enough to test. This is a talk about the necessary, practical and unique work of wrangling and debugging – illustrated with nearly-true tales from years of systems integration testing. Exploring **wrangling** (getting the system to *work*), we will dig into problems such as getting the plumbing right between system components, orchestrating authorisation for tools and testers, working with configuration data and managing interlocking shared environments. We will look into ways that testers can gain technical savvy and leverage political and expert champions to unlock problems. Focussing on **debugging** (getting the system to *work well-enough to start testing*), we will look into situations where experts rely on trial and error, and when the earliest actions on the system reveal fundamental problems in performance, security, and usability. We will cover how to design bulk experiments to find working paths, how to build data for variety and volume, and how to bring organisationally-distant colleagues together to resolve trouble with cross-disciplinary insights. Projects find discomfort in uncertainty – this talk will help you to anticipate problems and mitigate their effects. ## Supporting Articles [Not Testing, but DrowningTesters whose work is to get the system to work... do we help them to work better?![](https://www.workroom-productions.com/favicon.png)Workroom ProductionsJames Lyndsay![](https://images.unsplash.com/photo-1600828576920-cb1528d7ac66?crop=entropy&cs=tinysrgb&fit=max&fm=jpg&ixid=MnwxMTc3M3wwfDF8c2VhcmNofDJ8fGxvbmdob3JufGVufDB8fHx8MTYzNjc1Njg3NA&ixlib=rb-1.2.1&q=80&w=2000)](https://www.workroom-productions.com/not-testing-but-drowning-1/) [Wrangling: Not Testing but Drowning 2It can take weeks of work to get to a testable system. Here’s why.![](https://www.workroom-productions.com/favicon.png)Workroom ProductionsJames Lyndsay![](https://images.unsplash.com/photo-1583300376343-98f8b54b1fd7?crop=entropy&cs=tinysrgb&fit=max&fm=jpg&ixid=MnwxMTc3M3wwfDF8c2VhcmNofDE3fHxsb25naG9ybnxlbnwwfHx8fDE2MzY3NTY4NzQ&ixlib=rb-1.2.1&q=80&w=2000)](https://www.workroom-productions.com/not-testing-but-drowning-2/) [Debugging: Not Testing but Drowning 3Your system is working, but it’s on fire. You reckon it’s down to your test config / environment, not a mistake in the code.![](https://www.workroom-productions.com/favicon.png)Workroom ProductionsJames Lyndsay![](https://images.unsplash.com/photo-1583911026798-907ec9f2241b?crop=entropy&cs=tinysrgb&fit=max&fm=jpg&ixid=MnwxMTc3M3wwfDF8c2VhcmNofDF8fGNvdyUyMHZldHxlbnwwfHx8fDE2MzY3NTc1NjQ&ixlib=rb-1.2.1&q=80&w=2000)](https://www.workroom-productions.com/not-testing-but-drowning-3/) ### Hands-on, Tooled-up Testing URL: https://www.workroom-productions.com/hands-on-tooled-up-testing/ Last updated: 2023-09-13T15:38:55.000Z In this workshop, you’ll work with a simple system, taking several different ways to explore its true characteristics. You’ll dig into bugs, release notes, interfaces, configuration and changes, running thousands of exploratory experiments to reveal and understand the system’s behaviours. Starting with a single simple field, we’ll design small tests, and use recognised techniques to grow them into collections. We’ll look at equivalences and boundaries in input and output, at ranges and distributions, at collections to explore validation and at ways to manage the results of bulk testing. Moving to more complex elements with several interrelated fields, we’ll build exploratory data with combinations and refinements, looking at ways that we can manage the generation and interpretation of our tests. Participants will gain direct experience of test techniques suitable for exploring behaviours with bulk tests. Exercises are built to suit new testers, experienced testers, and people who manage testing work. This is a hands-on workshop; participants will need a laptop. To allow everyone to get to grips with the testing work, all the data generation tools and test runners will be supplied; participants do not need to read or write code, not to install tools. To allow ease of access, the test tools – and the system under test – will run in any modern browser. However, the content of the workshop will only deal briefly with presentation in the browser, and will dig deep into the behaviours of the underlying code and data. ### Working with Answers to Open Questions URL: https://www.workroom-productions.com/working-with-answers-to-open-questions/ Last updated: 2022-06-18T17:38:24.000Z When I've not got specifc questions, I prefer to use more-exploratory *requests* rather than *questions*. I'll say “Tell me about testing” rather than ask “Do you test?” What I get back is likely to be unstructured, full of gaps and irrelevancies, and phrased in local dialect. It’s harder to process than an answer which conforms to my structures, matches my bias, uses my terminology. I choose to spend the time and attention processing answers to open questions because I think that such answers are likely to tell me something true, and which I don’t know. If I want to make that value accessible, I need to make the procesing easier. How do I (currently think I) do that? I **prepare for processing** – I’ll generally have a couple of questions that I can start with. I’ll know some of my model in advance. I’ll try to read into the person I’m talking to – have a look at what tasks they have on a board, what contributions they’ve made to documents and to repositories, what meetings they run or regularly attend, their recent posts in company social stuff, their LinkedIn. I’ll have thought about how I’ll process their information, so that I can b clear with them. I **take notes**. I used to take notes on paper, and I prefer to, but my paper notes get lost. So I type as I listen, looking at the person I'm talking with. I **mark my notes**. I can’t mark with the same ease as when I’m scribbling, but I want to be able to glance back and see what’s needs a further question (marked `??`), a double-check (marked `!!`), what is a challenge to my assumptions and model (generally `!` or `**`). I **read my notes**. After a conversation, I take time to read and process. That means that when I book a 15 minute meeting with someone, I’ll give myself another 15 minus immediately afterwards to read over, fix any autocorrect awfulness, and summarise. At the moment, I’m separating things that are said and I want to record (\`described\`), things I’ve perceived(\`observed\`), things that seem to have evidence (\`inferred\`), patterns which seem to underly the conversation (\`working model\`) and things I want to do (\`action\`). I **update my model**. There’s no point in asking if you’re not going to listen. If I want my models to get better over time, I have to recognise that they are not yet as good as they could be. If I believe that no one is logging usability bugs, then I meet someone who’s logged a bunch, that belief can die. If A and B don’t mean the same thing when talking about ‘negative testing’, I need to manage that disparity myself, rather than impose my model on them. I **think about the conversation** – what did I learn? What did I miss? What challenged my assumptions? What do I want to remember? I **visualise my model** – it can be easier to add detail to a model expressed as a picture. Drawing several different pictures of the same model gives me different perspectives, and helps when a fundamental revision requires a new visualisation. At some point, I’ll typically **come back to aggregate** – sometimes I’ll know how I’ll need to do that in advance (so I can prepare), and sometimes it’s a surprise (so I need to work with what I’ve got). If the tech I’m using allows it, I’ll tag notes so that I can easily look for conversations about particular projects or ideas. I’ll often read my summary and from there I can go back to the information in the notes. ### Question Chaining URL: https://www.workroom-productions.com/question-chaining/ Last updated: 2022-06-18T17:35:05.000Z Here are a few ways I chain questions together **Move from facts, to reasons, to speculation**: “What happened when you sent the file for ingestion?” “Were you able to identify the specific differences in that file, and to link it with any logs?” “why do you think it behaved like that?” **Move from speculation to action**: “why do you think it’s not getting through?” “Can we try a different proxy, or a different tool?”. **Narrative** \- “what led to this situation?” “what happened next?” “How did we get here?” “ What are your options?". One might see this example as a time-going-forward collection. Constrast with the "[Five Whys](https://en.wikipedia.org/wiki/Five%5Fwhys)" approach, which goes the other way, from observed symptom to possible cause(s). **Consistency** – as I build a model of what’s going on, I try to note what no longer fits, what has changed at a significantly faster/slower rate from the rest, what is extra or missing. **Specifics from Generalities** \- as I hear general information, some might trigger my interest: “You said that you prioritise by risk – can you tell me about the risks that have influenced this release”. ‘Obviously’ is a great detector word for these. **Get Examples** \- “could you show me an account with this issue?”, “do you have a document for that?” **Some details are big questions** in themselves - “Tell me more about finding that you had different properties for the same identifier” ### Question Transformation 2 - Refocus URL: https://www.workroom-productions.com/question-transformation-2-refocus/ Last updated: 2022-06-05T09:07:02.000Z Following on from [Question Transformation 1 - Changing the Question Word](../question-transformation-1-changing-the-question-word/) All these are for the purpose of getting information about the thing you're asking about. If you're asking a general question i.e. in a job interview, or in a forum with an expert who's not in your context, then you're asking for different purposes than as a tester. We'll not deal with those, here. ## Changing the Pronoun "you" is so... othering. Switch the pronoun. > How do **we** track requirements? Indicates that this question is important to you because you both belong to the same group, and it's important to the group. > How do **I** track requirements? When asked out loud of another person, switching 'you' for 'I' is generally a straightforward request for instruction or facts. Switching the pronoun indicates that this question is important to you because you are about to take (some of the) responsibility for doing the action, or perhaps lets you take the focus off the person you're asking. When asked of *oneself*, it makes the question speculative, triggering the imagination and perhaps looking into the future. > How do **they** track requirements? Invites reportage, which can include judgement and speculation. ## Adding detail Here, I've added detail by including a limted example. > Can you show me what changes to requirements have been recorded over the last two sprints, and what effects those changes have had? > Can you show me the tool you use to track requirements? Note that these two questions have become requests. ## Specific vs General > How do we currently track requirements, here? > How are requirements generally tracked? ## Change a YesNo from contextual to personal Change "Do you track requirements?" to "Do you know whether requirements are tracked here?" to ## Placing in time / sensing trends > How were you tracking requirements back then? > How are we currently tracking requrirements? > How will we be tracking requirements when we get to that point? > How do you imagine we'll be tracking requirements? > How has the way we track requrirements changed? ## Inviting imagination > What's your ideal way of tracking requirements? > How could we get better at tracking requirements? > Do we need to track requirements? ## Making it personal Questions can get a bit abstract – particularly with 'you' in English, where it isn't always clear whether the question is about "you" as in the individual, or "you" as in your organisation. > How do you personally track requirements? > How do you feel about tracking requirements? > How do you feel about the way you track requirements? Compare the following transformations: - "do you track requirements" (may be personal, may be the team) - "do you track requirements yourself" (personal responsibility) - "do you need to track requirements" (an imposed task) - "do you need to track requirements yourself" (about delegation) - "do you feel the need to track requirements yourself" (a self-imposed task that is not delegated) ## Clarifying intent I may not immediately reveal my intent, but knowing why I'm asking a question will help me with follow-on questions – and those are likely to swiftly reveal my intent. If I'm purposeless, that will come across just as clearly. ## Things to avoid Be conscious about imposing a model. [Clean Language](https://www.cleanchange.co.uk/clean-language-questions-of-david-grove/) is a helpful (and inspiring) tool. If asking about feelings, ask how someone *feels*. Asking whether something is liked / disliked is narrow – and if I want to be that narrow, I get more by being specific; is something helpful / unhelpful, easy/hard. Avoid questions which ask someone to rank, if they've not already done so. You might ask for unusual examples, if you 're looking for outliers, but perhpas not best / worst. People tend to remember and sleect "*a recent*" more easily than "*most recent*", "*an excellent*" more easily than "*the best*". "*What do you dislike most about tracking requirements*" imposes a model, a judgement and asks for ranking. Prefer "*How do you feel about tracking requirements*", ## What I do Of these transformations, I typically : - Remove the interrogative to ask an open-ended question > Tell me about requirement tracking - Move to specifics > Tell me about how you currently track requirements on this project - Clarify intent of question for myself > Tell me about how you currently track requirements on this project (because I think that it's not working as well as it could do) ### Writing Tools URL: https://www.workroom-productions.com/writing-tools/ Last updated: 2022-05-23T21:37:42.000Z My handwrting and paper filing system gets worse, and most of what I write beyond jotting now goes via a keyboard. An [MX Keys](https://www.logitech.com/en-gb/products/keyboards/mx-keys-wireless-keyboard.html), generally. Here's my current set of tools, and their purposes: [OmmWriter](https://ommwriter.com) – for when I want to concentrate on writing, and only on writing. It looks lovely, has satisfying clicky typing noises, and I love its soundtrack of worn carriages on minor European branch lines. [Bear](https://bear.app) – for when I need something to be everywhere immediately, I flip open Bear on the iPad or laptop or phone. [Apple Notes](https://support.apple.com/en-gb/guide/notes/welcome/mac), too, when my phone's managed to offload Bear. [Scrivener](https://www.literatureandlatte.com/scrivener/overview) – for the heavy lifting. With rich metadata and vastly rearrangabe parts, I use Scriv for re-structuring articles, reviewing conference submissions, and writing big stuff. Scrivener's sister app, [Scapple](https://www.literatureandlatte.com/scapple/overview), is nice for throwing words down a diagram. I'm using both less, which makes me feel disloyal. And that's becasue of the next two... [Roam](https://roamresearch.com) – for building from notes, which means pulling ideas in plenty of different directions, and seeing where they hang together. As far as opposible thumbs for thinking go, this one has given me an all-new grip on my imagination. [Miro](https://miro.com)'s always neat. Ideas that breed in Roam go to Miro to show off. What's more, I can invite anyone to collaborate with me on stuff, in a medium that's rather more built for side-quests than a linear doc. However, if I'm being realistic, while I use all of these extensively, I write for this website *on* this website, and I'm not certain why. So, last and not at all least: [Ghost](https://ghost.org/help/using-the-editor/) – becuase I can publish straight from Ghost, I find myself working in its editor more and more. Yet it's not markdown, has fuzzzzy search, limited layout, leaves my typos unmarked, doesn't do change control, doesn't facilitate links to my other stuff, has iffy undo and if I happen to delete something I want, it's gone. I'm not sure if I use it out of ease-of-action or ill-dicipine. Maybe both. ### April stuff at Workroom Productions URL: https://www.workroom-productions.com/april-stuff-at-workroom-productions/ Last updated: 2022-04-22T21:19:24.000Z I built a [backlinks](https://www.workroom-productions.com/backlinks/) facility for the site – if I *explicitly* link one post to another, the linked page will *automatically* link back. [Backlinks are important](https://www.workroom-productions.com/why-are-backlinks-important-to-me/) to the way I'll build this site: bringing together old and new content, and trying to **stitch it into something fun to explore**. Use the "*what links here*" button at the end of the text on most pages (and let me know in the comments if it's weird for you). I bought in a bunch of stuff about **exploratory testing**: a rant on [Exploration without Tools is Weak and Slow](https://www.workroom-productions.com/exploration-without-tools-is-weak-and-slow/), and updates to the piece on [Exploring without Requirements](https://www.workroom-productions.com/exploring-without-requirements/). I tuned my notes on [Exploring the BlackBox Puzzles](https://www.workroom-productions.com/exploring-the-blackbox-puzzles/) and picked up something mostly from the archives on [Why Exploration has a Place in any Strategy](https://www.workroom-productions.com/why-exploration-has-a-place-in-any-strategy/). On the side of **exercises and teaching**, I posted an imagination-based exercise [Becoming Coverage](https://www.workroom-productions.com/exercise-becoming-coverage/). I wrote [Teaching Exploratory Testing with Code](https://www.workroom-productions.com/teaching-exploratory-testing-with-code/) about the way that I do it, and ran a partly-new exercise with people from the the [Exploratory Test Academy Slack channel](https://www.exploratorytestingacademy.com/). If you'd like me to run one here for subscribers, I'd love to. [Tell me when works for you](https://www.workroom-productions.com/run-exploring-with-bulk-data/). I put up two parts of my upcoming **[EuroSTAR Tutorial on Questions](https://conference.eurostarsoftwaretesting.com/event/2022/questions-questions/)**: [Changing the Question Word](https://www.workroom-productions.com/question-transformation-1-changing-the-question-word/) and [Questions about Metrics](https://www.workroom-productions.com/questions-about-metrics/). There's more to come on this, and I'll run the workshop in [several short chunks during May](https://www.workroom-productions.com/run-questions-questions/). This trial run is free for all; [sign up to come along](https://www.workroom-productions.com/run-questions-questions/). I also added stuff about **peer conferences**; lovely, safe [](https://www.workroom-productions.com/acid-test/)[LEWT](https://www.workroom-productions.com/lewt-the-london-exploratory-workshop-in-testing/) (which I set up and ran) and brutish, scary [Acid Test](https://www.workroom-productions.com/acid-test/) (which I may run one day). April has been hectic: I'm three weeks into a long-for-me job; three days a week at a great big institution, for three months, doing something I can't talk about at a place I can't name. That work, and this, had to fit around the kids' Easter break, which included plenty of travel and an all-too-brief return to the [London Bulgarian Choir](https://londonbulgarianchoir.bandcamp.com) for a packed-solid gig in Falmouth. To celebrate that, and more, I'm indulging myself with this photo, from that gig, of my favourite non-lanyard on-stage role... **Finally: [Support the people of Ukraine with something more than CSS.](https://www.workroom-productions.com/supporting-ukraine/)** Cheers – James ![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/2022/04/77d29f8e-739d-4ef2-babc-12a48a6b996c.JPG) ### Run your «Questions, Questions» workshop for me! URL: https://www.workroom-productions.com/run-questions-questions/ Last updated: 2022-04-25T14:01:24.000Z *free workshop!* We'll do stuff from my [Questions Workshop](https://www.workroom-productions.com/questions-workshop-at-eurostar-2022/), as a trial before EuroSTAR. I'll offer several 40-60 minute exercises during evenings in May. Put your name + when works for you in the comment below. I'll try to gather between three and ten people for each one. You can come to one, some or all. This is currently open to everyone. Subscribers can add their names below, and will get priority on spaces. If you want to come, but don't want to add your name here, email me. I'll update this page with dates and more. ### Exploring With Bulk Data – run your workshop for me! URL: https://www.workroom-productions.com/run-exploring-with-bulk-data/ Last updated: 2022-04-22T15:35:34.000Z Online workshop on exploring with bulk data _This post is for subscribers only._ ### Question Transformation 1 – Changing the Question Word URL: https://www.workroom-productions.com/question-transformation-1-changing-the-question-word/ Last updated: 2022-06-05T08:54:52.000Z *Background materials to the [Questions Workshop](https://www.workroom-productions.com/questions-workshop-at-eurostar-2022/)* Let's take a question, and unpick it. In this, I'm going to unpick the question by changing how I phrase it, and we'll see how that changes its use. > How do you track requirements? You might use this question when you're finding out about your new project, or you might have it as a checklist question for yourself. Perhaps you've asked it as an auditor, or as a prospective client. In all these senses, you're interested in *whether there *is* an answer*, and in *how requirements have been tracked *in this situation**. For different contexts, see the end. Let's look at some common changes. There will be some grammar, below. It needs to be English grammar, as that's my sole area of expertise. Those of you lucky enough to understand more than one grammar will be able to recognise similarities and might teach me about differences. Personally, I always have to think twice about *adjective* vs *adverb* (and *left* vs *right*, as it happens, but never *up* vs *down*), so let's hope that grammar-jargon below is minimal and mostly-correct. If you see an error, claim bragging rights and help me out by letting me know. ## Removing the Question Word Let's get rid of *how*. > Do you track requirements? This is now a [yes-no](https://en.wikipedia.org/wiki/Yes–no%5Fquestion) question – one which restricts the response to a simple choice. You ask this if you're looking for a brief and clear answer. Perhaps you've got a \[\[followup question\]\] ready to go. Perhaps you'd consider skipping straight to a follow-up question. For example, "*What requirements do you track?*" will encourage fact-based answers and open conversation, and it'll detect "none" just as well. You're also sending a signal that you're not interested in *listening* to the answer: you're giving the information that you're not (yet) going to put the effort into parsing anything at all complex, and that you've already got a model that is guiding your interview. > Tell me about tracking requirements. This is now an open-ended **request**. You've passed the conversation over, and indicated that you're going to listen to the answer. With so little guidance about your expectations, you'll need to work to parse the answer. Worse, you're open to misinterpretation, especially if you have a different understanding of vocabulary and intent. I asked (something like) this once, and was told (something like) end-user expectations for tracking a parcel. If your open-ended request relies on tone of voice in conversation, you'll need to make your request more precise when writing – and readers may skip the precision anyway. Worse still, you'll need to stop someone mid-flow if you realise that you need different information. You need project specifics – you get organisational strategy or a rant on requirements. It's your choice to stop, or to hold tight until the information you need is on offer. As with misinterpretation, this situation is worse in correspondence; your responder has already taken their time to gather irrelevant information and to write an unusable answer. Some might argue that, without a `?` or a question word, it's hardly a question at all. So I've called it a `request`, here. The phrase [open-ended question](https://en.wikipedia.org/wiki/Open-ended%5Fquestion) has been taken, anyway, and means something else. Let's move on. 💣 You might keep it as a question with "Can you tell me about..?" or "Would you tell me about..?". Be aware – you're indicating a YesNo, but 'Yes' is useless. Q: Can you tell me the time? A: Yes, I can. ## Getting more specific with a different question word Question words have a simple mnemonic, and it's tempting to think of them as the fundamentals. *Who*, *what*, *where* and *when* tend to lead to specifics – and so take the assumption that whatever you're asking about is done at all. Specific answers are factual, and so tend to be true (as contrasted with hand-wavy). Listen for doubt as well as information, and see whether answers match up across different people, and different days. Though question words look simple, their *use* in questions is more complex. I find that thinking of purpose helps more than thinking of vocabulary. We're dealling with the following here, and I'll expand to a couple of examples for each. Who – personnel, or responsibility, What – tools (as in mechanism), processes Where – tools (as in medium), trigger When – what situation, how often, trigger ### Who > **Who** tracks requirements? To identify the thing which takes the action. May be individuals or roles – and let's face it, we personify systems, so can be another system, too. Can help to detect something that's not done by anyone/thing, which can be important. I've been in situations where I've asked "do you X?", and followed up with "who does X?", and the 'who' question has revealled that the answer to the 'do' question was unquestioned wishful thinking. Unquestioned wishful thinking is *Not* mendacious bullshit. The person I was talking with believed that their organisation did X, but started to realise that, if they couldn't say *who* did it, perhaps *it* didn't. You *know* how that feels... > **Who** do you track requirements **for**? To identify who cares. May be individuals or roles, often distant or un-nameable. Works as a more graspable "why". This is a question which asks about audiences / consumers / stakeholders, so can help when you're understanding how a team sees its context. ### What > **What** requirements do you track? The most-obvious follow-on question to "do you X" is "what X do you do". It's handy to use to get > **What do you use to** track requirements? *What do you use to...* finds information on tools and methods. Can help > **What do you do to** track requirements? *What do you do to...* finds information on process and people. ### Where > **Where** do you track requirements? If you can map it, you can ask "where". Asks for a location – recognise that a location can be a country, a department, a room, a tool, a specific file – even (like when) a point on a calendar, or (still like when) a project context that triggers the action. "Where do you track requirements?" "Any work on new modules". May be equivalent to > **What (tool/capability) enables you to** track requirements? ### When > **When** do you track requirements If something can switch state, then time is important and you can ask "when". You ask *when* if you want a clock time, a calendar date, a relative day or periodic event – and you're also asking... > **Under what circumstances** do you track requirements? ## Getting more general with a different question word When you're done with the facts, go wide with *why* and *how*. ### Why > Why do you track requirements? This is the big one. Flipping to *why* from other question words makes your question about values and fears, rather than about actions and tools. "Why" questions also conjures existential dread, and can lead to fights (and hugs) if used in group settings. It's kinder, more personal, and gets clearer answers to adjust "why" to: > What are your aims in tracking requirements? Asking it may give a different answer every time, especially if you're asking different people. It's good for finding values, and for checking consistency of values across teams, between departments, up and down the food chain of management. If you're asking someone who thinks by talking, and who feel that they are making a difference, "why" is an invitation to improvise and invent. The person you're talking with may have new insights as they answer; they'll own the insight with delight, and thank you (on the inside) for bringing them to it. Listen for general principles, especially ones that are new to the speaker. If you're asking someone who doesn't know, and doesn't care, and who does what they do because that's what's expected of them, "why" may put them on the defensive. Listen for deference to authority. "Why" to a physicist (when talking Physics) invites a deep dive into fundamental universal principles. Software, and software testing, doesn't have natural laws. Organiations have history, and markets have regulations. Both are important – and are human-made interpretations of actions, fears and desires. As such, the source of the "why" is the action and its contexts, the humans who fear or desire. When facing a wooly model, you should trust clear constraints over hand-waving generalisations. It's crucial to ask *why*, but you may well find that you're the one who is expected to have the answer. ### How > How do you track requirements? With *how*, we're back where we started – with a question that typically asks for a process, or instructions. *How* asks for practicalities, for processes, for the resources which enable a capability. Be careful of asking 'how' and 'why' as rhetorical questions – you may genuinely not be able to see the mechanism or the rationale, but you're also imposing your model on the question. I use an expletive heuristic to swiftly trigger my better judgement – I imagine substituting '*how the fuck'.* If it feels all wrong, I'll retreat from *how* and explicitly say *in what way do you track requirements*. There are less-sweary alternatives if you feel the need to say this stuff out lound. ### How much > How much do you track requirements? Y0u're asking for a narrative about decisions and judgement, asking about limits and context. *How much* is a great question to elicit strategy and values alongside methods. But, as it implies some sort of a scale, doesn't fit everything. Be aware, when you're asking, of the differemt iuntis and scales you may be asking for. Are you asking for a quantity, or a frequency, an absolute or a rate of change? ### Questions about Metrics URL: https://www.workroom-productions.com/questions-about-metrics/ Last updated: 2022-04-20T23:29:45.000Z **Background materials to the [Questions Workshop](../questions-workshop-at-eurostar-2022/)* ## Basics - What is this a proxy for? - Why is this a good proxy? - What does this measurement tell us? - What situations suit this metric? - Does this metric lead you to enquiries, or offer answers? ## Alerts - Under what circumstances might it stop being a good proxy? - What does this metric leave out? - Under what circumstances might it decieve you? - How might someone else use it to to game the system, or to deceive you? ## Practicalities - How do you measure this metric? - Does it need to be measured regularly to make sense? - How frequently should the measure be taken? - Is it noisy, and how do we deal with the noise? - Does it aggregate several other metrics? - Is it delayed, and what it the lag? - Does the metric very by who is measuring it? - is this a relative metric, or absolute? - What are the units? ## Interpretation - What does “going up” mean for this metric? - What does “going down” mean for this metric? - Can it go up *and* down? - What does 0 mean? - Can it go negative? What does that mean? - Do non-integer values make sense? What do they mean? - Is everything it counts as a unit the same size? - Is there a minimum? What does minimum mean? - Is there a maximum? What does maximum mean? - What does no change mean? - What does slow change mean? What is slow for this measure? - What does fast change mean? What is fast, for this measure? - What might it mean if doing work, over a given period, makes no change to the metric? - Can the measure change little, but what is measured change a lot (i.e. there are still 10 failures, but they're all new) - What trends are interesting? Over what periods? - Is the change, or rate of change, more interesteing than the actual value? - What other measures generally change in the same way, or the opposite way? What might be happening in circumstances where this corrleation breaks down? - Do you have a model which attempts to explain this metric? WHere does that model break down? ### Exercise: Becoming Coverage URL: https://www.workroom-productions.com/exercise-becoming-coverage/ Last updated: 2022-04-20T23:31:40.000Z A game to play with colleagues to understand coverage more deeply. You’re going to pretend to be a metric! ## Logistics A handful of people can get useful value from this in an hour. This exercise can spin on for a long time, and lose people as it goes. Consciously choose to use a timebox to get to a fruitful end together. Gather a group for a known time around a whiteboard or other shared visual space. I use Miro if we’re online. Before the workshop, make everyone aware of this text, ask them to read (and extend / correct if interested) the overview, coverage types and questions. Extra kudos if you can ## Priming - 5 mins Briefly exchange examples where someone in the group has seen or heard the term ‘coverage’ in use in their work. ## Game - split available time evenly between participants: Take turns pretending. - One person pretends to be a specific meaning of coverage (see types below) - Other people ask questions about that meaning (see questions below) - When one is done, repeat with a new person and measure This is a conversation starter. Connection is more important than consensus. This not a test about who knows most: there are no right answers during the game. So: **please bear in mind this important note about *answers.*** *Convinced answers are *fine** *Uncertain answers are *fine** *Surreal answers are *fine** **“Don’t know” is also an excellent answer** ## Closing debrief - 10 mins Pick some from: - Share surprises - Share how you might make use of different kinds of coverage. - Share how you might help people to understand what each other might mean by ‘coverage’. - Discuss whether you need a shared definition of coverage - Complain about the misunderstandings you’ve experienced - Consider what action you could take to decrease confusion / increase clarity around ‘coverage’ as used in docs and conversation. ## Resources ### Coverage - overview Coverage is a proxy for Done-ness - people use it with the aim of assessing whether / when they stop. It is an assessment of the work, not (generally), a measure of the end product or its quality. However, while coverage indicates doneness of *work*, coverage trends can tell us that the *product* has changed, or may tell us that we have adjusted our idea of the *problem.* Coverage can go down as well as up – this is typically because we realise we need to work more than we had expected. This is good news, because we’ve learned necessary things about out problem. It is also bad news, because we while the shape of the problem is more clear, the solution is further away. Coverage means different things to different people, especially with different senses of ‘done;. Those differences may not be clear. Look out for situations where ‘coverage’ is passed around as a generally good thing, with no actual measure attached. A particular test technique may imply a specific kind of coverage (i.e. unit testing w/ lines of code, req testing with requirement coverage). Some authorities suggest that you should measure goodness of your testing by a different measure from the one that you used to drive your design (i.e. see if your req-based tests give you much code coverage). ### Sources - add your own! [Code coverage - Wikipedia](https://en.wikipedia.org/wiki/Code%5Fcoverage#Basic%5Fcoverage%5Fcriteria) [](http://www.badsoftware.com/coverage.htm)[Kaner: Software Negligence & Testing Coverage](http://www.kaner.com/pdfs/negligence%5Fand%5Ftesting%5Fcoverage.pdf) ### Types of coverage - add your own! - code - lines exercised by a set of tests - code - decisions - code - paths (combinations of decisions) - code - files / modules / objects / functions - data elements - crud(ms) - functionality - UI elements - Pages - Menu items - System parts - Transaction types - File types - Record types - Requirements - Acceptance criteria - Service-level agreements - Fixed bugs - Edge cases - Planned tests - Risks - Personas - Account types - Character sets - Platforms (Browsers / Operating systems) - Edge cases - Known errors - Error messages - Default values - Data combinations ### Questions about coverage - add your own! - How does one measure this metric? - What does this measurement tell us? - What does this metric leave out? - What situations suit this metric? - What does “coverage going up” mean for this metric? - What does “coverage going down” mean for this metric? - Is there a minimum? What does minimum coverage mean? - Is there a maximum? What does maximum coverage mean? - What might it mean if doing work, over a given period, makes no change to the metric? See [Questions About Metrics](../questions-about-metrics/) for more. ### Exploration without Tools is Weak and Slow URL: https://www.workroom-productions.com/exploration-without-tools-is-weak-and-slow/ Last updated: 2023-05-24T09:12:28.000Z *A rant, which I should have written years ago. Now's as good a time as any. Consider this unfinished until I've wrestled with it more: v0.2.* This is a core idea in the way that I practice and teach software testing. It seems *obvious* to me. So obvious that it hardly bears remarking. And I need to say it in pretty much every workshop, and to pretty much every client, at some point. So *I'm* the one who doesn't get it. And therefore, I find it hard to explain. Two fingers to the simple metaphors of *exploration*. Let's consider exploratory *testing.* ***You* are the tester. Not the tool.** *Your* speed and throughput of testing is dictated by your limits as a human. Is that human limit to do with being *slow at typing*? No. So get good at triggering your systems at the speed at which they can be triggered. Is that human limit to do with being *slow at designing*? No. Test design takes time, but can be BIG. A single obvious test can have thousands of interlinked moving parts. So get good at building tests which cover, rather than peck. Even for those tests whch are designed in response to what you've just seen. Is that human limit to do with being slow at *reading the output*? No. You can design a visualisation which can show you the results of millions of tests at once, use your extraordinary predator brain to pick out the patterns, then use the visualisation and your brain together to slice through the data and build an evidenced model to verify or refute your intuition. I watch some people exploring, and their clear limiting factor is the one-by-one experiment. Which is often a joy to do: in the best moments, it works like a precise karate move to the vulnerables, followed by a satisfying folding-up of the system under test. But that one move typically comes after several puzzling and unedifying ones. Your human limits are the limits to your imagination, your judgement, your cunning. For all that other stuff, there's a tool. If you're going to systematically push through some iteration over the real behaviours of a system, you need a tool to work at the speed of that system. You need a tool to set up every condition in a neat and reliable way. See it as a bonus that the tool will mean that you don't grind your mind to soup with the tedium. You need a tool to let you look through thousands of pages of aggregated crap like the omnipotent consiousness you are (by comparison with the machine holding the data). **Get off your arse, wiggle those magic fingers, and use a fucking tool already.** And, hold up, not one of those tools that does what a user does, but stupider. You need a tool you can drive, not watch. You need a data generator, a comparator, a grapher, a validator, a parser, plumbing and adaptors, data and data furklers, collections of rotten files, schedulers, sequencers and triggers. It's not going to be one tool (unless, at a pinch, it's Excel) but a quivver of tools, an arsenal, a pack. Each one a slice of superpower – and you're the one to weld all these things together into a cobbled-up one-off chimera to set up you, you tester, for that one slick move that leaves the system under test neat filleted and showing off the next flaw to be fixed. Go on, get to it. **The days of hand-cranked exploration are *done*.** ### Teaching Exploratory Testing with Code URL: https://www.workroom-productions.com/teaching-exploratory-testing-with-code/ Last updated: 2026-02-26T09:57:23.000Z *Copied mainly from a message on slack. Now has some links.* So, over on the [Exploratory Test Academy Slack channel](https://www.exploratorytestingacademy.com), [Maaret Pyhäjärvi](https://visible-quality.blogspot.com) wrote > I hear [@James Lyndsay](https://exploratorytesting.slack.com/team/U02ELPSDG8P) has been teaching exploring with automation, and was curious on how you ended up framing the session. I am really curious on the experiences of actively bringing code into space of exploratory testing, since so many people are framing it as learning on UI level and the jump to thinking in terms of what automation enables may be a significant one. and I answered (something like): ## How do I teach exploration with code? It's good to set the scene. As part of that interaction, I set out my relevant opinions, which help to frame the exercises. They are: - [exploration without tools is weak and slow](../exploration-without-tools-is-weak-and-slow/) (and dull) - tools give explorers powers of bulk data generation and of data analysis and of iteration and of reliable checking. All these are computational, so are better done by a computer than a human intelligence (so that's strong and fast and dumb covered) – and if they are, there's more for your human intelligence to grapple with (which seems like fun). - automating a fixed user journey relies instead on less-automatable work such as test design, broad-bandwidth observation, judgement I work with a couple of exercises to explore those ideas. - A [puzzle](../tag/puzzle/), or recently [raster reveal](../raster-reveal/) to look at exploring the behaviours of a software system. I use this to try to give participants a recent experience of having a revelation; a novel hypothesis based on aggregating information, rather than confirming pre-existing expectations. - A [tool which does simple conversion](http://exercises.workroomprds.com/converter%5Fv2%5Fgen/); number in, sentence out, along the lines of 4 -> “4m is 400 cm”. I use this to try to give participants an interactive experience of different ways into an artefact (behaviour, code, config data, tests, fixes, release notes). I also provide a bulk input facility – indeed, three; one which takes anything, one with pre-generated data, one with a data generation tool. I ask participants to design their exploratory tests to do many tests at the same time, to judge the results in aggregate / spot unusual oddnesses, then to dig in to surprises. Using this, their key automation is in data generation, with judgement to ‘eyeball’ any oddnesses in the output. - I consciously don’t set up an exercise to have an automatable UI, and I don’t ask participants to automate sequences of actions early in a workshop. I do that later, automating sequences of similar actions. I realise that I don’t think I’ve ever had an exercise in which participants are asked to automate login / data entry / check entered data – yet this is what I see many tests doing at clients. Perhaps I need to change. As we play, we'll highlight experiences of the group either in this exercise or in their work. I try to find moments to touch on: - the difference between long-lived automation expected to verify value over some fraction of the life of a product, and short-lived automation to reveal surprises - how existing automation can be re-purposed for exploration (i.e. take fixed examples and parameterise, switch bulk tests to approval tests, measure and aggregate/take trends of performance) - supporting judgement for automation – either in building judgement and using tools to get to a point where one can use it, or taking test approaches built around automatable judgement (fuzz tests, approval tests) - allowing themselves to ‘cheat’ - what ‘completeness’ might look like – and how one’s sense of ‘complete’ changes one’s aim and approaches - building one’s sense of product risk, and how that might inform one’s testing and choice of automation - automation within exploration as enabling experimentation --- Does that sound fun? Does it sound useful? If so, then know that's the kind of thing I teach. This bit has been part of what I teach since my \[\[Diagnosis Workshop\]\], and has existed in something like this form since \[\[Bulk Testing and Visualisation\]\], and has been turing up regularly in \[\[Insights into Exploratory Testing\]\] and some other workshops since. I ran these exercises in late March 2022 for infi.nl (see [Lee and Veerle's video](https://www.youtube.com/watch?v=n0KRmbs2HMI) for a review) and for the ET Slack channel. Subscribers get teaching notes. _This post is for subscribers only._ ### Range Inputs for Puzzles and Machines URL: https://www.workroom-productions.com/range-inputs-for-puzzles-and-machines/ Last updated: 2022-04-01T12:50:15.000Z My flash stuff used loads of sliders, for input and output. My CSS stuff doesn't – there's a browser-native slider, and I want to ue as much as possible, but it's not straightforward. Different browsers work differently, vertical/angled sliders are hard, slider annotations (ticks and numbers) are hard too. So here's a working one. Finally. For now. ### Ukraine CSS Heading Gradient URL: https://www.workroom-productions.com/ukraine-css-heading-gradient/ Last updated: 2022-04-01T12:21:54.000Z What's the line? ```css h1 { background: linear-gradient(0, rgba(255,215,0,1) 0%, rgba(255,215,0,1) 50%, rgba(0, 87, 183,1) 50%, rgba(0, 87, 183,1) 100%); } ``` Do it – and add the link, then do more. [DEC Ukraine Humanitarian AppealPeople in Ukraine are fleeing their homes to escape conflict. They need food, water, shelter and healthcare. Donate now.![](https://donation.dec.org.uk/favicon.ico)Disasters Emergency Committee![](https://donation.dec.org.uk/images/appeals/ukraine/HERO_DONATE_share.jpg)](https://donation.dec.org.uk/ukraine-humanitarian-appeal) ### Supporting Ukraine URL: https://www.workroom-productions.com/supporting-ukraine/ Last updated: 2022-04-01T12:17:50.000Z A [single line of CSS](../ukraine-css-heading-gradient/) puts the Ukrainian colours behind all `H1`s. Simple to do, gets in your readers' eyeline on every page, and makes fuckall difference to anyone who is hurting right now. So let's help the people of Ukraine with more than colours. [DEC Ukraine Humanitarian AppealPeople in Ukraine are fleeing their homes to escape conflict. They need food, water, shelter and healthcare. Donate now.![](https://donation.dec.org.uk/favicon.ico)Disasters Emergency Committee![](https://donation.dec.org.uk/images/appeals/ukraine/HERO_DONATE_share.jpg)](https://donation.dec.org.uk/ukraine-humanitarian-appeal) ### LEWT: The London Exploratory Workshop in Testing URL: https://www.workroom-productions.com/lewt-the-london-exploratory-workshop-in-testing/ Last updated: 2022-04-22T22:05:55.000Z [LEWT](http://oldsite.workroom-productions.com/LEWT.html) *ran from 2005-2012 ish. This is how I described it, then. I'll follow this article with some less-circulated stuff that I wrote when people asked how and why I helped LEWT to work.* *You need to know that anybody was welcome, until we'd filled the room, and that we facilitated the conversations, not the content.* LEWT is an exploratory peer workshop. We take the view that discussions are more interesting than lectures. We enjoy diverse ideas, and limit some activities in order to work with more ideas. Currently, the workshop is structured as a series of short talks, each followed by a longer discussion. The workshop is one day long. Most participants will make a short presentation, and talks and discussions are time-limited. New talks can be added at any time; participants prioritise the talks as the day goes on. Attendance is by application and invitation. People who have been to the previous LEWT have first claim on seats. Attendees are expected, but not required, to have a brief talk. We have space for two people with less than two year’s experience – they’re not expected to have a talk, but are otherwise full participants. We share the costs of room and food, and no one charges or is paid for their time or expenses. LEWT is run along approximately the same lines as LAWST™, particularly regarding intellectual property and publication. However, LEWT is not LAWST™. A number of the ‘basic format’ guidelines in the introduction to the LAWST™ handbook are superseded by LEWT’s local guidelines. The LAWST™ handbook is currently here: . [AST](http://www.associationforsoftwaretesting.org/)have a [LAWST™ page](http://www.associationforsoftwaretesting.org/drupal/lawst). Major differences between LEWT and LAWST™: - LEWT has many talks, and limits discussion in order to move to the next talk. - LEWT is an exploratory workshop. We aim to improve our understanding by sharing and discussing our experiences. The workshop does not necessarily share LAWST™’s aim to ‘crystallise conclusions, rules or techniques’. - Most LEWT attendees present a talk and answer questions. - We don't have a content owner - the group owns the content. - Some spaces are reserved for testers with less than two years experience. **What's the process?** LEWT has attracted interest within the testing community. This is a brief summary of the preparation for the workshop, and how the day is run. **Before the workshop:** - Everyone submits a title for a ten-minute talk. - We have a project site, where we can discuss ideas before and after the event. - It’s helpful to have abstracts for talks, and a brief biography. **During the workshop:** - All the talks we haven’t yet heard are stuck on a wall. Participants can add new talks at any time. - Everyone gets a limited number of sticky dots. These are votes – you vote for the talks you would most like to hear. Votes may be cast throughout the day. - The day is split into 90-minute sessions. Before each session, and with attention to the vote and the flow of the day, the facilitator choses a group of three talks to be covered in the 90 minutes. - A talk gets thirty minutes. The speaker present his or her ideas for ten minutes at most, preferably less. The rest of the time is spent on questions. When time is up, we move to the next talk. - During the talk, focus any questions on clarification. Leave most questions until the discussion. - The facilitator will handle the question queue during discussions, and keep track of time. **After the workshop:** - We’ll post recordings etc. on the project site. - Papers using ideas from the workshop should acknowledge the workshop and list participants. **Refinements:** - Times may change – typically reducing. - We will go to the next item if there are no more questions. - On request, we can move to the next item before time, or can add five minutes to the end of the questions. These decisions are taken collectively (action and majority are left to the facilitator’s discretion). - You can vote more than once for a talk. - You can vote for your own talk ### Acid Test URL: https://www.workroom-productions.com/acid-test/ Last updated: 2022-04-22T21:59:29.000Z *An idea about a peer conference.* I love peer workshops. And, to be clear, I see a peer workshop as an event where you take an explicit decision to treat everyone as your peer. I set up [LEWT](http://oldsite.workroom-productions.com/LEWT.html) on that basis in 2005, but after a while (2013?), though we were all fond of the thing, it kind of faded away. When I decided that LEWT really didn’t work any more, I took the principles of a lovely peer workshop, made an opposing format, and showed some colleagues. Many loved it. Me too. But I did nothing with it. So here it is, pretty-much I left it in 2016\. It’s not self-consistent, and though most elements are rational, I’ve not included my rationale. If you find it interesting, then it’s done its work. I may run it one day. Or not. My plan was to run it once, get kicked out, and watch as it careered away. Showering sparks, with any luck. --- ## Acid Test Peer Workshops *For testers who are speakers, and want to be better speakers.* ### Opening Description **Acid Test** is a series of winner-stays-on peer conferences. Winners get kudos. Acid Test is not a "safe" space. Some of you will lose, and you enter here knowing that. You will not lose friends, here, but you may take a ding to your status. That's the price of participating. It will, perhaps, help you to change. It might not; it might just sting. You have some skin in the game. Everyone here is deserving of respect for taking on that risk, for accepting whatever reward or learning there might be. The purpose of this is to raise our game – so we need to have stakes that can hurt. Winning, and losing, is public.. Acid Test is open to all *except* recent losers. You show us what's on your mind, *and* how you communicate it. We reward content, delivery and bravery. ### Setup We need a room, and an ante-room. Everyone brings a talk, and we have enough time for everyone to do the talk. We go in random order. Everyone gets 10 mins to present. We take 5 mins to vote, 5 mins to debrief. We do two talks (40 mins), break to the ante-room, do it again. ### Business We score each person's talk for **content**, **delivery**, and **bravery**. Voters cast asymmetric votes \[0, 1, 3? 0, 2, 5?\] - 0 should be your most-common score. 0 is for "alright". We don't have a score for "terrible" - You might not expect to give a 3 more than once or twice in an event - We may chuck all the votes from those voter(s) who awarded the most points, and those voter(s) who awarded the least. The process for doing this will be made public before the event. - Votes are *not* anonymous. Most will be, but several votes per talk, picked at random, will be revealed and the voters required to explain their vote. We may defer this until the end of the day. At the end of the day, we will publish: - Scores for those who scored - The aggregate voting record of each voter - ? Who voted "3" and what they voted for. Speakers with a standout scores in each of (content, delivery, bravery) have a guaranteed return to the next Acid Test. If, next time, they have a standout score in the same (content, delivery, bravery), they win *and* are booted out. People with low overall scores don't get to come back for at least two years. **Post-talk debrief.** We don't *discuss* a talk in the room. We may make individual commentary on what ?worked? – but that commentary should be brief, and will be accepted by the group without reply. The speaker may comment on any problems they had. We vote, announce (provisional) scores, (maybe later) pick several voters to explain. We do discuss outside the room, in whatever groups we fall into, because *of course we do*. ### Why are Backlinks important to me? URL: https://www.workroom-productions.com/why-are-backlinks-important-to-me/ Last updated: 2024-12-03T09:00:17.000Z *Polished once* As I build this site from new ideas and old content, I realise that I don't really know how *you* might use it. I expect that the themes of this site will become pretty clear. However, lots of that content has underlying links, and I want to be able to draw links between (for instance) the rules-which-allow-freedom in exploration, improvisation, and peer conferences. When I recognise a congruence, I can link two ideas togther. Hyperlinks go one way. Backlinks go back again, without effort from me. So if I want to build an interlined cloud of ideas from whcih to form interactions and more, I need links as much as content – and backlinks are an integral element to that network which is not integral to HTML. --- I make notes in [Roam](https://roamresearch.com). When I make a note, I can easily insert a link to another note. When I do that, Roam makes an adjustment to that other note – it adds a link to the not I'm working on. (I know this isn't actually an action, but it works OK as a metaphor for here). So that means, when I open an old note, I'll often be delighted by its explicit links to other, newer things. I've not had to work for those links. These are [backlinks](https://en.wikipedia.org/wiki/Backlink) – they allow you to start at the source, and to see what *refers back* to it. They're currently (2021-2) on-trend: you'll find them in Roam, [Notion](https://www.notion.so) and more. I've just built a [thing to allow backlinks](../backlinks/) for this site. When you read about [Digital Gardening](https://maggieappleton.com/garden-history), contrasting ["a garden" with "a stream"](https://hapgood.us/2015/10/17/the-garden-and-the-stream-a-technopastoral/) (the *stream* being the firehose of twitter / facebook / whatever ) , you're typically reading about people building a network of linked information. But Ghost? Ghost's a platform for serving content. Turns out that it's pretty stream-focussed: search is a plugin, and discovery is around reverse-date-sorted lists of posts in tag groups. Linking from page to page isn't facilitated, and building networks of information isn't easy at all. Handily, though, Ghost is built to open itself, rather than to close\*. So I can chuck in JavaScript, and get that JavaScript to grab stuff from the content database, then update the page. \* = within limts – I still can't easily upload code and assets independently of a theme. **Background:** I've found value in two-way references since I first got my hands on Hypercard, in 1987\. I've *think* I've been able to have automatically-generated lists of inbound links since getting a license for Tinderbox in, I dunno, 2002? For a while, google had a `link:` operator, whicih allowed me to ask "what pages link here"? Roam is slicker, quicker, and easier to get used to than any of its predecessors. ### Backlinks URL: https://www.workroom-productions.com/backlinks/ Last updated: 2022-03-24T17:19:11.000Z I wrote a dinky thing for this site. We're on Ghost, here. Ghost serves content. In *this* page of content. served by Ghost, there's a button. hit the button to see all the pages which link here. What's happening? The button runs some JavaScript. That code flips open Ghost's public API (with a public key, there in the JS for all to see), and grabs all the content for the site, as HTML. For each unit of content (a post, in Ghost terms), the code parses out the ``nchor tags, extracts the `href` attributes, interprests those link addresses as `URL`s, extracts the `pathname`s (the bit after the website address), strips any parameters just in case, and compares with the `pathname` of the address of this page. If it finds that one of the anchors links here, it keeps a note of the post. When the code has a list of posts which point to here, it gloms up a table below the button, and inserts a reference to each of those posts. Then it disables the button, because I'm not going to write another conditional to guess whether it's already run, and otherwise you get duplicate lists. Why not run the code as Ghost builds the page? It's wasteful. Why have it at all? Because I'm building a garden – and if a path leads somewhere, someone should be able to go the other way; from destination to source. Perhaps, as the site grows, I'll need to filter the extraction of pages. Right now, though, it's the wasteful opposite of what I wanted to do with a static site, and I'm happier about that than I had imagined. Ghost is a pain for links, compared with my go-to-thinking-and-gardening tool, Roam. I'm glad to have a half-arsed mostly-working hacked-together way to give you (and me) a more free way to navigate the ideas here. Here's more about [Why Backlinks are Important to Me](../why-are-backlinks-important-to-me/). ### New Exploration Exercise URL: https://www.workroom-productions.com/new-exploration-exercise-raster-reveal/ Last updated: 2022-03-17T09:48:26.000Z Email announcement of new Raster Reveal Exercise, and an invitation to a short workshop. _This post is for subscribers only._ ### Making the RasterReveal Exercise URL: https://www.workroom-productions.com/making-the-raster-reveal-exercise/ Last updated: 2022-03-17T09:09:21.000Z *A roughly-chronological story of how I ended up with the [RasterReveal](../raster-reveal/) exercise.* This is an example of what I do when making an exercise. For this, when most of the tech is taken from / written by another, you might imagine that I do more configuration and diagnostic work (because I don't know the thing I'm using) or less (because my code is built by one gadfly mind over years). I think it's about the same, though the work is perhaps more transferrable in this case. I've called this work [wrangling and debugging](../not-testing-but-drowning-1) (link will take you to the series I wrote on it in late 2021). I seem to spend ages wrangling and debugging, yet I don't find much about it in testing or coding literature. None the less, testers do it, coders do it, and we seem to spend ages on it, whether we're building a new thing from the ground up, or integrating some massive collection of legacy kit. Let's acknowledge that, and talk about *how* we can do it. --- ### Stumbling across an exercise The exercise turned up as I was looking at charting tools. I was playing with [Paper.js](https://paperjs.org/)'s demos. The [division raster example](https://www.paperjs.org/examples/division-raster/) was interesting because it was nice to do, and because that pleasure came from revealing something. A part of my mind is – unconsciously and constantly – on the look out for things which might help me to explain something. I thought that I could use the demo to help people get a swift feel for discovering something. I felt that it would be a better exercise if the image was less well-known, and better still if it was a random image. I wanted to integrate it with this site, too. I let the idea brew for a couple of days, then put a swift demo of my own together. I chose to work in [Tumult Whisk](https://tumult.com/whisk/) – I reckoned it wouldn't work straight away. Whisk gives instant feedback on my changes, so I could get swift feedback on my changes as I discovered what was possible, and learned how to make it work. ### Building my own out of parts To find out about the assets and dependencies, I started with the browser's DevTools, and got confused. I turned instead to [Working with Paper.js](http://paperjs.org/tutorials/getting-started/working-with-paper-js/), and dropped in the script exposed by the demo. I was glad to see that the demo site's dependencies on on jQuery and on codeMirror were not needed to get the thing running locally. I changed the image to a different local image. The experience of revealing the image remained pleasurable and a little surprising event when I knew the image, so I reckoned that it was worth pressing ahead with randomness. Unsplash will serve you a random image at . For more depth (you can specify quite a lot about your random image) take a look at . ### 0th go – security matters Random images failed initially and consistently with an error that seemed, CORS-like, to regard an [image from a different base URL to be inherently insecure](https://developer.mozilla.org/en-US/docs/Web/HTML/CORS%5Fenabled%5Fimage). *Initially and consistently*? I've since re-tried it and it worked fine, though ...um... demonstrating that one needs to curate photos for a positive learning experience. If you're doing it yourself, note that within paper.js, the URL needs a `/` on the end to avoid a generic "no photo" image from Unsplash, even though the un-slashed address works in the browser. And in a further update, I tried to repro a [working random image](https://exercises.workroomprds.com/rasterRevealRandom/) only to find, once again, that I'm stymied by *The canvas has been tainted by cross-origin data*. I suspect a hidden dependency, so perhaps there's a diagnostic exercise here. I chose to move on with a curated selection of locally-served images – which also addressed my concern about putting a randomly-chosen image in front of people I didn't know while they were at work – so picked up a selection of images which seemed delightful and tried them out. ### First go – pictures matter My experience was... not great. I learned that I needed images with large contrasting areas – flowers in the snow were no good. Details necessary to make sense of the image needed to be easy to find – images which recede to the centre and images with a gradient and a small point of attention were out, so I dropped a woodland tunnel and a drone shot. Images needed to play to our reality/TV-honed visual sense, so I dropped a couple which were more abstract or surreal. Having tried it out locally on my laptop, I put the the thing live on [exercises.workroomprds.com/rasterReveal](https://exercises.workroomprds.com/rasterReveal/). I put it there because I've got more-direct control over javascript and over local media there – it's my own server, rather than this service. From there, I used an iframe to embed it on this page. I tried it out on a real person (that's what kids are for), and it seemed OK. Then I tried it on an iPad, and it was *awful*. ### Second go – devices matter On the iPad, the image was huge. Worse, the hardware responded by *scrolling* the huge image. Scrubbing with a finger, as an analog to scrubbing with the mouse, lead only to nausea. Setting the iframe to `no scroll` made a difference, as did adjusting the frame size. Yet the experience was still awful; once the scrolling/sizing grimness was cleared, it became clear that revealing the image worked very differently, and I felt confused and disconnected as things I touched revealed nothing, but image details appeared elsewhere. ### Despair – have I got an exercise after all? Wondering whether this was to do with touch events, rather than mouse events, I went to the issues list, to find several issues from a decade ago which turned on the non-equivalence of touch events and mouse events. Similar problems had been fixed more recently, in 2018, but there were still problems. I couldn't see anything to do with event *position*, though. Aware that I was using an old library in new tech, and that I was stepping away from more-common use cases by dropping it into an iFrame, I was worried that I might have hit a bug in the library. I reinforced my dismay when I checked the original example at – it too showed the same behaviour on the iPad. ### Investigation as an antidote to despair However, studying what was actually happening showed me that the area being acted on was consistently not where I was touching: it seemed to be displaced proportionally to the top-left of the picture. For example, banging away at the centre tended to divide up the lower right quadrant, and anything much outside the top left quadrant made no difference at all. Within paper.js' site, comparing their [Voronoi example](http://paperjs.org/examples/voronoi/) on desktop and iPad made it even more clear. Oddly, this gave me some comfort; perhaps it was less likely that the iFrame was getting in the way. The *origin* seemed OK – perhaps the *scales* were wrong. My laptop is an old non-retina MacBook, and my iPad has much greater pixel density... The experimenter in me wanted to write a small, controlled bit of code to as ann example. The wrangler wanted to change the stuff I already had. The wrangler won. I made the following things the same size; the source image, the canvas on workroomprds.com, the iframe on this site. Didn't help. I took a step back; a new experiment would be long, and I had no guarantee of learning anything of long-term value. Was there anything else to change? Could code help, rather than aligning numbers? The paper.js code *looked* alright, as far as I could tell. Would I need to write a new set of code and introduce a scaling factor by device? *Could* I write it? What devices would I need to write it for? As I was thinking of this, I went looking in the HTML and CSS looking for other places to influence size, and noticed the following in the HTML for the `canvas` element. ```HTML ``` I [searched](https://kapeli.com/dash) in the [HTML5 docs for canvas](https://dev.w3.org/html5/html-author/#the-canvas-element). Finding that `resize` was not a documented attribute, I looked (there is no search) in paper.js's docs – and in the tutorial section, I found [Canvas Configuration](http://paperjs.org/tutorials/getting-started/working-with-paper-js/#canvas-configuration), which said: **resize="true"*: Makes the canvas object as high and wide as the Browser window and resizes it whenever the user resizes the window.* Worth a go. ### Resolution I removed `resize`, and found to my relief that, finally, revealing the image by scrubbing the iPad with my finger worked in the same way as it worked on a laptop, hovering with the pointer. I'll look at more devices when I dare. I note from the same Canvas Configuration tutorial, that I could also - sort out the css for the frame with \` `canvas[resize] {width: 100%; height: 100%;}`, - investigate the option to manage high DPI screens, - dig into the an `onResize` event which is called on resize – and which I guess calls the `onResize` event in the paper.js code I nicked from their demo. ### Building the software into an exercise I need to remind myself, on every exercise, that people need *clear* instructions. I'm not there yet – I've made a good second attempt, but the instructions will get better once I'ev taught it more. Good instructions tend to be pithy (verbs, sequence, answering what to do), to set constraints (time, tech, answwering how long / how to do it), and to give purpose (answering why). In something like this, there are layers to purpose: the purpose someone might have as they poke at a picture is different from their purpose as they think about their work (my instructions need to cover both), and different again from the purpose they might have as they enter the exercise, or as they bring in their colleagues (I might seek to inspire rather than instruct for that context).. I imagine that people will be generally learning without me on hand, which is ideal. I kept track of questions, and of revelations which beg a question, as I've explored various pictures. I've sorted those questions, and added them to the exercise. --- ### Emergent Behaviours Here are a couple of behaviours which emerged unexpectedly as I put working code into unfamiliar contexts: - the location where I acted on the screen was not, on at least one system, the point where the system acted on the image - working code was stymied by built-in security measures I should also acknowledge that the *exercise* depended on helpful images. That may not be an emergent behaviour of the software system, but it certainly emerges from studying the system of brain-and-hand-and-context-and-software. If we test the artefacts we build to the exclusion of how they are really used, we leave part of the job un-done. ### Raster Reveal URL: https://www.workroom-productions.com/raster-reveal/ Last updated: 2025-05-29T14:28:02.000Z ## Exercise Interact with the image to reveal more. Stop when you're sure what it is. Reload for a new one. You're [*exploring without requirements*](../exploring-without-requirements/) – you're discovering what it is, not verifying that it is something. Observe how that feels. Observe how you work. Run this as a lunchtime exercise, on your own or with colleagues. Take a minute to play with a picture, then stop and talk about what you found, and how you found it. Could you generalise your approach? Would a generalised approach work on another picture? Do it again. Compare. ## Questions The following questions may help you dig into your experience. Add others to the comments. ### About the picture – and about you - When did you know what you were looking at? How did that feel? - What can you see? What is the subject? What is the context – where is the subject, what is it doing? - How did you work out what the subject is? - How did you distinguish subject from context? - What areas did you give most attention to? - What did you refine? How did you refine? - How did you know you were done? - Did you sense a diminishing return – where, while in some sense incomplete, declaring you were done was preferable to doing any more? - What was the interplay between what you were doing, and what you perceived - How did your guess (of what you were seeing) change over time? - Were you surprised by what you revealed, or reinforced? - Could you have done this better / more swiftly / more easily without using the exploratory surface (moving over the picture) that you were invited to use? What other ways of exploring did you try? What other way might you try? Are some of those ways 'cheating'? Why? ### About testing, automation and judgement You’re a sophisticated visual perceiver with decades of learning about what images mean. A machine is none of these things – but it is fast, precise and relentless. - What parts of this work would be helped by a machine? - How would you decide which parts are worth building a machine to help with? ## Related stuff Here's my article on [how I built it](../making-the-raster-reveal-exercise). ### Reactions... Soon after this went live, [Alan Richardson](https://compendiumdev.co.uk/) / [@eviltester](https://twitter.com/eviltester/) posted his response: [Amorphous Blob Exercise](https://www.patreon.com/posts/63891711 ) . I'm delighted by his take, and I've quoted part of it below. > my immediate response to these exercises is to 'cheat'. > > By which I mean... there are multiple Domains in play with this exercise: > > \- The Exploratory Testing Exercise Domain > \- Self Reflection Process Domain > \- Gestalt Psychology > \- Image Cognition > \- JavaScript Canvas Web Domain > \- Web Technology > \- etc. > > There will be more domains. An additional exercise is identifying which domains you are aware of that you can bring to play in any exercise. > > Exploring each domain provides different insights and learning experiences. The process of modelling is a learning experience. > > By 'cheat' I mean, I jump to one of the domains which will provide an answer faster, and preferably a repeatable answer. > > Having done that, I can then explore the other domains more comfortably, knowing that I have a tactical approach to getting results if I need them. [Join his patreon](https://www.patreon.com/eviltester/) to see the rest – and to read more excellent (and frequently-updated) content. Subscribe to read my teaching notes, including how I treat the opportunity presented by people whon 'cheat' an exercise. [**Helena Jeret-Mäe**](https://www.linkedin.com/in/tetris/helena-jeret-mäe-643b4339/) prompted me to set up this [version with a familiar picture](http://exercises.workroomprds.com/rasterreveal%5Fhjm/). Subscribers get to read my teaching notes (including how to 'cheat' and why it's worthwhile) and will find liks for direct access to the photos _This post is for subscribers only._ ### Teeny Tiny Test Harness URL: https://www.workroom-productions.com/teeny-tiny-test-harness/ Last updated: 2022-04-22T22:20:46.000Z The more JavaScript I write, the more I use `console.assert` to test my code. It's my go-to teeny tiny test harness. I use it in two main ways: **Friction-free TDD** – I use CodeRunner to explore short, self-contained ideas. These proofs-of-concept fit onto a page or two. As these ideas are new toys, and as I want them to be self-documenting, I build with TDD to keep me on the straight. Here's an example of how I use it on `result` output from a test call to a target function. ```JavaScript console.assert(Array.isArray(result), "should return array") console.assert(result.length === 1, "should return a single value"); console.assert(result[0].slug === "thisOne", "should return an object with right slug attribute"); console.assert(result[0].title === "this one", "returned object should have a title") console.assert(result[0].updated_at === "somedate", "returned object should have updated_at") ``` If I choose to take this forwards, I would expect to adjust it into something more clear. I'll expand on that below. **Parameter Validation** – I miss typed languages, sometimes. So I write routines to validate the input to functions, especially when I've just wasted my own day by debugging something I should have avoided. A validation looks like: ```JavaScript check = console.assert; check(Array.isArray(inboundList), "buildListOfLinkedPages has not received an array"); check(inboundList.length > 0, "buildListOfLinkedPages expects non-empty content"); check(inboundList[0].hasOwnProperty("slug"), "buildListOfLinkedPages expects first element of array to have checkg") check(inboundList[0].hasOwnProperty("title"), "buildListOfLinkedPages expects first element of array to have checkle") check(inboundList[0].hasOwnProperty("updated_at"), "buildListOfLinkedPages expects first element of array to have property updated_at") ``` Do these two examples look similar? Why yes, they do – and that is because the output of the routine tested in the first is the input checked in the second. Subscribers get to comment, and can tell me if they'd like me to head down this particular rabbit hole. ## Enhancements to the Harness If I'm moving from teeny-tiny to just tiny, I'll grab something that makes things more clear, and more easy. I've done it, a bit, in the validation above: I use `check` as a synonym for `console.assert`. I'll use the synonym `example=console.assert`, too – this helps me understand which of my tests are the fundamental examples I want to keep. I might `let startNewTest=console.log`, as syntactic sugar to help me catch what I'm testing. Once I'm using synonyms rather than the native commands, I might put more in those calling functions; to let me keep counts, or toggle suites on and off. At this point, I recognise I'm building a test harness and so turn to Jasmine or Tape or my own libraries. The point is that I can start building tests with no imports, no plumbing, no necessary infrastructure to explain to myself or others once memory has faded. ## Enhancements to the Tests These examples aren't part of polished code – I'm working in a spike. However, if this idea seems useful, I'll need to move on. My test code will need to change to support that. I'd typically take a hard look at my TDD scaffolding, keep whatever seems scaffolding-like and might help me as/when I make changes, and bin most of the rest of the cruft. I'd rewrite some non-scaffolding cruft, and add more, to act as examples. I'd write (or at least note down) some failing edge cases, to help me remember limitations that I've left in. My aim in doing this would not be to check the code's behaviour, but to use working examples to clarify what I intended to make. I'd also take a look at dependencies – perhaps I can see a way to write new tests to join a few components. My aim is to set the groundwork to explore unexpected interactions. When coding, I make unit testing easy by favouring \[\[pure functions\]\] if possible, so I don't spend much time on dependencies for my tests. I might write probes and measures. Probes to hit whatever I'm testing with generated ranges, and to look programatically for simple stuff, or to plot the output in some way that I can eyeball. Measures to capture response times and more. My aim in doing this is to understand what I actually made. More manually, I might keep track of those measures and probe outputs, so that I would have a chance of observing changes in behaviour. I might use an existing or generated test set and capture the results as an approval test. My aim here is to support my flaky memory as I make changes over time, or as I use my stuff in different contexts. I expect to head down this particular rabbit hole in a later post. Encouragement from subscribers may help me write it sooner. ### Testing Without Requirements URL: https://www.workroom-productions.com/testing-without-requirements-3/ Last updated: 2022-03-01T13:28:11.000Z *Just a placeholder, for now* ### Exploring without Requirements URL: https://www.workroom-productions.com/exploring-without-requirements/ Last updated: 2025-01-15T14:46:25.000Z *This is not finished, but roughly sharable. I wrote the articles for (some) links.* ## *Why* to explore, without requirements Requirements are handy. They're also seen as unquestionable, complete, and so necessary to testing that testing cannot start without requirements. This fetish is a barrier to starting work, particularly where the context makes it hard or dangerous to *take* a decision to start work. Which is silly, because [testing reveals requirements](../testing-reveals-requirements). And because [your requirements are wrong](../imperfect-requirements). **Testing** involves judgement, and judgement involves comparison. If you typically judge against a model built from requirements, you'll feel uncomfortable without them. A decision not to test may seem reasonable. Your decision may not be as clear-cut enough to relax into, every time: I'll write about that shortly, in [Testing Without Requirements](../testing-without-requirements). **Exploration**, on the other hand, can proceed merrily without requirements. If you're holding off looking at a system because you don't know what it's for, you're denying yourself the pleasure of exploration. Hedonism aside, your're denying your organisation the your perspectives on what that software system is building / installing / using / changing / deleting / corrupting / allowing / hindering. Get over yourself. Exploring without requirements can be uncomfortable. Recognise that for what it is; try to understand your feeling. What happens when you accept that discomfort, and get on with the work anyway? --- ## *Ways* to explore, without requirements How *do* you explore without requirements? You work with what you've got. Exploration requires an artefact. That artefact will reveal itself to you in different ways, depending on how you approach it. Here are a few general ways to explore a software system: - Explore the system as it is now, to look at how it transforms information - Explore the system as it is now, to look at its parts and their interrelations - Explore the data that the system contains, how it is represented, and what it satisfies - Explore what the system naturally does, gather evidence about the intent of its makers and its apparent purpose. - Explore the underlying technologies and component technologies of the system, their capabilities, and the constraints that those technologies impose on the system - Explore for shortcuts that the system takes to be responsive / performant / low-impact, and the ways that it might be hardened to be usable, to be secure, to be resilient. - Explore the feedback and the regulators in the system. - Explore the limits and capacities of the system and its parts. - Explore the boundaries of the system – what is inside and what is outside. - Explore the ways that the system tells others about itself – manuals, specs, contracts, error messages, API descriptions, schemas, standards. An artefact is a made thing, which means it had a maker with an intention, it has a purpose, it has expectations and history and a meaning. These qualities go beyond what it is: as how it was made, how it will be used, how long it might be expected to last, what needs to be done to keep it valuable, what might be dangerous about it. We might see those qualities as to do with the relation between the artefact and its context. An ivory chesspiece is different whether one sees it as part of a set, or part of the current game, or part of the walrus it was carved from. A book is different whether one sees it as a collection of fine-grained paper, as a collection of ordered words, as a singular part of the author’s oeuvre, as a firelighter or as the guiding urge of a cultural movement. - Explore the system as it is now, looking at how it interacts with external things – users, other systems, filesystems. - Explore the ways that the system has been constructed, installed and set in motion - Explore the system's history – how it has changed, who made those changes, and why. - Explore the system's intended purpose. - Explore the system's use to those who don't attend to its intended purpose. - Explore the system in a human context – what do people feel about it? What do they use it for? How are they disappointed? How are they delighted? How do they misuse it? - Explore the system's end-of-life: what happens to its responsibilities, to its data, to its parts, to its users, to its licenses, to its secrets? - Explore the reasons that it might reach the end of its life. - Explore the ways that its connections to other systems and its transactions change over time. - Explore the ways that it is regulated and limited, and what controls or sets the limits and regulations. - Explore ways that it interacts with cycles that are out of step / much longer / shorter than its own. You'd find analogous approaches if you're testing a document or a map or a library or a building or a process. I expect you can add to this list, from inside or outside testing. Do add a comment with your own example. I may add it to this page. **We can all explore without requirements, and we can't help but build models of what we find.** **The trick is to start.** ### Testing Reveals Requirements URL: https://www.workroom-productions.com/testing-reveals-requirements/ Last updated: 2024-01-18T13:27:35.000Z *Moving on from a placeholder* Let's agree that [requirements aren't perfect](https://www.workroom-productions.com/imperfect-requirements/). Sometimes problems get spotted only when there's a deliverable to test. Sometimes the deliverable is a working system, sometimes it's a sketch. If you're not [looking for trouble](https://www.workroom-productions.com/looking-for-trouble/), you're missing your chance. When we test, we often are the first to feel the truth about what we're building. Sometimes, what we've *found* shows us that what we've *delivered* needs to change. > **Example:* our requirements tell us to aggregate all transactions: We have 20 transactions, and aggregated 17\.* When the demonstrable truth runs counter to a clear requirement, there's an easy bug to log. Sometimes, what we've found reveals that we need to change something deeper than the deliverable. > **Example:* we've aggregated all the transactions. We notice that when we include a reversed transaction, we count the transaction and its reversal as two transactions, and we see that the values cancel out. Both behaviours match expectations. We investigate averages for sets that include reversals, and they're seem low because we're counting reversed transactions. We make an example, with one transaction of value 100, and two pairs of reversed transactions: the average transaction value is 20 for each of 5 transactions. Do we need to change how we count transactions, change how we calculate averages, or change our expectations of what the counts and averages should be?* *We want to investigate what happens to daily averages, when the transaction and its reversal are in different days? What happens when the* reversal *is reversed?* We make systems for people. Perhaps some people want a *count of all transactions including reversals*, and others want an *average value of all unreversed transactions*. If we discover this in a deliverable, then we need the business expertise to understand that this is an inconsistency, we need the organisational knowledge to understand who cares, we need the explanatory skills to take our scenario to both in a way that makes sense and the facilitation skills help them to understand each other's point of view, and maybe be part of their decision. And then someone needs to be persuaded to put up the money to pay for the change that is needed and agreed. The interesting thing here is that we need to change our decisions about the system, before we work out how to change the system. Something unexpected has been revealed about the problem. Testing has revealed a requirement. If we do what we might hope for, and discover this before the system is built, we'll still need all those skills, but perhaps the change is swifter to deliver, or cheaper and simpler to deliver. We're testing the requirements, not the deliverable – but still purposefully exploring those requirements and judging what we find. Did I say *we*? Do *you* do that? Do *all* testers? Are testers *expected* to do that, in your organisation? Or are testers expected to make a known set of simple measurements to check off requirements? Surprises come from all directions. You'll find them in the code, in the data, in system behaviours, in people's reactions. Their roots are in our assumptionad and expectations, in our requirements, forced on us by our architecture or bought in by a dependency. If we [look for trouble](https://www.workroom-productions.com/looking-for-trouble/), we'll stand a better chance of reacting productively to surprises. ### Exploring the BlackBox Puzzles URL: https://www.workroom-productions.com/exploring-the-blackbox-puzzles/ Last updated: 2023-09-18T21:29:19.000Z *This is more polished, awaiting feedback... and follow-up articles.* My BlackBox Puzzles help you explore the ways that you build and test *models*. To help people focus on model they are building, I don't give requirements. Indeed, I explicitly ask people to to work *without* requirements. Working without requirements can be counterintuitive, and is uncomfortable for some. Nonetheless, I think it’s a [worthwhile skill to gain](../exploring-without-requirements) – and the puzzles give you a safe and responsive place to play as you learn. *You can't easily *test* one of these abstract puzzles – what would you judge it against? So put that task to one side for now, *explore* until you have a model, then *test your model*.* ## Focus Yourself Set aside time to play, and decide on your purpose for that time. You'll need more than 5 minutes, and I wouldn't want to spend more than 30\. You'll want to guide your time with some purpose. If you don't feel comfortable with setting your purpose, here are some hints. ## Exploration is Play with Purpose *[Alan Richardson](https://twitter.com/eviltester) put this thought into my head. Read* **What is Exploratory Testing?** in his [**Dear EvilTester** book](https://leanpub.com/DearEvilTester) *to get closer the the source*. *I expect I'm conflating* intent *with* purpose*, here.* In exploring a system, you’re looking to develop a mental model of relationships within the system. I use and teach several techniques which lead to sharable models. You’ll have your own, I expect. If not you’ll start to develop them, right now. Here are three broad approaches which work well with the puzzles. - List components, seeing what can be worked with directly, what reacts, and keeping a record of what you observe. Many people, faced with a UI, list recognisible UI components. That's a fine place to start: You might choose a different interface. You’ll be *making a map* of *input->output*, looking at *data transformations*. Perhaps you'll see some *equivalence classes* in your records – set of inputs or outputs which behave in *symmetrical* ways. - Seek to *model states and events* – observe behaviours, consider events that seem to have an effect, look out for collections of behaviours that seem to persist together, and what makes a change. Many people, faced with a UI, look at how the subject's UI reacts to them. That's a fine place to start: You might look for reactions that aren't in the UI, and for events that you don't personally trigger. - *Look inside*, and go digging for anything that might help you in the code. You can *see structures* in the code, *read my UI labels and function names*, and *check resource use*. Maybe you'll *use the debugger* to track what's going on in execution, or *check the console* and local storage, or mess with with the HTML and CSS to *understand the browser’s perception* of what’s happening in the DOM. It's javascript, so it's all open – unless the puzzle talks to a server somewhere. Do you reckon the puzzle is *running anything remotely*? Each italicised phrase above is something to guide your exploration. Can you think of other ways you might explore? Can you coax those techniques into a shareable model? As these are games, be aware when you're stymied by a sense of *not wanting to cheat* – does a particular method seem inappropriate because it reveals too much? Why is that a problem? Would it be a problem in work? Let's be clear: **I invite you to rename 'cheating'** ***, and to do it anyway***. Being simple and purposeless, my puzzles don't respond well to several common alternatives that you might find in a commercial setting – subscribers get to see those (and to comment) below. ### Making Notes moves you from Good to Great Some testers trust their minds to recall all the salient parts, and to discard the irrelevant. I don't trust mine, so I keep notes. My notes – when I keep them – let me step out of, and step back into, my exploration; I often regret *not* keeping them when I'm bounced out of an exploration that I had carelessly slipped into. Notes help me to recall and refine more reliably, to see new perspectives, and to manage distractions more easily. If you're not keeping notes, consider how you're mananging those aspects of your exploration. ### Hypotheses Arise At some point you’ll find yourself building models and hypotheses. You may not recognise them until they are well-formed – play with the [Raster Reveal ](../raster-reaveal/)exercise get a feel for their arrival. Those which turn up before you've engaged aren't to be trusted as readily as those with some evidence. Mine tend to arise unbidden from my subconscous, and they improve when I work with them. Our mechanism of discovery influences the models which are formed, and those models in turn suggest alternative approaches. Assumptions and shortcuts may help you or lead you away – you need to choose how and when to follow them. You're making a model of a software-based system, so you’ll probably consider where you see dependencies, and whether you think that you’re seeing deterministic behaviour. You’ll wonder what history is kept, if any, how data is created / updated / deleted, and think of one-to-many relations. You’ll model what is being stored, and what is being consumed – seeing an API ( if it's available) might give some clues... ### Done? Hopefully, you feel done before the time is up (you did set a time, at the top?). If you can describe the system in a tweet, you'll probably know that you can. If you want reassurance, or congratulations, email or DM me, and I'll respond. If you want kudos, try teaching someone on your team how to solve it, in general, and see what you both learn about testing. Then go stand up in front of a group and do it again. Do let me know. If you don't feel you've achieved much in the time you've wasted, then do take a break. You may find that your massive pattern-processing kit needs a moment without being fed new stuff, so it can process what it has. Look over your notes in a day or two and see what turns up. If youve got insights, write them down in a place where you'll see them later today, tomorrow morning, and a few more times over the next week or so. That way, they'll stick and you can say that you taight yourself by playing with a puzzle. ## Testing involves judgement These puzzles are made so that you can summarise what the system is doing in a sentence or two – for my own discipline, I've described each puzzle in a 140-character tweet. I believe that each description can allow someone else to predict the behaviour of the unique parts of each puzzle. Ask me nicely, and I'll share. To reach your own summary, your exploration will need to move from measuring to condensing those measurements into a model, and you’ll need to build tests to verify that model. You'll spot symmetries and patterns, depedencies and hidden information. You'll verify your observations and dig into areas which seem obscure to you. In doing that, your're testing. **Those tests will verify (or refute) your *model*.** ## The Limits of a System If you find yourself seeking bugs in *my code*, I’m delighted – please let me know what you find and I'll fix the ones I can. However, I built these puzzles so that looking for code bugs is not (necessarily) the most interesting thing to do. I want explorers and testers to move **towards understanding the behaviours** of a system, and **away from easy bugs**. There is a temptation to declare that your work is done when you find an easy bug. If you accept that temptation, that's your choice – I regret leaving in easy bugs precisely because the stop people exploring. I’ve built these tiny systems to explore: By design, you won’t be able to judge much. But I hope you’ll feel your judgement turn on and off; it will guide your exploration, and it’s helpful to recognise when it’s doing that, as you may be being guided by an assumption. As these puzzles are built into web pages, perhaps you’ll feel that without requirements, the only thing you can *test* are web standards. You’ll act towards those areas you can judge, and you'll compare your observations with your own complex internal model – does the thing resize well, does it degrade gracefully, is it open to known attacks? Do run then puzzles through code analysers or cross-device browsers to see what standards are being broken, and what I’ve failed to code for. At that point, you’re working with my artefact, as a test subject, not trying to find the patterns in the system I’ve tried to build. It’s a subtle distinction – use these machines, if you like, to discover your preference. *This post was inspired by a question from Trisha Chetani, who asked the following question:* > how I would approach doing exploratory testing on puzzle 29 & puzzle 31? Subscribers get to see ways that *don't* work on the puzzles... _This post is for subscribers only._ ### Exploratory Testing: Lessons from Fact Checking URL: https://www.workroom-productions.com/exploration-and-fact-checking/ Last updated: 2025-02-13T12:17:58.000Z In 2015, [Mike Caulfield](https://hapgood.us) popularised the concepts of web content as stream and garden. You can see more in his [lecture](https://www.youtube.com/watch?v=ckv%5FCjyKyZY&feature=emb%5Flogo) and [post](https://hapgood.us/2015/10/17/the-garden-and-the-stream-a-technopastoral/). His primary work is in collaborative education, and you might also have chanced upon his [SIFT approach to sorting truth from fiction](https://hapgood.us/2019/06/19/sift-the-four-moves/). His open-source book [Web Literacy for Student Fact-Checkers](https://webliteracy.pressbooks.com) won a [MERLOT](https://www.merlot.org/merlot/) award. In that book, he proposes ‘four moves and a habit’ for people seeking the truth. His instructions connect with me, as an exploratory tester – they remind me of what I do, and exhort me to do it better and to explain it more clearly. He writes: > Moves accomplish intermediate goals in the fact-checking process. They are \[strategies\] associated with specific tactics. Here are the four moves this guide will hinge on: > **Check for previous work:** Look around to see if someone else has already fact-checked the claim or provided a synthesis of research. > **Go upstream to the source:** Go “upstream” to the source of the claim. Most web content is not original. Get to the original source to understand the trustworthiness of the information. > **Read laterally:** Read laterally.\[1\] Once you get to the source of a claim, read what other people say about the source (publication, author, etc.). The truth is in the network. > **Circle back:** If you get lost, hit dead ends, or find yourself going down an increasingly confusing rabbit hole, back up and start over knowing what you know now. You’re likely to take a more informed path with different search terms and better decisions. His ‘habit’, since you ask, is to respond to strong emotions with fact-checking. When we explore, we're trying to learn. We want the truth, we want it fast, and we want it to be relevant. I imagine that fact checkers have similar motivations. My thoughts are that I have a similar habit, and that I use similar moves. To put my approaches into Caulfield's terms: - Habit: **I look out for emotional reactions**: intrigue, bubbling curiosity, amusement, smug satisfaction, surprise, irritation. Those emotions (not an exhaustive set) are telling me something – and I believe that if I listen to them, I'll find out something I don't know. I make this a habit, because if I listen to my emotions and then give my mind something to chew on, I might enhance those deep responses over time – and perhaps that's one way to be a better tester. - Move: **I check for previous work** – is this something I already know? Has someone else talked about this? I can check the manual, other parts of the system under test, the bug logs, field reports and customer views. I can consider things I'm not testing, and dig into whether they'd provoke a similar reaction. - Move: **I go upstream** – I try to find out about the data, or the code, or the systems involved. I try to think about the edges of what I'm working with, and whether my reaction is to something that is within a specific part, whether it is to do with an interaction. If there's an inconsistency between what I expect and what I perceive, I wonder whether my reaction is to do with my model in my head, or the system – and if outside my head, and outside anything I can influence, what that means for further time spent exploring. - Move: **I work laterally** – typically by using diagnostic approaches to judge the causes of what I've found, and also by thinking of the potential impacts and where else similar problems might lurk. I don't follow Caulfield's actions to read what other people say, as I'll have already looked for that earlier and I'm often the first person with specific focus on the problem. - Move: **I circle back** – I look out for more emotions, here: confusion, rejection, frustration, ennui. Those tell me that I'm stuck. I might circle back by seeking out a rubber duck, by revisiting what I was doing, by putting a flag in my notes that allows me to forget, by re-assessing where I was, by clearing back to the start. I've written about this in [Handholds Framework](https://www.workroom-productions.com/handholds-framework/). 😖 **I published this here in 2022, but published it wrong, and it languished, unfindable.* ### Puzzle 36 URL: https://www.workroom-productions.com/puzzle-036/ Last updated: 2025-12-03T20:43:51.000Z ### Supported by these lovely people: **[Julie Gardiner](https://twitter.com/cheekytester?lang=en)**, **[Pascal Dufour](https://twitter.com/Pascal%5FDufour)**, **[Joep Schuurkes](https://www.twitter.com/j19sch)**, **[Huib Schoots](https://twitter.com/huibschoots)**, **[Jim Holmes](https://twitter.com/aJimHolmes)**, Ioana Chiorean, Kristine Corbus, Ide Koops, Peter Houghton, Adun Urke, Christine Yen. Help me make more and I'll put your name on the list: Use [Patreon](https://www.patreon.com/workroomprds). Close When the top circle is red, the machine has crashed. The three circles have a simple connection. Discover that connection with experiments and learn how to reliably and efficiently crash the machine. Want a hand? Here's an article on [Exploring the BlackBox Puzzles](../exploring-the-blackbox-puzzles/) Inspired by a story told by [@testObsessed](https://twitter.com/testobsessed) – watch her [video](https://www.youtube.com/watch?v=9FKY1Is0lgs). Supported by these lovely people Enjoy this? [Support another!](https://www.patreon.com/workroomprds) Built by James Lyndsay - [@workroomprds](http://twitter.com/workroomprds) © Workroom Productions 2022 ### Use Your Imagination (1) URL: https://www.workroom-productions.com/use-your-imagination-1/ Last updated: 2022-02-09T15:39:47.000Z *I wrote this for EuroSTAR's 'Little Book of Testing Wisdom', published on their 25th anniversary. Subscribers get audio!* Systems testing is fascinating – an open-ended, wicked problem that demands the truth about the surprising things we build. We need to use our imaginations frequently, and consciously. I find that, when I do, I gain insight faster, find more relevant information, and engage more closely with my customers. Here, I'm going to share some specific ways that I work with my imagination to improve my testing. I hope that, if I share my approaches, you might, too. ## I feed my imagination. My imagination loves ephemera, but my brain has a hard time holding on to all the bits. So I try to catch fleeting insights and chance details in notes and scribbles. When I'm testing I try to capture something from every document I read, every chat I have, every collection of tests, every diagram I parse. I make those notes to free my mind. Making physical marks on paper helps me to think. Making meaningful marks helps me to re-think; to check for sense, to extend into unexpected areas, to seek both precision and inspiration. Keeping meaningful yet of-the-moment information outside my head lets me forget all my half-baked testing ideas, yet use the best bits on demand. Making pictures helps me to express and refine testing ideas, and gives my brain a different opportunity to make connections and useful simplifications. Sometimes, I use props – lego and cutlery, pens and cups, to make a diagram real, to allow other people to pick stuff up, move it around. I make diagrams of system components and boundaries, of data relationships, of timelines and organisational change, of who sits where and how teams talk, of environments and shared resources, of race conditions, of lifecycles and response times and dependencies and containers. My imagination needs, and uses, many tiny clues. I build probes and datasets to dig around behaviours at the edges of what's possible, so that I can absorb some of the oddnesses, and drop them into some sort of subconscious processing that delivers insights later. I have a better chance of remembering collective behaviours if I can find a way to accommodate all the measurements into a glance, so I build pictures that aggregate thousands of test results into scatter plots and wavy lines. More in the next post Subscribers get to hear me read it... _This post is for subscribers only._ ### Questions Workshop at EuroSTAR 2022 URL: https://www.workroom-productions.com/questions-workshop-at-eurostar-2022/ Last updated: 2022-06-17T19:41:59.000Z I'm delighted to be on the programme for EuroSTAR2022, in Copenhagen. I'll run a [half-day workshop on Questions](https://conference.eurostarsoftwaretesting.com/event/2022/questions-questions/). Here are [the materials](../questions-eurostar2022/). Here's an [aggregate page of everything I've labelled 'questions workshop'](https://www.workroom-productions.com/tag/questions-workshop). [Questions, Questions - EuroSTAR ConferenceIn this video from James Lyndsay, he details what he will cover in his tutorial Questions, Questions. We look forward to welcoming you to EuroSTAR 2022 Session Speaker James Lyndsay Test Strategist – Workroom Productions, United Kingdom I’ve been working as an independent test strategist for over 20…![](https://conference.eurostarsoftwaretesting.com/wp-content/uploads/2021/05/cropped-es-fav-icon-270x270.png)EuroSTAR Conference![](https://conference.eurostarsoftwaretesting.com/wp-content/uploads/2022/01/James-Lyndsay-ES2022-Tutorial-Rectangular-Feature-Image.jpg)](https://conference.eurostarsoftwaretesting.com/event/2022/questions-questions/?wvideo=9yvc5u3dus) EuroSTAR's workshop page ### Exploring while Unit Testing URL: https://www.workroom-productions.com/exploring-while-unit-testing-timers/ Last updated: 2022-01-13T16:49:24.000Z I want to share ways that I explore what I've built, as I build it, with unit tests. Let's use a real example. This morning, I've been writing JavaScript for some new BlackBox Puzzles / teaching Machines. Here is part of it. ```JavaScript return_true_possibly: function(supplied_chance) { // to make it clear that different instances behave diffferently local_chance = supplied_chance; true_sometimes = function() { return (Math.random() < local_chance) } return(true_sometimes) }, return_true_afterDelay: function(delay_time) { returnState = function() { return state; } changeState = function() { state = true; } state = false; setTimeout(changeState, delay_time); return returnState ; } ``` This code returns *functions*, not values. So my unit tests need to execute those functions. *Note: The snippet above is roughly the version of the code that I spent my time exploring. It grew into this, as my initial confirmatory tests. Subscribe to read more about my approach to TDD in this situation. I know that these names don't currently describe the code well – and may sort this later.* ### Unit Testing for Randomness The intention of `return_true_possibly` is to return a function which returns `true` or `false`, randomly. The balance between `true` and `false` is set by a parameter. Here are my tests: ```JavaScript assert(return_true_possibly(1)() == true); assert(return_true_possibly(0)() == false); ``` It's a minor pain to test stuff that produces random numbers. I *could* explore this by running it `n` times and reporting what proportion return true. Is it worth it? Not right now. So long as I can see that it can reliably produce `true` and `false` based on the parameter set, and so long as I can inspect the code, I can be confident-enough that setting a parameter between 0 and 1 will do... something. I may test this later, and when I do, I'll pop that in the subscriber's bit below. ### Unit Testing a Timer The intention of `return_true_afterDelay` is to return a function which returns `false` until `delay_time` has passed, and then to return `true` for ever more. I got a bit more detailed with this testing – and that detail has allowed me to explore, and to learn. Here is my (initial) test (roughly-remade, may have syntax problems): ```JavaScript assert(return_true_afterDelay(500)() == false); ``` Here's a second test. It sets up a 500ms timeout, and checks it after 1000ms. ```JavaScript timeFn = return_true_afterDelay(500); setTimeout(function() { assert(timeFn() == true); }, 1000); ``` I combine those, and things get clearer. ```JavaScript testable = return_true_afterDelay(500); assert(testable() == false); setTimeout(function() { assert(testable() == true); }, 1000); ``` ### Starting to Explore I recognise that there's stuff to play with – specifically, the length of my timer, and the tolerance of the wait for it to finish. I have no idea what might typically fail, so I don't really know whether my test is useful, and by playing with it I'll have a better idea about what the timers can do in JavaScript. So I change that `1000` in the `setTimeout` to `800`. Test passes - as expected. 800ms is longer than the 500ms timer. I set it to `300`. Test fails - as expected. 300ms is shorter than the timer. ### Code to Enable Exploration I sense I could write code to make my experiments easier. I change my test code to: ```JavaScript delayTime = 500; testBracket = 20; testable = return_true_afterDelay(delayTime); assert(testable() == false); // at start setTimeout(function () { assert(testable() == false)}, delayTime-testBracket); setTimeout(function () { assert(testable() == true) }, delayTime+testBracket); ``` What does this do? It runs three tests – the first to check that the thing always returns `false` immediately. The next two look either side of whatever time I set. My intention is 1) to see how close I can get to the delay time with my checks and 2) whether that closeness depends on the delay time. A `testBracket` of 20 works with a `delayTime` of 500ms. So at 480ms it reports `false`, and by 520ms it reports `true`. I try a `testBracket` of 10\. It works. I try a `testBracket` of 1 – to my surprise, it works. That means that a test of my function at 499ms on a 500ms timer returns false, and at 501ms, returns true. There's scope for misinterpretation here. Nonetheless I am surprised to find that the tests do what I'd *ideally* expect. Obviously, it's a `testBracket` of 0 next. With both tests looking at 500ms, the test looking for `false` fails, and the test looking for `true` passes. My expectations are met, yet my eyebrows are still raised – with my initial 40ms test, I'd imagined poorer precision. I could write an extra test, looking for `true`, bang-on the time end. I don't, because I can learn as much as I need to without it. Let's mess with `delayTime`. I put the `testBracket` back to 1, and set the `delayTime` to 200ms. The test looking for `false` at 199ms passes, the test looking for `true` at 201ms passes. `delayTime` 20ms. Tickety-boo again. `delayTime` 2ms? Also passes my tests. Blimey. ### Reaching my Limit, in a stop-start kind of way Should I try setting the `delayTime` to 1? No – at 2ms, the first is already measuring after 1ms. Setting the `delayTime` to 1ms means measuring at 0ms. I don't know what that even means, so how would I use what I find? But, hey, it costs less to write-and-run than to write about. The function returns `true` when measured at 1ms with 1ms delay, and further exploring reveals that I can set a negative delay. I might choose to write some validation code in my function. I might not. I look into negative delays, and find... something which needs further investigation. More in the subscribers bit. What have I learned? No only that my function works, but also that the javascript timers (on my machine, in this toolset, on a sunny day like today) are accurate to within 1ms on short times. Short times? Let's try a `delayTime` of 10000ms. I set the `testBracket` to 1, and get bored – nothing is reported (because `assert`s and confirmatory testing – and also because I appear to have a boredom threshold of less than 10s). To see the end of the timer, I need a failing test, so I set the bracket to 0\. As expected, the test looking for `false` at 10000ms fails – but it just passed at 9999ms. Should I go further – write more tests, wait for longer timeouts? Naah. I need to write this article, then get back to building out my Diagnostic Exercises in HTML / CSS / JS. ### Decisions and Discoveries I've got unit tests for my code. I've taken decisions about how far I'm going to go with those, and that decision in both cases involves looking at opportunities and *not* taking them. I've used my unit tests to explore what I've built, linking my experiments to answer a succession of questions as those questions arise. Again, I could go further, and I choose not to. My confirmatory tests give me the confidence to refactor and reuse my code. My exploratory work has given me practical and evidenced insight into what I'm building, and what I'm building upon. I hope that helps you see how, as a tester-who-codes, I get to help my stuff to behave consistently, and I get to learn about the ways it can be surprising. --- Free subscribers get to comment, and can read more: - A note on TDD, Test harnesses, and Test Runners - More on "unit" testing random generation - Something brief on Pure functions - Reasons why I've not tested the function, and instead have tested the function-which-returns a function - Does JavaScript have a minimum length-of-timer? _This post is for subscribers only._ ### Exploring a page with Selectors URL: https://www.workroom-productions.com/using-queryselectorall-to-explore-with-selectors/ Last updated: 2022-01-12T21:34:18.000Z ### tl;dr use `document.querySelectorAll` Use `document.querySelectorAll('...your selector here...')` in your browser console to identify what elements it will pick up, in a real page. I find CSS hard to interpret. I can't easily see bugs, and I get burnt when I tweak existing CSS. I've learnt my selectors, and have a fair idea of what formatting one can change, and how it can be done. I even have a CSS mentor to explain the obvious problems, and to help me understand where some thing perhaps can't be helped. Nonetheless, unexpected behaviours (changes to things I thought unrelated to my changes) rock up regularly. Let's parse some CSS: ```css #fromPageContent, .front-page-section p { font-size: 1.25rem; } ``` CSS selectors are those parts that come before the {formatting rules}. The browser will apply the format `font-size: 1.25rem;` to every element which matches the selector `#fromPageContent`. The `#` means this is a class selector, so the browser will format *all* elements which have the class `fromPageContent`. --- ## An Example Here's an example of how I explored to find out. It worked on the front page of this site (at the time of writing): I've got a CSS selector that I want to double check. In the CSS, it reads `.front-page-section p { ...some formatting ... }`. I can read that as *"paragraphs within elements of class `front-page-selection` should look like ...some formatting..."*. To double-check it, I'm going to use my browser's **DevTools**, and I'm going to write a bit of **JavaScript** to tell me what, in the browser, should take ...some formatting... There are plenty of ways to open the DevTools – generally, I right-click and choose "Inspect", which works on Chrome and Safari. Once there, I head to a console tab, and look for the `>` caret. At the `>` caret, I type in the query I'm interested in. Here, it's `document.querySelectorAll('.front-page-section p')`. I hit return. The console hands me back a bunch of stuff. In Safari (but not in Chrome), it's a list of paragraphs. 19 of them – and looking inside, I can see that the list includes things that I don't want. So I know my selector is wrong. I re-interpret the selector query as *"*all* paragraphs within *all* elements of class `front-page-selection` *and their children* should look like ...some formatting..."*. I just want to affect the paras immediately within that selection. I reckon I need to use the `>` selector, which selects children, and not more-distant descendents. I could re-write my selector as `.front-page-section**>**p`. I can try that out with `document.querySelectorAll('.front-page-section>p')`. It picks out what I need. Having experimented, I can change the CSS with a bit of confidence. Which is handy: I'm in a theme and changes need to be recompiled and re-uploaded, which means trial-and-error takes about 3 minutes per try. ## Use as a Tester As a tester, I use selectors to pick out specific elements from web pages (and XML, come to that) to work with in tools. HTML/CSS layout is everywhere, and if you need to deal with a UI, you'll often be pulling out parts with CSS selectors. Trying the blasted things💣 out before I put them into a test helps me see when a query brings back more than one thing, or when it brings back stuff I didn't expect (often along with what I wanted). More generally, trying things out before I make them permanent appeals to me as an experimenter and explorer. I've got a shallow understanding of broad technologies; using the console to give me feedback lets me learn directly, and (by surprising me) exposes my "how hard can it be" assumptions in a way that I can learn from. ## Footnotes (kind-of) 🗣️ I think I first saw this handy trick demonstrated by Alan Richardson and Viv Richards at their [MoT workshop on DevTools in Summer 2019](https://www.ministryoftesting.com/events/london-tester-gathering-workshops-2019). 💡 HTML elements (stuff inside a `stuff`) can have none, one or many classes – and the classes don't have to exist. 💣 CSS selectors don't fit well in my head; I've had to learn them by heart, several times. \*\* You can pick out the same element in different ways. It's good to have several choices, because (for testers who want their code to last) some selectors are brittle. Having brittle elements in your tests means that pretty ordinary changes to the system-under-test can be fatal to the plumbing of your test code. I've been stung by selectors which use IDs, and deep nests of selectors. I've had successes when I combine page structures and classes, especially in situations where the pages are built with HTML that makes semantic sense. ### Encourage Me... URL: https://www.workroom-productions.com/subscribers-get/ Last updated: 2024-09-22T19:31:23.000Z I need encouragement – we all do. When you subscribe, you'll make me happy. When I'm tempted to deliver less, or less-often, I'll feel guilty. Carrot *and* stick. There's something in it for you, too. You'll get to feel part of an ongoing effort to take my corporate work public. Lots of it is already out there, and I know it's useful. The more we share, the better we, as a collective, learn. How altruistic of you to subscribe. There's more: You can use my stuff to explain testing to yourself, and use it to explain to others. I'd be delighted if you did, and I'd love to hear about it. As a subscriber, you'll get to comment on posts. I'll reply, of course. You can opt in or out of the regular newsletter, announcing new content and this week's workshop. I'll probably set up more than one, split by interest, and I'll ask susbscribers for what they're particularly interested in. It's up to you: [free](https://www.workroom-productions.com/free-subscription/) or [paid](https://www.workroom-productions.com/paying-subscribers/) subscription? ### Free Subscribers get... URL: https://www.workroom-productions.com/free-subscription/ Last updated: 2022-01-11T13:19:37.000Z ...more depth. I want everyone and anyone to be able to play with toys and puzzles and exercises, subscribed or not. I'm no great fan of paywalls. However, I want to be able to identify people who use my stuff, and who want more depth. So I'm keeping some in-depth stuff back for subscribers only. I can always release it more publicly if you need me to. Indeed, I'll probably switch some stuff in and out of public access to catch the public's interest. When you subscribe, you can expect: - **Notes** on [exercises](../tag/exercises) - **Ideas and questions** to take you further on [videos](../tag/videos/) - **Teaching notes** on course content - **Hints** on puzzles (if you want them) - More **commentary** on talk [outlines](../tag/outlines/), [articles](../tag/articles/), [events](../tag/events/) and more. - I'll run occasional **online get-togethers** for all subscribers. - You'll find that you can **comment** on posts. - You'll find that you can opt-in / opt-out of a regular **newsletter**. - ...and I'll fold away that "subscribers" section on the home page. --- \`\* free to you. Costs me a small chunk. ### Paying Subscribers get... URL: https://www.workroom-productions.com/paying-subscribers/ Last updated: 2025-01-07T11:48:04.000Z [15 minutes interactive exercises](https://www.workroom-productions.com/workroom-playtime/) a week with me. And also... I'll be surprised and delighted. You'll feel that you're helping me to contribute to our community, and you'll probably have my attention in a different way. ### Limited-time offer… I think that I can promise that early subscribers get to keep the rate they signed up at, forever. So I’m offering the current pricing to the first 25 paying subscribers – it’ll go up after that as I seek a sustainable level which fits with my other offerings. **To lock in my lowest-ever pricing** [***subscribe now***](https://www.workroom-productions.com/#/portal/signup)**.** ### Paid Member page URL: https://www.workroom-productions.com/test-paid-member-thing/ Last updated: 2022-06-06T21:11:37.000Z Testing paid membes _This post is for paying subscribers only._ ### Fuzzy Search URL: https://www.workroom-productions.com/fuzzy-search/ Last updated: 2026-03-31T15:01:36.000Z I turned on search. Which was easier than expected. I dropped an API key into the header... and Roberta's my aunt. Immediately. Everywhere. Fork me. Ghost has [plenty of options](https://ghost.org/docs/search/) for search, but doesn't provide an API for a *back-end* search of my content on its servers. The theme I'm using (a customised [Liebling](https://github.com/eddiesigner/liebling), for now) does a [*front-end* search](https://support.algolia.com/hc/en-us/articles/4406981933457-Searching-from-the-front-end-or-the-back-end-What-do-you-recommend-). What does that mean? It uses the Ghost API, so to pull site content into your browser, slams through it for an index, rubs that lamp with your search term, and returns page titles as you type. It uses [fuse.js](https://fusejs.io/concepts/scoring-theory.html), which is a fuzzy search (implementing a [Bitap algorithm](https://en.wikipedia.org/wiki/Bitap%5Falgorithm), apparently). What might that mean? **Let's experiment**. Subscribers get to see what I found. _This post is for subscribers only._ ### Wicked Problems URL: https://www.workroom-productions.com/wicked-problems/ Last updated: 2022-06-06T21:17:14.000Z Subscribers can explore this topic further with questions and further reading. _This post is for subscribers only._ ### Test Data Generation – building XML with Python URL: https://www.workroom-productions.com/test-data-generation-building-xml-with-python/ Last updated: 2021-11-20T00:13:23.000Z > *Hands-on workshop, where everyone will have the chance to use Python to build meaningful test data in XML.* In this short workshop, we'll use Python to generate XML, looking into namespaces, schema definintions / DTDs, data types and templating. We'll see how simple Python scripts can help you to flexibly and swiftly deliver custom and controlled XML which is valid, well-formed, and useful to testers. This workshop will suit you if you need to generate data for your tests, or of your tests involve XML ingestion and manipulation. The scripts and sample data for this workshop are open-source. You'll get most from this workshop if you bring a working Python environment on a local machine. If you don't have that, you can use one of the workshop's cloud python environments, accessed through a browser. - *Python libraries and methods to build XML* - *Specific considerations when building XML* - *Dealing with test data, and how it relates to test design* Submitted, not yet delivered. ### People need Strategy. Automata need Plans. URL: https://www.workroom-productions.com/people-need-strategy-talk/ Last updated: 2021-11-20T00:08:41.000Z > *How to steer a group towards a goal* *Match your method to the things that will be doing the working.* *Come play with control, constraints, resonance and emergent behaviours. We’ll contrast with harmony, hacking and improvisation, and try to judge some ways we might steer groups towards goals.* - *Strategy helps people work coherently at a distance. A plan gets stuff into the right arrangement in time and space.* - *A strategy sets out what you plan, and how you plan it.* - *Strategy rests on shared values and contextual awareness. Plans rest on organising parts.* - *You use play to change a group into a team.* - *Culture is the collection of assumptions that we choose to keep.* Went to AgileTestingDays 2019, and CultureCon 2019\. Involved group singing. Absolutely bombed at both. I've got videos, and I'll watch them one day. Hey ho. ### Keynote eXtreme URL: https://www.workroom-productions.com/keynote-extreme/ Last updated: 2022-01-11T13:35:31.000Z *Session went out at AgileTestingDays 2021\. That's the title ATD gave it.* Here are the titles we (Bart and I) used. We reckon that these titles could be spoken about, by most testers-who-speak, without preparation for 5-15 minutes. You're welcome to use them yourself. Subscribers (free) get more: the principles that guided us towards these, the background, some of our notes and submission stuff, how we ran it on the night, and what we'd do differently. Sign up or sign in. ## Titles - Taking Responsibility - Trust and the Modern Tester - Things that Scare Me in Testing - Testers, Teams & Me - Tools, Toys & Testers - Growing Testers - Kindness and Quality - Shift Where? Shift What? - Guided by Use - My Testing Story - My Machinery of Testing - Moments that Changed me - Testing the Unicorn - Risk and Me _This post is for subscribers only._ ### Wrangling, Debugging and Testing (conference talk abstract) URL: https://www.workroom-productions.com/wrangling-debugging-and-testing-conference-talk-abstract/ Last updated: 2022-06-18T17:44:51.000Z Superseded by [Wrangling, Debugging and TestingDo we spend months getting our systems to a point where we can test? We do. Here’s why, and what we can do about it.![](https://www.workroom-productions.com/favicon.png)Workroom ProductionsJames Lyndsay![](https://images.unsplash.com/photo-1484729191033-ab703f3eac3a?crop=entropy&cs=tinysrgb&fit=max&fm=jpg&ixid=MnwxMTc3M3wwfDF8c2VhcmNofDEwfHxjb3clMjBidWd8ZW58MHx8fHwxNjU1NTczMDM3&ixlib=rb-1.2.1&q=80&w=2000)](https://www.workroom-productions.com/wrangling-debugging-and-testing/) > Abstract for a 40-min talk. Seems a bit unlilkely, to me, that I can fit this content into that time. So it needs refining. Before a system can be verified with even the simplest automation, it must work, and work right. Getting to that point can involve testers in months of system wrangling and debugging. Yet our industry pays less attention to this necessary task than it does to more controllable or satisfying testing work. This talk seeks to redress the balance with stories from integration projects testing systems critical to large businesses and to national infrastructure. Exploring **wrangling** (getting the system to *work*), we will look at situations where the system under test could not and would not work in any test-accessible environment, and at situations where testers could not gain adequate access to critical elements of the system. We will cover problems such as getting the plumbing right between system components, orchestrating authorisation for tools and testers, working with configuration data and managing interlocking shared environments. We will look into ways that testers can gain technical savvy and leverage political and expert champions to unlock problems. Focussing on **debugging** (getting the system to *work well-enough to start testing*), we will look into situations where subject-matter experts and technical experts rely on trial and error persuade the working system to process data, and where the earliest actions on the system revealed fundamental problems in performance, security, and usability. We will cover designing bulk experiments to take the slog out of trial and error, creating configuration data and consumable data in enough variety (and in enough volume), and will talk about ways to get access to logs and transaction information. We will cover ways to bring organisationally-distant colleagues together to resolve trouble with cross-disciplinary insights. Throughout the session, the speaker will give information about what worked to mitigate some of these issues. We will look briefly at how time spent wrangling and debugging can impact an organisation expected test activities and metrics. We will not cover ways to disguise debugging and wrangling in project plans, and the speaker will resist the urge to offer platitudes and simple metrics. We will not cover how to plan testing to include this often unanticipated work in a way that doesn’t scare the project board. --- ## **Use at work** Some may look at their work of wrangling systems in order to test them, and think "We can do better, and this is how" --- ## Takeaways - Stories from real systems integration projects - Common pitfalls when testing several interlinked systems - Some worthwhile approaches to testing several interlinked systems --- ### Questions for Testers URL: https://www.workroom-productions.com/questions-for-testers/ Last updated: 2026-05-29T14:57:09.000Z card 1 idea from [sussed](https://sussedcardgames.com "Sussed Card Games") by @workroomprds [web](http://workroom-productions.com) [twitter](https://twitter.com/workroomprds "James Lyndsay (@workroomprds) on Twitter") Deck and questions yours on [github](https://github.com/workroomprds/QuestionsForTesters "GitHub - workroomprds/QuestionsForTesters") 1 ## Questions For Testers ### To trigger conversations and build connections. Make a group. Take turns. Each turn, one of you reads a question, and its options, out loud – then chooses one option, privately. Harder decisions are more interesting. Talk. The reader may give points to anyone who predicts (or sways) their choice. Points are pointless. ← swipe → Members get to read a little about how I ported the cards into Ghost. _This post is for subscribers only._ ### Questions, Questions URL: https://www.workroom-productions.com/questions-questions-abstract/ Last updated: 2022-03-21T09:38:16.000Z > this is an abstract for a half-day conference tutorial. If it feels unfinished, that's because it is. I've submitted it to a conference, and I'll polish it here. I'll add it to GitHub, too, so that I (and you) can track changes, This interactive workshop will help you ask the right testing questions, of the right people, at the right time. We will identify questions that have been useful to participants. From those questions, we will recognise types of questions, the contexts in which you might use them, and the people you might ask. We will look more closely at types which catch the attention of the group. Examples could include: - questions which give you context for new work - questions which confirm (or refute) an existing model - questions which reveal underlying values and fears - questions you should should be able to answer, as a tester We will look into ways to adjust questions for more specificity, or to allow greater freedom in the answer, to help the person being questioned, or to act as links in a chain. We'll look at ways we can use questions poorly: loaded and leading questions, questions designed to trick or trip, questions which go over the same ground, and other anti-patterns. Finally, we will spend time on when to stop asking, how to process information, and what and who to ask nex. The presenter will share their lists of questions, built over several years consulting. --- ## Use at work - use the patterns described to bring structure and depth to their own questions - see their typical approaches, and improve - use the techniques to obtain more useful information, more swiftly and more willingly --- ## Takeaways - Ask more purposeful and focussed questions - See the patterns, gaps and opportunities in one's own questions - Know how to stop, and what to do next. > Next steps for James – more examples, more content, less repetition ### Agile Testing Days 2021 URL: https://www.workroom-productions.com/agile-testing-days-2021/ Last updated: 2022-02-10T22:54:22.000Z ## Nov. 15 – 18, 2021 • Potsdam, Germany My first in-person event for ... a while. [ATD](https://agiletestingdays.com) looked after everyone carefully, and it was *still* a blast. Bart Knaack and I ran our *[Speaker Prep Day](https://www.agiletestingdays.com/2021/session/speaker-prep-part1/)* workshop Bart and I also facilitated *[Keynote eXtreme](./keynote-extreme)* *Kevin Harris and I are postponing TestMaster until next year* ### Online Exploratory Testing Workshop URL: https://www.workroom-productions.com/et_workshop_01/ Last updated: 2021-11-12T14:08:19.000Z I've taught exploratory testing since 2002 – I've made a name for teaching by testing custom-built software, real systems, and my clients' products. I started to teach corporate clients online in 2018\. I've never taught a public workshop online. It's time to try that out. I'll use some of my existing coursework and exercises. Each will get a short introductory video, an exercise and follow-up questions. I'll release one a week, to paying subscribers, over eight to ten weeks. We'll will start in late November and continue into the New Year. Each week, we'll have an online real-time conversation to talk about the exercise, and to help us exchange and reinforce learning points. Free subscribers can see the video and the exercise. Paid subscribers get the follow-up questions and access to the conversation as well. I've put a course outline, below – subscribe to see it! ## Tentative outine Week 1 – exploration and purpose Week 2 – emergent test design 1 Week 3 – emergent test design 2 Week 4 – parsing an interface, inferring intent Week 5 – exploration surfaces Week 6 – diagnosis Week 7 – bulk testing Week 8 – exploiting automation ### Debugging: Not Testing but Drowning 3 URL: https://www.workroom-productions.com/not-testing-but-drowning-3/ Last updated: 2022-04-22T21:24:46.000Z *Debugging* is where you have a system which is, at last, responding to you. However, it is responding with unexpected errors, or intermittently failing to respond, or merrily telling you that everything is fine while you can see that, in fact, everything is on fire. When debugging, you may be running plenty of data (if you have it) to see what’s weird; digging into system diagrams to work out what could be to blame; opening filesystems and running diagnostics to see not only what’s going on at a more granular level but also whether those system diagrams and docs are currently accurate; looking in libraries and supporting systems and manual for error messages; working with configurators and coders on fixes, waiting for and wrangling those fixes into place, seeing whether the fix has made the problem go away (?made others turn up?); persuading and advocating, listening and learning to gain leverage on getting fixed; talking to people who might use the system about what your prioblem might mean, and whether it’s a real problem. Where you might turn to a system integrator or a DBA in wrangling, for debugging you’re turning more often to a coder or designer or business user. You’re certainly designing tests and exploring a system when debugging. However, those tests and that exploration are to address a current risk, rather than to add much of long-term use. You’re holding not two but three models; what you imagine the system should be doing / what the builders found it would do / what it’s actually doing (this is the part which is factual, but unknown). ### Wrangling: Not Testing but Drowning 2 URL: https://www.workroom-productions.com/not-testing-but-drowning-2/ Last updated: 2022-04-22T21:31:03.000Z When a team is trying to get a system to work at all, I’d characterise their actions with the word *Wrangling*. Testers wrangle systems in a test environment. Their work is made harder if they don’t have permissions for that environment. They need to give that environment, and their system, the right configuration to gently hum with no input. They need tests to observe that steady state / ready state / several states - typically low-volume, happy-path stuff which often is called ‘smoke testing’. Ideally, but not always, they need their systems to be able to run those tests without restarting - which may need data to be re-generated, or (worse) environments to be purged, databases to be cleared or counters to be reset. They may need information, and in the absence of information, to iterate towards something which may work. Over the years, I’ve worked with several teams which have spent weeks trying to get one transaction through their system. The typical pattern of bottlenecks involves: gaining access to they system – aligning test accounts (may be more than one needed if a key flow ie approvals needs more than one actor); gaining permissions through the layers of system hierarchy to act; gaining permissions through the layers of system and organisational hierarchy to observe; blocking off or stubbing outputs to other systems; mocking round-trips to other systems; using trial and error to discover the correct sequence of actions needed to get they system to do something; using trial and error to find the right information - lining up record validity, internal consistency, business rules and coherency with pre-existing data; working round and through points where the current knowledge is inadequate; working around and through points where the current knowledge is wrong; discovering output points for correct stuff; keeping track of different error messages and debugging towards those; observing different outputs and judging not whether the system is working or not, but whether the test is; writing automation to do what a person can do; much more (inc. limits of system here). Each of these deserves an example - perhaps I’ll add those at a later edit. You don’t especially care about business stakeholders, when you’re wrangling, and they don’t particularly see what you’re doing. You don’t necessarily need a model of what the system *does* in terms of value to the organisation - but you’ll need a model of what the system *is* in terms of how its parts link together. ### Not Testing, but Drowning URL: https://www.workroom-productions.com/not-testing-but-drowning-1/ Last updated: 2021-11-12T22:58:42.000Z Sometimes, I work with a client’s test team. And their work is to get the system to work. Are they testing? They are. They’re paid to test, they identify as testers, they talk about their work as ‘testing’. And they’re not. They’re not demonstrating requirements, they’re not digging into system behaviour, they’re not judging the system. The information they gather is generally swiftly forgotten – acted on if necessary, otherwise sunk. These testers – or at least their actions in this context – are not served by requirements tracing, system exploration, test management and reporting. They’ve been left behind by tools, left behind by standards, left behind by great chunks of our industry. Yet their work is still needed, their hard-won expertise still highly valued within their teams. So much so that the chattering end of our industry is only tangentially relevant to teams whose time is spent in getting systems to work. Working to get something to work has been a part of testing since I first professionally tested a thing which was not my own, in 1986\. It’s been a part of testing at pretty-much every client (at least those clients who had something to test) since then. Perhaps my clients form a diminishing cohort of antiquated organisations. Let’s think, then, about the transferrable skills of *wrangling* and *debugging* which underly the daily work of so many testers. I’m not generally a definition-maker, but to help me (and us) think about this, I'll write next about some of the aims of these activities. *(This is likely to be a four-part post – this one, Wrangling, Debugging, and a last one on what we can do in Testing to help our wrangling, debugging colleagues to move on to work which produces relevant and valuable information)* ### Exploratory Testing Notes URL: https://www.workroom-productions.com/exploratory-testing-notes/ Last updated: 2024-01-30T11:29:11.000Z A suite of short papers, detailing my thoughts on questions that turn up frequently in my Exploratory Testing classes, and when working with ET teams. - ET Notes - [Why Exploration has a place in any Strategy](https://workroom-productions.com/papers/Exploration%20and%20Strategy.pdf) - ET Notes - [Scripting and Exploring](https://workroom-productions.com/papers/ET%20Script%20or%20Explore.pdf) - ET Notes - [The Importance of being Judgemental](https://workroom-productions.com/papers/Judging.pdf) - ET Notes - [What to Record](https://www.workroom-productions.com/what-to-record/) (may change) – [What to Record](https://workroom-productions.com/papers/Record.pdf) (orignal pdf) - ET Notes - [Software testing diagram 1 and variants](https://workroom-productions.com/papers/SWT%20diag%201.pdf) --- ### Video: Wicked Problems URL: https://www.workroom-productions.com/video-wicked-problems/ Last updated: 2021-06-27T21:48:18.000Z Explore this topic further by using one of the tasks / questions below. > ### Go deeper > > Here's [Wikipedia's entry on Wicked Problems](https://en.wikipedia.org/wiki/Wicked%5Fproblem) – which definition (Rittel & Webber / Conklin) resonates most with you? Which quality in particular? > > Do you think that a problem has to meet every qualification in a definition to be *wicked*? > > Are all unsolvable problems wicked? Are all wicked problems unsolvable? > > Is it useful to distinguish between wicked and tame problems? Why? What other distinctions do you use? > ### Application to your work > > What examples of wicked problems are meaningful to you? Can you think of an example from your own context? > > How do you approach problems in your work where you don't know the scope? > > Can you think of a problem in your work where the solution was not obvious – until you found it? > > Can you think of a time that you started work on a problem as if it was tame, and then realised it might be more wild? > > What problems do you have, in your work, that might be combinations of tame and wicked problems? > > What problems do you have, in your work, that seem particularly wicked? Why? > > What problems do you have, in your testing work, that you cannot experiment with? How could you change your context to allow you to experiment? > > As testers, we *find* problems. Do we find *wicked* problems? > ### Further Reading > > Stahl and DeGrace wrote "Wicked Problems, Righteous Solutions" in 1990, about wicked problems in software development. Research it – is it still relevant? > > In this short [Guardian article](https://www.theguardian.com/social-enterprise-network/2012/jun/08/wicked-problems), Robert Ashton ascribes Jeff Conklin's principles to Tim Curtis. Consider the "four barriers" listed in the article – do you see them? Do you see others? > > Have a look at Tom Wujec's [Draw How to Make Toast](https://www.drawtoast.com) exercise and TED talk. ### Exercise: Other People's Code URL: https://www.workroom-productions.com/exercise_other_peoples_code/ Last updated: 2022-04-22T21:48:48.000Z *Consider your own code through the lens of other people's code.* **Purpose:** to reveal how to we might make our code more open to other people. **Method:** Read other people’s code. Identify what might make parts of that code easy or difficult to understand. Decide what you might adjust in your own code. **Logistics:** 40-60 minutes, working group 2-10. ### Sequence 1. Pick your language / tech stack / area of interest 2\. Find some unfamiliar code which someone else has written. You might: - Search in GitHub - Look on Rosetta Code - Pick up some code from elsewhere in your organisation 3\. Spend 10 minutes reading the code and any surrounding text. **Confusion is your marker** – keep track of the points where you’re wondering, whether you resolve it or not. **Try to understand (several or all of):** its structure / its purpose / what it parameterises and returns / what keywords it uses / how it has named its data and functions / how it gets around tech stack weirdnesses. 4\. Come back to the group with one or more of the following - An example of **something that you (initially?) struggled to understand**, and (if you can) the changes / comments / context needed to reduce that struggle - If you understand it all, then an example of **something that you would do differently**, what you would do, and why you prefer your way 5\. The group will pick a few. 6\. Using those examples, consider what you might do to make any code you work on to be more-easily understandable. ### MoT Exploratory Testing Week 2021 URL: https://www.workroom-productions.com/mot-exploratory-testing-week-2021/ Last updated: 2021-12-09T13:17:49.000Z ## Live Testing Session at MoT's Exploratory Testing Week An hour-long live exploratory testing session. I took on the Python interpreter with some Python automation, and failed to use ~~used~~ [Roam](https://roamresearch.com) for notes. [Bart Knaack](https://twitter.com/Btknaack) kept me on track, and [Richard Bradshaw](https://twitter.com/FriendlyTester) lurked productively in the online hinterland. GitHub has the [Code](https://github.com/workroomprds/TDD-ish%5Fexplore%5Fpython3%5Fvs%5Fproperty%5F-identity-%5Ftests). MoT has [video and more details](https://www.ministryoftesting.com/dojo/lessons/experience-report-live-exploratory-testing-a-product-with-james-lyndsay). ### Puzzle 35 URL: https://www.workroom-productions.com/puzzle-035/ Last updated: 2026-06-30T15:57:49.000Z The buttons and lamps obey a simple principle. Supported by these lovely people **[Anne-Marie Charrett](https://twitter.com/charrett)**, **[Julie Gardiner](https://twitter.com/cheekytester?lang=en)**, **[Joep Schuurkes](https://www.twitter.com/j19sch)**, **[Jim Holmes](https://twitter.com/aJimHolmes)**, Ioana Chiorean, Kristine Corbus, Ide Koops, Peter Houghton, Adun Urke, Christine Yen, Huib Schoots. Ide helped me test the beta! Help me make more and I'll put your name on the list: Use [Patreon](https://www.patreon.com/workroomprds) Enjoy this? [Support another!](https://www.patreon.com/workroomprds) Built by James Lyndsay - [@workroomprds](http://twitter.com/workroomprds) © Workroom Productions 2022 ### Live Exploratory Testing – Python URL: https://www.workroom-productions.com/live-exploratory-testing-python/ Last updated: 2021-11-21T00:56:02.000Z In Challenge 3 for 2021's Exploratory Testing week from Ministry of Testing, I tested something live. I chose to test the Python interpreter, using test automation in (naturally) Python. I'm not deeply skilled with Python, and had no idea what I'd find. I found interesting information that came as a surprise, but I'm not *sure* I found a *bug*. --- Here's the **challenge** [Exploratory Testing Week - Challenge 3 - Exploratory Testing a Product@alexm one question we didn’t get to during your experience report, @aldila asks: Is it common for you to create such reports on daily basis(for one feature for example)?![](https://aws1.discourse-cdn.com/business7/uploads/ministryoftesting/optimized/2X/2/211dd653c50d7a102809cba62411a3bbf1b3d1df_2_180x180.png)The Clubmwinteringham![](https://aws1.discourse-cdn.com/business7/uploads/ministryoftesting/original/1X/bfef53af7c8a466fee2ee4f7f718639a65e9a9e3.png)](https://club.ministryoftesting.com/t/exploratory-testing-week-challenge-3-exploratory-testing-a-product/48603/5) Here's the **video** of me doing my thing: [Experience Report Live: Exploratory Testing a Product with James LyndsayWatch as James Lyndsay takes on Challenge 3 of the exploratory week live, the Exploratory Testing a Product challenge. He ops to test the Python Interpreter. ...![](https://www.ministryoftesting.com/assets/favicon-4b6ba0a4118ce928dfcf3d3230cdbb09347a5a9b8d98f244c79d0d56384758bb.ico)MoT![](http://www.ministryoftesting.com/assets/dojo-open-graph-01-16c70360eeba3459c3ffdfbefa6fdce584fec1a642891b65575446b35a8ec22d.png)](https://www.ministryoftesting.com/dojo/series/exploratory-testing-week-2021/lessons/experience-report-live-exploratory-testing-a-product-with-james-lyndsay) Here's the **code** I used [GitHub - workroomprds/TDD-ish\_explore\_python3\_vs\_property\_-identity-\_tests: For MoT’s Exploratory Testing Week April 2021For MoT’s Exploratory Testing Week April 2021\. Contribute to workroomprds/TDD-ish\_explore\_python3\_vs\_property\_-identity-\_tests development by creating an account on GitHub.![](https://github.com/fluidicon.png)GitHubworkroomprds![](https://opengraph.githubassets.com/61f6596ed15dc9215962f99601763d738e407ffabc624a91f6c2b05cc4526d83/workroomprds/TDD-ish_explore_python3_vs_property_-identity-_tests)](https://github.com/workroomprds/TDD-ish%5Fexplore%5Fpython3%5Fvs%5Fproperty%5F-identity-%5Ftests) Here's **more** from the week [Exploratory Testing Week 2021Watch all the action from Exploratory Testing Week 2021![](https://www.ministryoftesting.com/assets/favicon-4b6ba0a4118ce928dfcf3d3230cdbb09347a5a9b8d98f244c79d0d56384758bb.ico)MoT![](https://d2h1nbmw1jjnl.cloudfront.net/series/og_images/000/000/213/original/Exploratory_Testing_Week_Rewatch_2021_OPENGRAPH1.png?1621608257)](https://www.ministryoftesting.com/dojo/series/exploratory-testing-week-2021) ### Ansible and DigitalOcean: setting up URL: https://www.workroom-productions.com/ansible-and-digitalocean-setting-up/ Last updated: 2022-04-22T22:33:06.000Z > Here's my most recent post from [blogspot](http://workroomprds.blogspot.com/), copied here as simply as possible with a simple copy/paste from the webpage. Some of it hasn't worked (in-line monospaced especially - I've fiddled with that and with the list at the end, too) but it's not awful. I need to look at bulleted lists, too. Perhaps I can make a better pipeline with Markdown. Compare with the Original at The playbook in my previous blog gives an idea of what *I* do, but it doesn't get *you* working, and misses out lots of configuration. Let's fill in some gaps by looking at the infrastructure round the playbook. But before that, some rationale so you can see some of the reasons for my decisions. When I need to do precise work over and over again, I look around for a tool to help me do that work faster, more accurately and more repeatably\*. Plenty of tools exist to help in setting up servers; here's Wikipeda's [big list](https://en.wikipedia.org/wiki/Comparison%5Fof%5Fopen-source%5Fconfiguration%5Fmanagement%5Fsoftware). Chef and Puppet are common choices. I've chosen [Ansible](https://www.ansible.com/) over [Chef](https://www.chef.io/) or [Puppet](https://puppet.com/) because it instructs servers over ssh, so doesn't need to install client software before communicating. I'm told that Ansible is easier to learn. I've leant heavily on other guides – chiefly \* DigitalOcean's [How To Use the DigitalOcean API v2 with Ansible 2.0 on Ubuntu 14.04](https://www.digitalocean.com/community/tutorials/how-to-use-the-digitalocean-api-v2-with-ansible-2-0-on-ubuntu-14-04) and [How To Configure Apache Using Ansible on Ubuntu 14.04](https://www.digitalocean.com/community/tutorials/how-to-configure-apache-using-ansible-on-ubuntu-14-04), \* Ansible's [Getting Started — Ansible Documentation](http://docs.ansible.com/ansible/intro%5Fgetting%5Fstarted.html#your-first-commands). Crucial gaps were filled in by \* [Raymond Yhee](https://gist.github.com/rdhyee)'s [Ansible playbook to launch a digitalocean droplet and then configure it to run Minecraft](https://gist.github.com/rdhyee/7047660) and \* [Lorin Hochstein](http://lorinhochstein.org/)'s [Variables and Facts](https://www.safaribooksonline.com/library/view/ansible-up-and/9781491915318/ch04.html) chapter from his book [Ansible: Up and Running](https://www.safaribooksonline.com/library/view/ansible-up-and/9781491915318/). View them as definitive, and my account below as flaky. Most of the following is done in the terminal of my MacBook under OS X.11.5 (El Capitan), with an admin user. I used Atom as my primary editor. First, we'll need Ansible. I want to run it locally. I used homebrew to install Ansible, installing 2.1.0.0brew install ansible Irritatingly, there's a [bug](https://github.com/Wiredcraft/dopy/issues/41) in the tool that enables Ansible's communication with DigitalOcean. While working on the playbooks, you may see an error `NameError: name 'DoError' is not defined\r\n`. To get over the hump, I needed to *downgrade* the version (0.3.7 for me) of `dopy` (DigitalOceanPYthon) that comes with Ansible. `sudo pip install 'dopy>=0.3.5,<=0.3.5'` I set up a base directory for this project, deep in my home folder. I've not set any special permissions on the new directory. When I run an Ansible command, I run it from that directory, and I keep all my Ansible configuration, playbooks and templates there\*\*. In that base directory, I've put following `ansible.cfg` file to instruct Ansible to use nearby files to read inventory and write logs. `[defaults]inventory = ./hostslog_path=./ansible.log` Please note: here, and below, I'm using `./` to indicate explicitly to the system (and to you, reader) that we're starting from whatever directory we're in – which is the base directory, most of the time. The inventory file holds information about the servers you're asking Ansible to manage. I'm using Ansible locally, so my inventory file `./hosts` looks like this: `[local]localhost ansible_connection=local` You might imagine that I'd need to put all my DigitalOcean servers into this inventory file. I can't, because I don't know them. We'll need to use Ansible's "in-memory inventory" instead. So that's Ansible set up. Let's hope. Ansible has a concept of [playbooks](http://docs.ansible.com/ansible/playbooks.html) – short, readable files that basically link a list of hosts and set of instructions about what to do with them. These playbooks are written in YAML – they're readable, but not as easy to build as to read. I've got a short posting coming on .yml files for Ansible. Ansible's [example playbooks](http://docs.ansible.com/ansible/playbooks.html) use [roles](http://docs.ansible.com/ansible/playbooks%5Froles.html) – which is good for reuse, but harder to read. My playbooks share three sets (currently) of configuration information. Each of these are in their own file, so that they can make changes in one place only, and so that I can strip the information out from change control. They're all together in a `./vars` directory. `./vars/sensititve.yml` – the API key (which I don't want to share) `./vars/sshInfo.yml` – the ssh information (which I want to share temporarily) `./vars/droplets.yml` – the list of servers to build (which will include different configuration options). Let's get the bits together to allow my setup to identify itself to DigitalOcean as a valid account owner, and to the servers as a valid controlling account. The DigitalOcean API key will identify my account from any playbook that sets up (or destroys) servers. I got it from [DigitalOcean - API Tokens](https://cloud.digitalocean.com/settings/api/tokens). I don't want to share it with anyone, ever. If I do, I need to revoke it and get a new one. `./vars/sensititve.yml` looks like: `` --- sensitive: do_token: *« 64 characters of hex signal that looks like noise but isn't. »*`... `` My ssh key information is used in any playbook that communicates with servers. If a server has the public key, and I have the private key, Ansible can log in over ssh without a password. `./vars/sshInfo.yml` looks like this: `-- sshInfo: do_ssh_key_name: TestLabEuroSTAR2016 local_private_ssh_key: ~/.ssh/TestLabEuroSTAR2016...` This needs a bit of matching infrastructure, and a rationale. I want to have the option of sharing the public key with TestLab people, so it needs to be separate from my usual keys. Sharing means I don't want to use my default keys, so I need to make and name it for this task, and specify it explicitly. I made a custom-named key like this `:ssh-keygen -t rsa -f TestLabEuroSTAR2016` which builds two files in `~/.ssh/` for my key. The public part is `TestLabEuroSTAR2016.pub` and the private `TestLabEuroSTAR2016` Note: this command will ask for a passphrase. Don't forget the passphrase; OS X (not Ansible) may ask you to enter it if you're re-using the key after a couple of days. I uploaded the public part of the key to DigitalOcean (via: [DigitalOcean - Settings](https://cloud.digitalocean.com/settings/security) ), so that DigitalOcean can put it on any new server. I gave the key a name on upload, and it's that name I'm using in `do_ssh_key_name` above. When Ansible sets up a new server, it will ask DigitalOcean to add this uploaded *public* key to the server as it's being made. When my tool communicates with the new server to load and configure software, it will use the *private* key `~/.ssh/TestLabEuroSTAR2016` to get in via ssh. If I don't specify this in my playbook, Ansible will quietly default to id\_rsa.pub, and that won't get in. While we're considering shared information, here's one more. I want to have a script that destroys the servers I set up (and only those servers), so I'll need to share information about those, too. My `./vars/droplets.yml` file looks like: `--- droplets: - name: TestLab01 - name: TestLab02...` I can add more servers as I need them. I'm currently limited to 50 droplets. I expect that I'll add more details to each server (an indented list for each - name: line) as I differentiate my servers. When I get my servers set up, I want each to be doing something that differentiates it from its neighbours. I'd also prefer not to be faffing around with IP addresses. I've set up a template to build a web page for each server. Each server's page will have its own name at the top of the page, and a list of named links to the others.I've put my template at ./siteStuff/index.html , and it looks like ``` Basic HTML Template

James's TestLab stuff for {{WPL_server_info}}

Ansible set up this index from a template

``` This is a jinja2 template, and will make bare HTML. The set up index task in my playbook will generate an index.html file for each of the hosts we set up in the newServers group in the in-memory inventory. Look back to see that I set up a bunch of host variables for those servers – the template substitutes the stuff in {curly brackets} with those host variables. It uses the server name (which came from the names in the list of desired droplets), then builds a list of links to all the servers in the group. The playbook uploads the built page to the server. When I run the makeDroplets.yml playbook with ansible-playbook `makeDroplets.yml`, I get plenty of information about what's happening. I won't paste it here. Occasionally, one of the post-server-creation steps fails – I may need to add a `wait_for`. However, Ansible is [idempotent](https://en.wikipedia.org/wiki/Idempotence), so if a step fails I can simply run the playbook again, and it should fill in the gaps. Be aware:- Ansible and DigitalOcean take about a minute to set up each server.- Each of these smallest-possible servers costs $0.007 an hour to run. Which is piddling, until you fire up 50 and forget to destroy them. I can check my handiwork by browsing to one of the IP addresses. I hope to see a page with links to all my newly-minted servers – and when I click through, I hope to observe that the server name changes – and so does the IP address. \* And with my soul intact (but my [yaks shaved](https://en.wiktionary.org/wiki/yak%5Fshaving)). \*\* here's an edited listing to give you an idea of the shape ``` ./ansible.cfg ./ansible.log ./hosts ./destroyDroplets.yml ./makeDroplets.yml ./siteStuff/index.html ./vars ./vars/sensititve.yml ./vars/sshInfo.yml ./vars/droplets.yml ``` ### Why Exploration has a Place in any Strategy URL: https://www.workroom-productions.com/why-exploration-has-a-place-in-any-strategy/ Last updated: 2025-10-15T09:59:23.000Z *Written swiftly, in 2006 or earlier, this unpublished paper was [available as a pdf](https://workroom-productions.com/papers/Exploration%20and%20Strategy.pdf) and intended only as a supporting piece to my workshops. Several people like it. I'm re-posting it here to make it simpler to access, more reliably located, and easier for me to refine. Consider this old, mostly finished and polished for its time.* When I test, I try to think of three aspects of my test. Let's call them, for the sake of argument, ***Action***, ***Information***, and ***Observation***. By ***action***, I mean things to do. Things that are done. Things that have been done. Actions: everything is steady in some sense, until something *acts* on something else. Often, I'll think of actions that the tester takes - but I'll also think of actions that the system takes, that a part of the system takes, or that are taken by something outside the system. By ***information***, I mean things that are. Things already in the system, the state that system is in. Not just the stuff I choose to give to the system. I could call it *data* \- and it often is, but the label is over-loaded and limiting. By ***observation***, I mean things that are perceived. Whatever tool captures information after an action - my eyes, a database monitor, the smoke detector connected to the office sprinkler system - that information has to be *perceived* to be part of the test. --- It's easy to script ***actions***. Many written scripts dictate action rather precisely. ***Information*** is harder to write down completely; most scripts leave at least some information to the tester - or to chance. Scripts can suggest what to ***observe***, but there's simply no way to describe all potentially useful observations in advance. A three-ring-binder scripted test leaves important elements *of* the test *to* the tester. Even if the ***actions*** are set in stone, the ***information*** may change - by choice or by circumstance - and the tester can broaden their ***observations*** if they choose. Exploration is an important part of the skilled manual execution of scripted tests: The tester influences the test design during the test by choosing what to observe, how to treat the information, be the actions ever so precisely prescribed. The word *prescribed* is a lovely word for the Latin lover. *Scribere* meaning 'write', *prae* meaning 'before'. A test script is exactly that; a writing-down-before, of the test. The amateur etymologist in all of us will notice also that *precision* shares the prefix. Its root, however, is *caedere* \- to cut short. --- It is possible, and common, to write a test script that is much more precisely prescribed than a manual test script. Automated tests are, by necessity, both precise and prescribed. That's not to say that a sophisticated automated test script has only one path through the system, only one set of data; an individual script can be re-used for a range of similar tests. Its observations, however, are precise. They are cut short. An automated test won't tell you that the system's slow, unless you tell it to look in advance. It won't tell you that the window leaves a persistent shadow, that every other record in the database has been trashed, that even the false are returning true, unless it knows where to look, and what to look for. Sure, you may notice a problem as you dig through the reams of data you've asked it to gather, but then we're back to exploratory techniques again. I have a lot of time for a certain pure ideal in automated test scripts. They should run any time and often, and without hassling the people who run them. This applies to a great regression test pack as much as it applies to a fine suite of test-first unit tests. They're written to run time and time again, at a greater and greater distance from the point where they were designed, and still bring us useful information. That information is information about **value**. It's not just that the build survived the latest changes, it's that the system does everything we've asked it to. It's not that the regression tests have found no new bugs, it's that the users can still rely on the system to support them in their work. We *need* precisely prescribed scripted tests. They're great. They tell us about value - the value that exists, and continues to exist, in the artefact we're testing. If what is valuable stays much the same, then so do the scripts - write once, run forever. --- It is also possible, and common, to test in a way that is not prescribed. An exploratory test needs no script, no chosen set of actions. Choice of actions is up to the tester, at the point of testing. Choice of information, of observation, is limited not by pre-existing design, but by opportunity and resource. Moment to moment, the tester chooses what to do, what to do it with, how to check what's happened. Interesting things will be examined in more detail, weaknesses tried, doorknobs jiggled. The tester chooses to try two things together that between them open the system to a world of pain. The tester chooses to use *this* information, with *that* action, not to just to see what's desirable, but what's *possible*. The exploratory tester focuses on **risk**. We *need* exploratory tests. They're great. They tell us about risk - unexpected, unpredicted, emergent - that goes hand-in-hand with the system that has been delivered. Exploratory tests are immediate, of the moment. The risk is known - and you'll not need to test for it again until you've addressed it, and *written* a test to show you that it's gone. --- Let's summarise, briefly. Some tests are designed to find risks. They're made on-the-fly and run once. Some are designed to tell us about retained value. They're made once, and run forever after. You need *both*: they tell you different things. --- So much for the ideal. Now for a different perspective. We've all seen diagram 1 of software testing: ![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/2022/04/Two-Circles-1.png) Let's look at the left hand circle; our expectations. These things we expect, those we don't. If we're going to write a set of tests, we're going to scan through the finite set of things we expect. Let's look at the right hand circle; the deliverable. These things we've got, those we haven't. If we're going to do a set of tests, we're going to scan through the finite set of what we've got. The diagram splits the world into four regions. First, let's deal with the overlap. This chunk is things we expect, that we know we've got. Both sets of tests will identify that our requirements are met by our deliverable. That's duplicate work, but let's not judge (just yet). There's the region outside both circles. That's all the stuff we didn't want, and haven't got - let's hope not too much work was spent there. It's not exactly a finite set, whichever way you look at it. Let's consider the left-hand arc. We expected to get this. We didn't get it. The deliverable is less valuable than we'd hoped. Now the right-hand arc. "We found the system does *this* and frankly, Bob, that's a bit of a surprise". Is it risky? Yes, in some abstract sense - but now we've found something in that region, we can discover just how risky. ![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/2022/04/Two-Circles---Value---Risk.png) --- *Scanning through the finite set of things we expect* can be done before we ever have something to test. It relies on a good understanding of our expectations. If we're going to get the job off the critical path, we need that understanding to be stable. Then we can write down the steps in testing, before we ever test. When we do the test, we'll know about what's there, and what's not. We'll know about value. This is **scripted** testing. ![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/2022/04/Two-Circles---iterate-over-expectations.png) *Scanning through the finite set of what we've got* requires, as its first step, something to scan. That means we're on the critical path, no way off. We'll need to be fast, and to be fast, we'll have to be prepared, and skilled. We'll find risks, and risks need to be assessed and addressed, so we'll need to be really good, really soon. This is **exploratory** testing. ![](https://storage.ghost.io/c/7e/30/7e30843b-2abb-494a-ab80-0e931d8ae9a9/content/images/2022/04/Two-Circles---iterate-over-behaviours.png) We *need* both, for a good understanding of risk and value. ### What to Record URL: https://www.workroom-productions.com/what-to-record/ Last updated: 2024-01-30T12:20:16.000Z *Written in 2004... still current? This is one of my swiftly knocked-out* [*ET Notes series*](https://www.workroom-productions.com/exploratory-testing-notes/)*.* A particular phrase has rung through my life as an experimenter. I can remember the day I first heard it, and it's followed me round ever since. Let me set the scene. It's Double Physics. I'm sitting in a classroom - not our usual one, with the high lab desks and the ticker-tape timers, but a smaller one. One where a precious video recorder can be connected to a jittery television without interference from crocodile clips and galvanometers. It's small, and we're all a bit perched and huddled. Not something thirteen-year-old boys enjoy, particularly ones with a touch of Physics in their souls. On screen, a succession of experiments. Bits of metallic stuff are being dropped into dishes of water. Before each experiment, we're shown chemical symbols, the periodic table. The voice from the screen says 'Write it down'. We do. We're shown the weight of the stuff, its colour, the ambient temperature, the volume of the dish, the air pressure. Each time, the voice from the screen tells us to 'Write it down'. We do, navy-blue hardback lab notebooks balanced on knees and the tops of seatbacks. We write it all down. The stuff drops into the water, nothing happens. We write it down. The camera zooms in; nothing. We write it down. Another experiment, more stuff. Still nothing happens. Still we write it down. And so we get used to the nothing. When the metal looks odd after a moment or two's submergence, we write it down. When a sheen of tiny bubbles gradually creeps over it, we write that down too. When the first, single bubble escapes its hold and rushes to the surface of the water, we write it down. The next surrounds itself promptly with a silver sheath of gas, and as it bobbles on the bottom of the bowl, we realise we really should have used a stopwatch. The geekier and richer are using their digital watches. Another experiment fizzes like an aspirin, the next positively leaps about. We're shown more metal - it's yellow grey and skinned with oil - the voice tells us it is Cadmium. We write it down. There's barely a moment after it hits the surface of the water, and the bowl explodes. Steam rises, water pours from the shards, soaking the black cloth that covers the studio table. We write it down. Oddly enough, I've no real idea if that last bit of stuff was cadmium. The choice of subject seems odd, now I think of it, for a Physics lesson. Perhaps it was Double Chemistry - but I've always hated Chemistry. I'm twenty-five years older, and writing it down hasn't helped me retain many of those facts at this distance. The lesson itself, however, has followed me round, whispering 'write it down' in labs and libraries, concerts and car journeys, wrapping my fingers around a pen or pushing them over a keyboard even as my eyelids drop and I'm called in to bed. I've used notepads, jotters, exercise books, Moleskines, dictaphones tape and digital, Palm Pilots, laptops. Everywhere I've gone, everything I've done, I've written it down. --- My life, then, is filled with scraps of rubbish paper, tapes, files in obscure formats. It'll come as no surprise that I've been writing notes throughout my time as a tester. For me, recording what I do is fundamental to doing it. I believe that I do a better job, just because I'm making notes. My mind is clearer, my concentration better, my decisions more justified - and sometimes, more surprising. It's a pain to find that I'm halfway through something, and I've lost my notes. It's worse to find I've not been making any - because although common, writing stuff down is hardly my default behaviour. Last year, I made a loaf of sourdough bread every week for six months; not a note to be seen. I started to make notes - had to find paper, made less bread, got wet flour in the laptop - but the bread got better. Way better. Why did I kid myself by not bothering? --- Sometimes, I teach people to test - and sometimes, I teach the systems benath the peculiar magic that is exploratory testing. Indeed, I teach session-based exploratory testing, and recording sessions seems to be a particular problem for my student explorers. I've coached good testers, who show me three lines of post-test scribble to describe ninety minutes of exploration. I've worked with interested and well-informed teams, with only a buglog to show for their efforts. I've had a class full of people look at their pretty session templates, and write not one thing - not a bug, not a plan, not a target for testing, not their name or a date in the labelled boxes at the top of the paper. It strikes me, at these points, that perhaps I'm getting something wrong in my teaching. I guess the most immediate thing I want to explain at these moments is why one might want to keep good notes. Better still, how keeping good notes can help. Lets do that, just to tick them off: - You think more precisely - You're more likely to make decisions based on the information, not on habit or expectation - You don't have to remember everything - just get it down and move on with a clear mind - You'll remember more when you come back - You can show someone else, any time you like - You can show someone that you've actually been working - Notes don't decay over time ...and more. I go through all this, and my students nod and seem to understand. Indeed, they often suggest the reasons themselves. All compelling, all important, none making but a dribble of difference. Perhaps there are a bunch of 'why not's. If you loaded someone with truth serum and asked away (not something I've yet tried - but that serum's tricky stuff to administer in a classroom), what might a noteless tester say? - I don't have time - Note-taking disrupts my testing karma - I don't want to get caught doing a bad job - Doing the work is more interesting than keeping notes - I don't think it helps, so I'm not going to try, even in this classroom, even after you've pleaded with me to give it a go, even after you've told me my every effort is worthless without the backup that notes provide --- Plausible? Perhaps, though disappointing. But I have another idea. I think that some testers might not have had that voice in their head for most of their adult life, telling them to Write It Down. Some testers just haven't been Writing It Down. Indeed, it is possible, I believe rather patronisingly, that some testers really aren't altogether sure What to Write Down. --- Me? I did the Physics until I finished growing up. Then I segued neatly into Testing, which I've done for years, too. I never really knew what to write down, but I wrote it down anyway. A quarter of century of notes. If you're a compulsive writer-down of unconsidered trifles, I suggest you need read no further. On the other hand, if you'd like a shortcut to my personal take on the Secret Stuff that should be Written Down, I crave your attention for a few paragraphs more. --- If you make a **plan**, write it down. If you're just tootling along all planless, you need a **strategy**, an **approach**. A sticky note will do. There are no excuses - accept no substitutes. You'll want to remember the **actions** you take, the **data** you use, your **expectations**, your **observations** \- including **the time**. Don't necessarily limit yourself to exactly what you're testing - you're working in some kind of **context**. You'll get better at this over time; there's an instinct that comes with practice that lets you separate the wheat from the chaff. There's always going to be a bit of chaff. Keep track of **things that repeat**. Even if nothing happens. Dullness is a virtue in most working systems. And without track of dullness, how will you notice . . . **Surprises**. Is that a goat among the sheep? If you didn't expect it, it's worth writing down. If someone else wouldn't expect it, it's a **bug**. Perhaps you've seen an **exploitation**. Have you a **hypothesis**? Are you making a **model**? And when you've supported your hypothesis, found a potential bug, had a surprise, or the dullness is just too much to bear, you need to . . . . Make a **Decision** \- many people get so used to testing by instinct, or by the book, that they don't notice they're making decisions. Worse, they've no idea what the decisions might have been. Scripted testing can be decisionless, but decisions are key to exploration. When you decide to take a different approach, to try different data, or just to consciously do exactly the same thing again, but watching more closely this time, you're taking a decision. Make a quick note. --- A lot to keep track of? Sure - but that's why you write it down. You can't keep track of all this stuff without a bit of paper by your side, Superman. Just as continuous test design has strange and positive effects on the quality of your code, continuous note-taking works wonders on your testing - and on your thinking. A bare minimum? I always have a spare moment for a bare minimum. For me; strategy, data, surprises, decisions. For you, something else. Keep notes, and you'll be there in no time. It's not hard, it's not dull, but it needs a little persistence, a little focus, a little discipline. So; if this article has triggered one new idea, a decision, an observation, a single spark of intent or insight, I urge you to stop reading, right now, take pen and paper and . . . **Write it down.** 💡 ****Notation** **`-`** Item **`*`** A more important item - sometimes used for 'return to this' **`!`** One you'll want to remember at the end of the test. `!!` is typically a bug, sometimes qualified with a spare `?` or `¿`. **`[`** An aside - a thought or observation that needs to go down, but that isn't in the flow`]` **`¿`** Something I'm not sure of - may need more tests **`?`** A question for someone, or something **`Plenty of arrows and circles`** \- not forgetting diagrams, underlining, tables etc. ### STARWest 2001 URL: https://www.workroom-productions.com/starwest-2001/ Last updated: 2023-06-19T10:43:56.000Z *San Jose, Oct 29 - Nov 2, 2001* Not my first [STAR](https://www.techwell.com/software-conferences), I think. But the first I spoke at. My talk was *Better Data, Better Testing.* The organisers renamed it. I met Elisabeth Hendrickson, Steve Splaine, James and Jon Bach, Cem Kaner, Ross Collard. [Software Testing Analysis & Review (STARWEST) 2001 Concurrent SessionsSoftware Quality Engineering (SQE) develops quality testing training for developers, testers, quality assurance, and management to improve software quality![](https://conferences.techwell.com/archives/sw2001/favicon.ico)![](https://conferences.techwell.com/archives/sw2001/images/lg-sw00-top.gif)](https://conferences.techwell.com/archives/sw2001/concurrent-f.html?day=f)