Wednesday, July 20, 2016

'Get in the groove': Something I assume nobody will bother implementing ever

You know how when you're driving a long the highway, and you get to a nice stretch where for minutes at a time there are no repairs or the slightest pot holes and you just drive along listening to the beat of the seams of the concrete slabs, and the grooves embedded in the concrete give a certain pleasant drone to the drive? And then you get off the highway and the grooves stop and life is somehow just that little bit more boring?

What makes that droning noise is the tires spinning along the pavement and the tone or pitch of the drone comes from the width of the grooves.

And that is the start of the idea.

The grooves are created at the time the concrete is laid  (this works for a concrete highly which will 'hold' the grooves, on a local street the asphalt is too soft to maintain the grooves). A different groove separation width (not the grooves themselves, but the distance between the grooves) would produce a different tone with the speeding tires.

So the suggestion is to modify the groove tracer mechanism, which scores the soft concrete being laid down linearly by the machine that lays the concrete, modify it so that the distance between tines is not fixed but is instead modifiable. Some sort of caliper action for the whole width of tines. And have this calibrated so that particular notes could be created in succession, allowing melodies to be embedded into the highway. Hey, go wild, you could have two part harmony if you groove differently the left half vs the right half of a lane.

So you could have say the highway leading into the airport playing the melody for 'Stairway to Heaven' or route 95 in northern NewJersey with the view of the Manhattan skyline playing 'New York, NewYork'.

Of course there are issues. Musically, the pitch is determined by groove width and to some extent by tire tread markings. The speed of the car doesn't alter the pitch considerably because the groove width is parallel to the tires' direction. The tempo would certainly change. There's the safety issue of the distraction of hearing this subtly melodic droning as you drive, but it's not as distracting as billboards. There should definitely be some signage alerting drivers to the sound so they don't think they're going crazy, can't get that tune out of my head.

I think this would only be realistically feasible for long stretches of road like a highway. If local streets were paved in such a manner, intersections and other slower cars would impede satisfactory 'playing' of the melody.

So now how to create a monetization model out of this. Commercial jingles, I-57 sponsored by Anhaeuser-Busch with the Budweiser theme? Christos superscale artwork?

Thursday, July 14, 2016

Science is all mental

Science is all about figuring things out. How do I move this big rock? Which way do a turn the wheel when parked on a hill? Why is that guy such a jerk?

That's all dealing with real world things but we're doing it with our thoughts. The non-obviously obvious ways of doing it are:

Looking or using memory - You can't just make things up (which memory often does sometimes), so you have to look to to make sure you're not remembering wrong. Your mind may make ideas, but you should check them against reality to make sure you're not wrong/crazy.

Naming - we use language to communicate what we have ideas about with others. But frankly, just for ourselves, giving a name to something, using that name with something the same, giving another name to something that is different, those are all mental tools even for yourself.

Guessing well - names don't always fit perfectly or are vague, but start with one word and if that doesn't fit, then use another or create a new one.


Saturday, June 25, 2016

Quantity vs quality - big litter vs few children - product strategy

There's a set of biological characteristics around number of children that seem correlated.

(picture source)
(picture source)

Some organisms, like mice, salmon, and dandelions, have lots of children in a litter. This is correlated with those children (or eggs or seeds) having little parental maintenance, small body size, short-lived as adults, not particularly robust, having a low probability of any individual one reaching maturity. Others have small litters or even just a single child, tend to large body size, long life expectancy, and lots of time and energy is spent in making it a high probability that the offspring lives to child bearing age.

The theory behind it is called r/K selection theory (I only described a vague phenomenon, not any reason or mechanism behind it). The r and K come form a simple equational relation those as probabilities to the number of offspring.

The theory says that the evolutionary pressure that's behind all these correlations is stability of the environment. 'r-species', those with quantity offspring like mice, tend to exist in unstable, unpredictable environments with a lot of room in their niche (high 'carrying capacity'). 'K-species', quality offspring like elephants, usually occur in stable (low population change) environments or close to filling the niche.

That said, there is a loose, not perfect, analogy with product design. Consider a widget making company. The company can spend all its energy creating lots of different widgets of acceptable quality. Some may work out some may not, but having a lot of then increases the chances that at least one will work out in the end.

Another company with the same resources may spend them all on making very few widgets but very refined specially made ones. Each of these few widgets are very likely to become successful because of the resources put into them

So, just like with animal reproduction frequency, it's a tradeoff. For constant resources, low probability survival of individuals can be offset by many individuals, and high probability of survival requiring lots of resources offset by having very few individuals.

For business, the corresponding explanation would be that in an unstable unknown business environment, having lots of products with slightly different features, can increase the chance of viability (eg web game companies have many games on offer but only one or two become famous). In a very stale business environment, with few companies selling few alternatives, a lot of company resources should be spent on just a few very large well-known products.

So this is just theory, and an analogy of a theory, and one can imagine many examples for which the explanation works, for which it practically works, and for what it just makes no sense.

'Agile' and 'lean' and 'fail quickly' are recent manufacturing trends. They are usually in contrast to the software development 'waterfall' planning method. Agile and waterfall don't necessarily apply well to the same kinds of software. Waterfall usually works best for a large scale well-understood design situation, and agile for a small project where all issues are not well-understood before hand and quick changes need to be made based on changing circumstances.

The analogy from biology to business to software development is not perfect but there are enough similarities to be recognizable.

One more analogy: scientific publishing. Historically (Europe), publishing of scientific knowledge started off as monographs (eg Aristotle, or at least that's all that has survived) and during the renaissance included personal correspondence which morphed into privately bound collections of papers and now (early 21st c) the library-industrial complex of journals. Textbooks and monographs are still considered the pinnacle, but are infinitesimally small by weight of paper, or more modernly, number of bytes.

Also, research has turned, during and after WWII, a government supported endeavor (where before only rich people had the means to do it).

There is so much research going on and papers being produced and science being fractured into narrower and narrower domains, that many papers are published (after being peer reviewed) and never read by anyone else. I'm not saying this is a bad thing. It sounds bad because that effort seems wasted. I'm just pointing out the analogy with quantity vs quality; the recent trend is towards quantity. Actually I don't think quality has necessarily bee traded off, just that the resources has exploded, allowing lots of failures to exist alongside really great advances that would never have come without the resources.

Tuesday, June 21, 2016

Comments on "10 things that sound entirely true but are false"

A recent lecture by Neil deGrasse Tyson is entitled

"10 Things we have heard and Re-told but are completely False"

Great lecture on mistakes in elementary astronomy  (only 2 slides for 10 minutes breaks the many slides rule in just the right way). You only need to know elementary physics and astronomy that you pick up in elementary school and life to understand how these are both right and wrong. Here is the main slide:



I want to comment on the flip between 'obviously right -> explanation -> so terribly wrong' presentation for each one, what the nature of the rightness and wrongness is. The generalities are that almost all of them are conceptual mistakes, but some are context mistakes and others naming mistakes. That is, some are actual errors in understanding, some are errors in misplacing the context intended, and others are mistakes in using the names themselves wrongly. No actual trick questions.
  • What goes up must come down - this is a context dependent. In the context of a person throwing a ball in the air, yes, it must come down. In our limited experience that is entirely the case. But rockets are basically high powered throwing. And (most) satellites never come down
  • The Sun is yellow - context. The little time we can actually look at the sun it is near the horizon when yellow is the predominant color. But most of the day it is up high and is blindingly white (it -could- be yellow but happens not to be).
  • Weightless astronauts left Earth's gravity - with respect to say the ISS they loo weightless. But no they are falling around the Earth because of gravity along with the ISS. Also as far as words go, gravity goes forever, some satellites have 'left' = won't be pulled back in/not elliptical path.
  • The North Star is the brightest - this is just a factual error, or rather an error of just assuming that an important thing excels in all aspects. Polaris is actually pretty faint in comparison to other main stars.
  • On a dark night you can see millions of stars - an error of words (and counting?). there are a lot of stars you can see with the naked eye away from the city on a moonless night. But millions? If you count by the area of the sky, given average human eye acuity, there are 1000's. But millions is over reaching, just using the word 'millions' to mean 'a lot'. The Milky Way galaxy you say? Sure you can see the galaxy, but you're not actually able to pick out individual stars making up the milky band.
  • Total solar eclipses are rare - This is a context error. Sure they're rare for your location, but not for the Earth. The band of darkness happens every couple years. Also, rare is relative.
  • Days get longer in summer, shorter in winter - This is a naming error (also a little conceptual). Summer starts (by official name) on June 21st the solstice which is when the days start getting shorter. So if you include June in your meteorological summer, then yeah for part of a month the days do get a little longer before it goes backwards.
  • At noon the Sun is directly overhead - conceptual: sure in them olden days, that was the definition of noon, when the Sun was at its highest (I'm not going to get into the complexities of directly overhead). Because of time zones and daylight savings, the clock time of 12pm has been set so that the Sun is mostly near the highest point, but further off depending how close you are to the border of a timezone. It's hard for the eye to tell how far off things are in the sky. 
  • The Sun rises in East, sets in West - a naming problem, somewhat pedantic. Directly due East is a point on the horizon, and the sun only rises there on the equinoxes. Dates further away it is further away on the horizon (furthest at solstices). All it takes is looking at where the sun sets on June 21st. People just tend not to do that.
  • The Moon only comes out at night - I think this fits all three. 'Comes out'? You only notice it at night. The moon is always there somewhere in the sky.  We've all seen a half moon during the day.

Monday, June 20, 2016

What are really the problems with EHRs

There are a lot of complaints about EHRs (2016). Too much useless typing, too many clicks to get what you want, records are not really available. scanned documents are a pain, release forms take forever.

The intended benefits of an EHR are obvious. Data gathered about a patient should be available to everybody who needs to see it, quickly and seamlessly, just like all the other rocket science apps that track our dating.

I see two major problems: data sharing, and user experience.


  • Data sharing - electronic health records was never a community service. When a doc or medical situation of a small team needed an IT solution, it was solved only for that particular team or doc. Nothing was intended to be shared. This creates the data silos. It would be a perfect metaphor except real silos can exchange grain so easily just by trucking it over. There is also the other turn of phrase, standards, of which there are, comically, many. There's no universal heath record, or even univesal health patient identifier.
  • User Experience, both data entry and retrieval. The pencil used to be the universal recording medium. It was infinitely creative, hobbled a bit by legibility. Typing is so ... easy... that you're expected to do it constantly, but you can't draw. For retrieval, the current EHRs have at best the most rudimentary search. The EHR for a single patient reads like an electronic phone book: if you know what you're looking for you can find it, but it doesn't tell you what the town is like. Everyone complains that you get a lot of data but you just don't get what is happening to the patient.
It's annoying to hear complaints with out solutions. For once I feel I have some.
  • Data sharing - the world is going to have to spend some time and money making a universal health record, just like a utility. It's not difficult to do, it just takes some desire and money
  • UX - there's a lot of deeply -thought out UX design that could happen. But really just a quick modification to the 'facesheet': add a 3 line text box for a few notes about current status. A lot more could be done but that would change things radically.
These problems are not rocket science. Technologically they are simple. Maybe labor is involved but not much thought.

Going away speech

People often ask you at the end (well, really a change) of a career if you have any regrets. Is there something you would do differently, now that you know the consequences, your deathbed confessions of life changing decisions, minor twists that inordinately changed the direction of later events, what you wish you had done or or what you wish you had not done or stopped doing over and over and over.

You will regret many things. But of all those individual things, you will regret the many times you've had to regret things and you will regret that you failed to regret the things you've done and haven't done. You will probably regret hearing this too.

But now that you are a half step out of our daily lives, I'll tell you this: at your new place, things will look up, things will look down and you may miss all those good ol' days and familiar likeable and happy faces, just feel comfortable in that, once the door has closed behind you, whatever happens back there, those that are left behind will end up blaming you for it.


How to sound native in a foreign language

How do you sound native in a foreign language. First, there's obviously the accent, how you get those weird vowels and inflection just right. There's just plain grammar, which words are masculine or feminine, is the the past perfect continuous participle or the simple hortative passive? Then there's just plain word choice: a sentence could be translated one-for-one but you just don't say it that way in the other language.

Most grammars/instruction of languages will give you this. They give you all the conjugations of all the regular and irregular verbs. They teach you the right preposition. And maybe you can give a lecture in high-speed particle physics or order cake at a coffee shop or even discuss basic politics with a taxi driver. And all this fluently.

But they don't tell you how to be influent like a native. How to make the mistakes a native speaker would make. Hemming and hawing and slurring and skipping unnecessary words and adding the slightest of hints at words that change the entire meaning of a sentence, all like a native.

So here's a list of things, some barely linguistic grunts, in English that you should learn in the other language of your choice to add that bit of informal fluency.
  • Uh, um - /u/, /um/ just filler until you can think of the next word
  • Hm - /h/ I'm thinking
  • Hunh? (recently considered to be a language universal, the similar phonology) 
  • What? (I didn't hear your or I didn't understand you)
  • Uh hunh (yes) u hÅ«
  • Unh unh (no) Å« ?Å«
  • Right? (wasn't I correct) also No?
  • Hey ('watch out!' or 'look at me!')
  • Pfft /f/ - expression of disdain
  • Ha 
  • just plain laughing - different in every language
  • ow! (that hurts!)
Of course, these may not translate well, or translate at all. German has 'je', 'doch' etc, Chinese the sentence endings 'a', 'ba', 'ne'. What are the universals?

---

It goes without saying (but I'm actually saying it and that's contradictory. Does that go without saying?) that profanity and other taboos are a large part of sounding native. I hesitate to add that as another category of things to learn. First, because they are almost by definition part of an informal language that is not very public and so not that necessary for communication. And second, it's bad enough when a native uses it, but it's extra awful when someone with even the slightest hint of an accent mouths off; it's rude _and_ they didn't do it right!. That said, here goes:


  • ow! (I hurt myself)
  • dammit! (I made a mistake)
  • Damn you! (you made a mistake)
  • You bad person! (insults)
  • Leave me alone
  • You suboptimal person! ()
  • Taboo body parts and functions (sex, death, family, religion, and excrement), stand alone or in combination.

I've gone almost beyond mincing, but I think these are universal situations for which you can give canonical examples and extrapolate from there. Some languages have their own idiomatic domains (e.g. Quebecois seems to only use taboo terms that are also perfectly fine vocabulary of the Catholic Church, Arabic seems to favor comparisons to animals). There is surely a lot of overlap in the list, and many possibilities for each one. Expanding to full phrases may involve all of them.


Tuesday, May 31, 2016

Deep Learning: Not as good, not as bad as you think.

Deep Learning is a new (let's say 1990, but common only since 2005) ML method for identification (categorization, function creation) used mostly in vision and NLP.

Deep Learning is a label given to traditional neural nets that have many more internal nodes than ever before, usually designed in layers to feed one set of learned 'features' into the next.

There's a lot of hype:

Deep Learning is a great new method that is very successful.

but

Deep Learning has been overhyped.

and even worse:

Deep Learning has Deep Flaws

but

(Deep Learning's deep flaws)'s deep flaws

Let's look at details.

Here's the topology of a vision deep learning net:

(from Eindhoven)

Yann LeCun
What's missing from deep learning?
1. Theory 2. Reasoning, structured prediction 3. Memory, short-term/working/episodic memory 4. Unsupervised learning that actually works

From all that, what is it? Is DL a unicorn that will solve all our ML needs? Or s DL an overhyped fraud?

With all such questions, the truth is somewhere between the two extremes, we just have to figure out which way it leans.

Yes, there is a lot of hype. It feels like whatever real world problem there is, world hunger, global warming, DL will solve it. That's just not the case. DL's are a predictive model machine, very good at learning a function (with lots of training data). The function may be yes or no, or even a continuous function, but still it's take an input and give an output that's likely to be right or close to right. Not all real world problems fit that (parts of them surely do, but that's not 'solving' the real world problem.

Also, DL's take a lot of tweaking and babysitting. There are lots of parameters (number of nodes, topology of layers, learning methods, special gimmicks like autoencoding, convolution, LSTM, etc etc with lots of their own params). And there are lots of engineering methods that have made DLs successful, but these methods aren't specific to DL. Lots of better data, better software environments, super fast computing environments, etc etc.

However, there are few methods nowadays that are as successful across broad applications as DL. They really are very successful at what they do and I expect lots of applications to be improved considerably with a DL.

Also, for all the tweaking and engineering that needs to be done (as oppose to the comparatively out of the box implementations of regression, SVMs and random trees), there are all sorts of tools publicly available to make that tweaking much easier: Caffe, Theano libraries like Keras or LasagneTorch,  Nervana’s Neon, CGT, or Mocha in Julia.


So there are lots of problems with DLs. But they're the best we have right now and do stunningly well.

Kinds of Data: there are more than just the basic four

The science of statistics specifies that data points come from four basic types:

  • nominal - these are incomparable labels like truth (yes, no), color (red, blue green), country (UK, France Germany, Italy). There is no relation among these elements there other than that they are in the same set. All you know about them is their names and that a name is different or the same as another.
  • ordinal - only a rank is known (1st, 2nd, 3rd...) and nothing else (we don't know how far ahead 1st is from 2nd), just the order, like finishing order in a race.
  • interval - we know the distance between any two elements A - B  like the height.
  • ratio - we also know the ratio of two numbers where for example A can be twice B, like half-life of an element.
Notice how I describe these both mathematically and conceptually, because often a set selected from a mathematical domain, like the reals, can be interpreted in any one of these. For example, from the reals, they obviously have a ratio by division, and a distance by difference, can be ordered by 'less than', and can be categorical by using cutoffs, say >= 0 for yes and < 0 for no.

Of course, as with most systematizations, this list came after years of using methods that were created to work with whatever data was at hand, and then when the data just didn't work with those methods, new analogous method were created, or entirely new methods created for quite different purposes.

Statistical procedures seem geared to work with one of these types. Chi-squared on contingency tables are good for categorical data. Wilcoxon signed ranks for ordinals, t-tests for integral data, Poisson for count data. But mostly there are just two kinds discrete and continuous which fall to nominal/categorical statistics and pretty much all the rest of statistics respectively.

Existing science isn't as deliberate as a current systematization, as monday-morning quarterbacking/textbook-writing may make it seem. It's more incremental, and filling in gaps as needed rather than laying out the system ahead of time. You have a problem and you use a tool that works good enough right now, you develop that tool incrementally until it metastasizes well beyond it's initial conception. Contingency tables are great ways of summarizing tabular data, but you may want to do a significance test like all the t-test guys. 

---

Any kind of systematization is an oversimplification, forgetting possibly irrelevant details to make different things look alike, and placing a particular item into that systematization is also forgetting possibly irrelevant details to make it look like one of a few categories. But sometimes those details are not so irrelevant.

Binary data is a subset of nominal data, with just two categories. Two by two contingency tables and logistic regression are especially designed to deal with them.  Some multinomial categories will have some minimal relationship, say geographic location with countries, or wavelength for colors (colors are very complex because the brain processes them by multiple systems involving the wavelength, opponent process pairs, or beyond. Rank data is ordinal by definition, but when encoded as numbers, can be processed as interval or even ratio data (depending on the interpretation desired.

These four data types work very well for statistics. But it seems underspecified. We're used to measuring quantities or counting objects so all those categorical and interval methods apply so well. But there's so much more structure to the way things can be measured. Not humanities-style vague, wordy, qualitative description. Perfectly exact, just not necessarily a number.

There is an existing method for description of data. A very rich description method. It's mathematical notation. If data should be treated continuously, use R. If a vector over integers, Z^n. If an ordinal set, then that's a total order. If categorical, then you have a simple set. If the elements are related to each other one on one but in a complex restricted manner, then maybe a graph is the way to notate things. if the elements allow certain operations but not others, then maybe it's from a particular algebra, a Hilbert algebra, or instead a Banach algebra. 

Measurement is not always in the elementary numbers we count or measure or weigh with. There can be quite a bit more structure in the measurements than just a number.

There are no exact synonyms

There are no exact synonyms.

That may sound a little extreme, especially given that thesauruses exist.

There is no pair of words where one can replace the other in all circumstances.

'Bail out the canoe with a bucket'

Can you replace 'bucket' with 'pail'. Of course. But can you say 'kick the pail'? No, of course not, that would be wrong. You can't always replace a word with its purported synonym.

Well, OK, there are some circumstances where there are exact synonyms. In technical circles, especially the sciences and math, there is a special way of attaching a word to a definition. In technical areas one 'stipulates' a definition of a word. That is, you give a word a definition that is simply a shorthand for replacing a word with its definition. You're stating authoritatively that for a word, you must treat it like an exact replacement. Often these technical terms are supposed to be evocative or metaphorical, supposed to give you a good idea of the intended meaning. But you can have whatever mental connotations that help you remember the true meaning but the true meaning is what has been stipulated, it doesn't matter, A=B and that's all there is no more no less.

But with non-technical words, there is no stipulation. A word is just a trigger for some associations. And if it sounds different, then there is no way it can be identical in all situations. Different stimuli can give different responses.

I will even go so far as to say that even a given word is often not its own synonym because all words have multiple meanings.  I'm not even talking about homophones (words that are spelled differently, but sound the same, like 'horse' and 'hoarse') or the other side homographs (words that are spelled the same but can have different pronunciations and meanings, like 'bow' a knot in a ribbon or tie, and 'bow' the front of a ship). I mean a word that is spelled and pronounced the same but has a different but related meaning. For example, 'run' is a verb to move fast by your legs, but is also a noun for a  long rip in a stocking or a small stream.

The point is that if you desire a synonym, you can get that from a thesaurus, but it may not slot in perfectly as a replacement. And even a single word may have many associations and alternate meanings that it is not good in that slot itself.

Thursday, May 5, 2016

Free will?: Science assumes determinism

Whatever the philosophical decisions made about free will versus determinism, science (or factual knowledge) attempts to discover everything that is deterministic, and to that end almost assumes determinism.

The only thing counter to this presumed determinism is its literal negation, non-determinism, which is modeled using probability. And probability is just shorthand for what we don't know or can control yet. This applies all the way from physics to sociology.

Wednesday, May 4, 2016

What's the point in these theorems?

Sometimes math is weird. Often you know exactly why a particular math thing is interesting. Like it's so obvious that algebraic geometry is there to help figure out where really weird multinomials intersect. But other times, even for simple things for which there's lots of research and historical precedence for concern, I just don't get it. Here's a list of things I just don't get. I don't understand the point of pursuing them. I understand the mathematical process, I just don't get the point:

  • Craig's Interpolation Theorem in logic, if a implies c, then there exists b such that a implies b and b implies c and b only involves the intersection of vars from a and c
  • Curry's paradox and Löb's theorem - I have trouble following the elementary proofs of these. They seem to say you can prove anything "'if X is the case then Santa Claus exists' proves Santa Claus exists' or something
  • Herbrand's theorem - proves universals using examples?
  • the Deduction theorem - it just seems so obvious. It's just Modus Ponens, right?
  • quadratic reciprocity - allows computation of square roots in modulo arithmetic. Why you would want to do that, I don't know


I want to understand these things, and I can (usually) follow step by step manipulations, but I just don't get what they are for and what the point is.


Two-pass forward-backward approximation systems

A number of approximation methods work in a two stage cycle, a forward pass to compute the test on the system, and then a backward pass to update the system with the error of the test with respect to the supervised true answer.

In the backward pass, un update function moves in the opposite direction of the edges, update weights as it goes a long.

Neural NetworksDeep Learning - a directed (usually acyclic) graph with weighted edges used to compute a function. In the forward pass, the values at any node are computed as the dot product of the value at the source nodes and the edge weights, then a simple threshold function (in topological sort order). starting from the input nodes (no in-edges) This computes the values at the output nodes (no out-edges). Traditional neural nets have a layer of input nodes, a single layer of hidden (interior) nodes, and a layer of output nodes with no edges directly from input to output. Deep neural nets have more layers. Arbitrary undesigned graphs are not usually not very successful.

Expectation-Maximization approximation of parameters of a statistical model. First the expected value of the likelihood function is calculated, then the model parameters are calculated to maximize that function

Primal Dual linear programming for maximizing a target function restricted by a set of linear constraints. The difficulty is dealing with the possibly large set of constraints and large set of dimensions.

Kalman filter successive measurement refinement. This is usually applied to position measurement with slightly fallible sensors. From an initial (fuzzy) position and direction of an object at time t, the position/direction at time t+1 is predicted y combining the prediction of movement from time t plus a fuzzy sensing at time t+1. Combined, the variance is lessened.

Are any of these even more alike than just the forward-backward pattern? Are there any other algorithms that are superficially similar, have a two-step iterative process?

Update (5/24/2016): this blog post seems to give pointers to how backprop, primal-dual LP, and Kalman filters are interderivable

Sunday, May 1, 2016

My pet language peeves/non-peeves


It really bugs me when other people say (my inner prescriptivist):
  • for 'often, pronounced 'off ten' instead of the correct 'off en'
  • 'Between you and I' instead of the correct 'between you and me'
  • pronouncing 'forward' as 'foh ward' (no first 'r')
  • 'comparable' pronounced 'com `pair able' not ' `com pruh ble'
  • Dwarfs roofs baθs instead of dwarves, rooves, baðz
  • pronouncing 'processes' as 'prah cess eez' instead of the correct 'prah cess ehz'
  • 'my bad' and 'back in the day'
  • whilst/amongst
  • Nutella as New-tella. I prefer  Nuh-tella

Conflicted
  • 'irregardless' - obviously a sign of not caring about words, but it still takes me a second to register it as 'wrong'

It really doesn't bother me at all to say:
  • 'Hopefully' to modify a sentence
  • Feb you ary

I only recently learned how t pronounce correctly:

  • awry. I used to say 'aw ree instead of uh 'wry


Wow, is that it? I was sure these lists would be a lot longer.

Friday, April 22, 2016

How hard is it to learn a foreign language?

How hard is it to learn a foreign language?

If you're like me, you grew up speaking a language. Since you're reading this in English, that was either your native tongue, or you've learned it at school. (If you don't understand this then thang nabbit presumably on the slindy intercongruenda, but that's neither here nor there)

But in the current cultural situation, native speakers of English (NSE) are at an intellectual advantage over speakers because it is the preferred language of diplomacy, science, popular culture. I could make a case that this is a disadvantage for NSEs but for the moment let's just assume that an English speaker is about to learn another language.

Learning a language involves a number of separate domains which are somewhat independent. There's:
  • grammar (syntax), how words are ordered and modified to tell you their function with relation to the other words in a sentence
  • vocabulary and sayings, the dictionary of words, what you call that thing or action
  • pronunciation, the accent, those weird sounds that make up the words. 
  • writing and spelling, how the language goes on paper or screen. Though writing is not really language itself (language was spoken first and then individuals created writing after the fact), it is often the easiest entry into language learning and offers higher volume of consumption per unit time. Non-roman alphabets or complicated spelling rules make it difficult to learn how to pronounce things from writing (certainly for beginners).
  • culture, the supporting tools in your environment that ease language learning, books, media, acceptance. Not traditionally thought of as a part of learning a language, but makes a huge difference in learnability. The internet is making a lot of things just universally available: news and movies, ordering books.
Here is my assessment of how difficult it is for NSEs to learn very specific foreign languages, (in misleadingly quantifiable but semi-educated subjective terms: 1 very easy (little to no effort) , 2 easy, 3 medium, 4 hard, 5 hardest) As usual it is easy to disagree after the fact (after you've spent time learning), but these are what I think the effort will be for those who don't know yet, within 1. Also, 'easy' means easy with respect to many languages. You'll probably still struggle with even the cognate words at first. But some languages don't have even that to fall back on. Also, there are two kinds of difficulty. One is that there is a new distinction, a totally new sound that is hard to do or remember, the other is the scale, usually just a number of things to remember. So there might be many new one-off distinctions (a new 'r'), or there might be a single distinction with lots of instances (like gender on nouns), or both (Russian verbs).

So here is a structured list of those issues for each language/language group.



French - pretty easy overall. French (and Latin cognates) form most of the educated language.
  • grammar - 3 conjugation of verbs in a few tenses but very regular, gender of nouns to remember. Word order mostly like English
  • vocabulary - 2 once past the basic words, educated vocabulary is almost identical. Very easy to guess meaning.
  • pronunciation - 2 mostly the same as English, no distinctions but a few strange sounds to get the accent right (u, nasals)
  • writing/spelling - 2 Latin alphabet, spelling very regular, a few accent marks to be aware of, a few unspoken letters at the ends of words. 
  • cultural - 1 taught by default in the US/England. In the US, products (few of them) that have bilingual packaging are in French for Fr Canada). Easy to get French things to read for learning.


Spanish/Romance - pretty much the same as French except slightly easier, because of close cultural connections (in US) and simple grammar
  • grammar - 2.5 conjugation of verbs but very very regular, gender of nouns to remember, verb conjugations but that's it. ser/estar, por/para only difficulties
  • vocabulary - 2 once past the elementary words, educated vocabulary is almost identical, part of the European group, easy to guess meaning.
  • pronunciation - 2 mostly the same as English, no foreign distinctions but a few strange sounds to get the accent right (g, j, h, r)
  • writing/spelling - 2 Latin alphabet, spelling very regular, a few very regular differences. Very wysiwyg
  • cultural - 1 large subpopulation of Spanish speakers. Currently the default foreign lang taught in  in the US. Lots of media that is Sp available.



Latin - The grammar is a lot: conjugations, declensions, gender, agreement. But the vocabulary is very recognizable. Or if it is not recognizable at first, then there will probably be some mdeical, legal, erudite term in English that corresponds.
  • grammar - 5 conjugation of verbs but very very regular, gender of nouns to remember
  • vocabulary - 2 once past the elementary words, educated vocabulary is mostly English looking. Half of English tech vocabulary is Latin or Latin-derived French.
  • pronunciation - 2 since it's mostly used literarily, pronunciation doesn't seem to matter much. But if you have to, it's very organized. When in doubt sort of like modern Italian. 
  • writing/spelling - 1 Roman alphabet (the original), spelling very regular, a few very regular differences.
  • cultural - 2 The source of European culture. English law and medical vocabulary. Lots of erudite literature. Very few children's stories ('Winnie Ille Pu' is about it) and no news. Even though it is not used in the real world as such, makes law, medicine, and all romance languages very easy to learn.



German - you think it is going to be easy because English and German have close historical roots, but it's not that easy.
  • grammar - 3.5 conjugation of verbs sounds like English, gender and 4 cases of nouns to remember, word order/prepositions just slightly different enough to English to be annoying. 
  • vocabulary -  3 a lot of similarity with elementary words, educated vocabulary is loan translations from Latin (that is, not Latinate, but the pieces translated to German pieces) so visually don't look right, but are made up of squashing basic words together. Vacuum cleaner = Staubsauger = dust sucker. Easy to guess.
  • pronunciation - 2 mostly the same as English, no distinctions but a few strange sounds to get the accent right (r, ch, pf, v/w/f)
  • writing/spelling - 1.5 Latin alphabet, spelling very regular
  • cultural - 2 not part of US/English linguistic consciousness. Not taught at schools. Large subpopulation of people with German heritage (both US and England) that is entirely ignored as a heritage. Yiddish, a variant of HG, has a larger cultural heritage (from a much smaller population). But some things are in the fabric of English culture (Grimm's Fairy Tales, general northern European culture)

Russian - for IE, hard to learn basics, hard to master because of so many rules.
  • grammar - 5 conjugation of verbs sounds like English, gender and 7 cases of nouns. And even more exceptions to these. Getting the basics is hard and then it steadily gets worse. 
  • vocabulary - 4 there is a lot of borrowing for educated vocab directly from English, but no where near as much as English from Latin. So most vocab unfamiliar
  • pronunciation - 3 lots of consonant clusters and weirdly placed palatalization, and 
  • writing/spelling - 3 Cyrillic alphabet, spelling mostly regular, except for 'o'
  • cultural - 3 not part of US/English linguistic consciousness. Not taught at schools. Small subpopulation of people with German heritage (both US and England) that is entirely ignored as a heritage.

Irish - This is the most surprisingly exotic. It is definitely Indo-European, but it changes as many parameters as possible within that system to make it just plain weird. Nothing seems to be regular at first sight. Also at second sight. But eventually it is very regular, just there are so many context rules that are hidden or look like they overlap, but actually aren't
  • grammar - 4 if only gender were the only difficulty, but there's so much more. VSO, prepositions that seem to conjugate like verbs, lots of initial sounds changes with exceptions on exceptions. The rules are regular but getting them is hard for a beginner. Oh, and 5 declensions for 2 cases with nouns.
  • vocabulary - 3 not really cognate, except maybe numbers (but to complicate it, two sets of numbers that lenit or elide depending on the number). Plurals and verb roots don't seem regular at all.
  • pronunciation - 3 the sounds aren't strange (except for 'ch' and 'dl' and 'gh' and dh' and slender vs broad and etc etc), it's the sound changes based on grammar that are insane.
  • writing/spelling - 3.5 Latin alphabet, spelling is very regular but in the most misleading way. Lots of letters (vowels and consonants) sound nearby but different enough their English counterparts, if they're sounded at all. Should be 4,  except it really is rule based.
  • cultural - 3 Taught in Ireland, not in US. Very small world community. Huge subpopulation of people with Irish heritage in all former English colonies...where Irish is entirely ignored as a language heritage. Not much out there to read.


Persian - This is the most surprisingly non-exotic. No gender, very little conjugation of verbs straightforward word order.
  • grammar - 2 no gender, little conjugation of verbs
  • vocabulary - 3.5 some cognate words, but educated has a lot of Arabic borrowings. But mostly unfamiliar.
  • pronunciation - 2 a couple of uvular throaty g's (gh and kh) but not the sore-throat of Dutch or Arabic
  • writing/spelling - 3 Arabic alphabet, no vowels. Hard to get pronunciation unless heard before. Also it seems like all typesets of Persian/Arabic are much much smaller for a given fontsize
  • cultural - 4 very little connection with the West. Mostly Islamic. 


Arabic/Hebrew - exotic grammar and writing, but at least it has an alphabet.
  • grammar - 4 conjugation of verbs and nouns similar not by endings but by vowel changes within the root, gender different enough to English to be annoying. 
  • vocabulary - 4 entirely different. some direct English borrowings but not enough to make a difference.
  • pronunciation - 3.5 a few extra throaty sounds and s/t/th's that they distinguish but don't exist in English
  • writing/spelling - 3 Semitic alphabet (no vowels) (Arabic and Hebrew are similar like Roman and Greek), so you have to know things first to read properly (makes learning difficult)
  • cultural - 3 Hebrew has a lot of cultural connections (Christian and Judaic vocab). Arabic has the difficulty that MSA (Modern Standard Arabic) is not really used in the street or at home, and each country has its own dialect (which two countries over may not be intelligible). Why Arabic and Hebrew together? Salaam/Shalom should be enough to convince you


Chinese - exotic, learning writing takes as much work as an additional language. I'm referring to Mandarin, not Cantonese or Shanghainese, which are unintelligible foreign languages (very similar but unintelligible as English is to German). But just as hard (or harder if you think having many more tones is harder. which it is).
  • grammar - 2 no gender, no conjugating, no declining, no nothing. Straightforward order (slightly different from English), few exceptions. 
  • vocabulary - 4 entirely different. But like German, new words are created from pieces of smaller words. Semantically it's not always obvious how though. Words and phrases made from these short simple VC words have this otherworldly feeling that you can't get a hold of them, maybe because they're shorter or all sound a like.
  • pronunciation - 3 tones are exotic, and distinguishing some palatals is exotic. but that's it. lots of near homophones (see pronunciation) so lots of room for misunderstanding or puns. Not as bad as it seems. 4 for tones and palatals but 2 for everything else. Cantonese has like 12 tones, so count yourself lucky.
  • writing/spelling - 5 Ideographs, every syllable has a picture. Like learning an entirely additional language. It's not as terrible as that; most characters actually have two distinguishable parts, one as a hint to pronunciation and one. Supposedly only 2000 characters needed to read a newspaper. Only?
  • cultural - 4 very different cultural history so interpersonal expectations can be very different

Japanese - exotic, shares ideographs with Chinese but the spoken language is entirely different
  • grammar - 4 (my vague impression)
  • vocabulary - 5 entirely different. (my vague impression)
  • pronunciation - 4  (my vague impression)
  • writing/spelling - 5+ Chinese borrowed ideographs, plus two syllabaries (hiragana, katakana), and oh plus maybe another alphabet. So like Chinese but worse.
  • cultural - 5 very different cultural history (my vague impression)

One take away is this, that because of extensive borrowing from nearby or imperial cultures, languages that are different grammatically may have learning eased by having some vocabulary overlap. So European languages tend to be easy to learn among each other for pronunciation and vocabulary (and if one Romance is your native language then all but Latin are all 1's), Arabic and Persian with vocab, and Chinese and Japanese with respect to writing.

Of course this is an entire oversimplification.  There are much more refined metrics on language learning: a scales of language proficiency like ILR (0 to 5), CEFR (A1 to C2), or ACTFL (novice to distinguished). And each level has published expected hours of instruction to reach it. So these are admittedly very subjective assessments on my part.

The best way to learn a language is to be born into it. If you can't do that, have relatives in your household who speak it and nothing else so you can't rely on the one you already know. As those strategies are not always at hand, immersion is the next best thing (but also where you can't fall back on the crutch of your native language). As that often poses a chicken and egg problem, your most realistic first attack is study habits, lots of listening and reading. And the latter is what I think I am 'scoring' for English speakers.

It's only fair to go the other direction. One might think that it is a simple inverse of the above (or really since difficulty is a distance, exactly the same numbers as above. But that's not exactly the case. There are language to language difficulties (or simplicities) and universal ones. I will oversimplify considerably and talk about all languages trying to learn English.
  • grammar - 2 basic grammar as simple as can be. No gender, little inflection or irregularity. prepositions and phrasal verbs have lots of exceptions and ambiguity. I'd give it a 1, but I'm trying to account for familiarity bias.
  • vocabulary - 2 easy for Europeans, not so much for outside of Europe. 
  • pronunciation - 3 should be given 5 for 'th' alone which hardly any other language in the world has. Short i and diphthongs are weird but you'll get over it.
  • writing/spelling - 4 For a logical roman alphabet, its spelling is atrocious. There seem to be more exceptions to exceptions than the rules themselves. Almost like ideographs, you feel like every word has to be learned by itself.
  • cultural - 1 Because of English colonialism, American economic strength, entertainment media, and commercialization, English is everywhere. Movies, music, TV, the internet, everything you buy, something about it is going to be English. Everything is available everywhere in English. Even if it is exotic to you, it is so in your face you can't help but know about it. People elsewhere know more about English speakers (mainly Americans) more than they do themselves. There's so much opportunity for learning English.
If it weren't for the spelling part, English would be ideal for learning (only Chinese and Japanese are more difficult to learn in orthography). Grammar, vocab, pronunciation are easy for everybody.

Sorry if I left yours out. I'd like to add in other representatives, Slavic (Czech), sub-saharan African (Swahili), Indian subcontinent (Hindi and Mayalayam) and southeast Asia (Thai and Indonesian), to get a better world coverage.


Handy table to compare (and more easily see where you disagree with me)




LanguageGrammarVocabularyPronunciationWriting/SpellingCulture
French32221
Spanish2.52221
Latin52212
German3.53222
Russian54333
Irish43333
Persian23234
Hebrew/Arabic44334
Chinese25455
Japanese4545.55
English22341

Wednesday, April 20, 2016

Weird pronunciations in medicine

- pathology - obviously pronounced puh-thology (in IPA /pÉ™ 'θɑ lÉ™ dÊ’ij/) But all I ever hear when docs say it is path-ology (IPA /'pæθ 'É‘ lÉ™ dÊ’ij/), the first syllable not unstressed sounding like 'path'. Since everything in life must have a reason, I wonder what it is. Is it an attempt to differentiate it from something that sounds similar? Do they just want to emphasize it somehow?

- patent - obviously pronounced pa-tent (in IPA /'pæ tent/). Like patent attorney, patent leather shoes, patently false. But docs use it to refer to, say, a vessel or duct or that is not collapsed, that is full of liquid or air keeping it mostly cylindrical-ish, never mushed or squeezed down (blocked or occluded or limp). And when they do so, they say pay-tent (IPA /'pej tent/)

There are others. Doctors are weird.

Foundations of Math: Sets vs Categories vs HoTT

There's been a controversy lately about the foundations of mathematics. It was controversial when it first was discovered in the late 19th/early 20th century when the foundations were being developed, controversial mostly because most mathematicians didn't feel its necessity and also because as a young field things just hadn't been resolved. But some of the greats (Hilbert, logicians) thought foundations were needed.

The development of foundations started almost entirely with logicians and philosophers (Boole, Cantor, Peirce, Frege, Zermelo, Russell) creating the axioms of set theory all based on the concept of set membership. This was not particularly in the center of mathematical development. It was essentially completed in the 1920's, then controversialized by Gödel in the 1930's, but was refined by him and other logicians (Church, Tarski, Putnam,) and set theorists over the next 30 years.

Then in the 60s, an alternative foundations was proposed and that was category theory all based non the concept of maps betwen objects (morphisms) representing similarity. It felt like category theory could model most any mathematical field that way, even if objects and morphisms were not central, and the diagrams it created seemed to show similarities among the disparate fields of math. Category theory itself was created in the 40s by MacLane and Eilenberg but Lawvere began in the 60s to think of it additionally as an alternative foundations.

More recently, the 2000's, higher order type theory (HoTT) has been proposed and developed as an alternative foundations (mostly by Voevodsky). I don't know much about HoTT other than it seems like a subset of Category Theory.

Without going into the details of each of these foundations, it seems to me, as a semi-educated outside who knows these fields very superficially, that the word 'foundations' means something different to these different alternatives.

For set theory, it is a foundation for reasoning, for specifying proofs, for how you can justify things. For the others, the foundation is for content and similarity among branches of mathematics, what is known in mathematics. That is, these aren't really exclusive foundations, they both have their purposes and they are non-conflicting they're just different kinds of formalisms rather than competing for the same ideas. It is also an intra-mathematical result that they are interpretable in each other; the concepts in one are translatable to concepts in the other. You can do HoTT with set theory as the underlying basis and similarly you can do set theory with elements defined in terms of HoTT.

My preference is for set theory. Another reflection on this is that set theory and axiomatics is easy to explain early (to those with less mathematical maturity) using very basic mathematical analogies, and category theory and HoTT is more difficult needing lots of higher math to understand and use as non-trivial examples. Also, set theory includes logic (somehow, I'm not sure if is as additional machinery or if it is considered a consequence of set theory), and I don't know what kind of inference mechanism is allowed with HoTT or category theory.

Of course this is entirely biased of me because I think that Turing Machines form the best foundation because it is very simple to understand and is entirely operational, and is a metaphor for all computation which is all that mathematics is good for anyway!

Tuesday, March 29, 2016

Can Deep Learning be applied to Automated Deduction?

Deep learning is just a deeply layered neural network, and by deep they mean more than one internal layer. It has recently gained much attention because of its successes with images and go and speech and all sorts of things plain old NNs weren't doing so well at.

But what about automated deduction? That's not rhetorical, this is entirely speculative and answer free. I am wondering if DL (or any ML technique) could be thrown at AD (or ATP (automated theorem Proving) whatever you'd like to call it).

First the optimism: wow wouldn't that be cool, a method that would prove really hard mathematical theorems (the ATP community has to work hard and do a lot by hand to do things like flyspeck) or even . But desiring the outcome doesn't say anything about the implementation.

Next the pessimism. By analogy, images and speech/text are iffy optimization problems. Lots of perturbations get you the same thing. But logic (and similar combinatorial problems) are all or nothing. It is either true/proven or false/wrong. how could an approximation optimization method translate to logical combinatorial problems? Where do you get the scads of supervised instances required by DL? Millions of tagged images from MNIST but from TPTP literally only thousands.

So, I don't know. But just because I don't know how to do it doesn't mean it can't be done.

Monday, March 28, 2016

A small nitpick about a small p-value problem trope


Lately the p-value has been getting a lot of press, almost entirely bad (tip of the iceberg). Whatever it is and means has been up for discussion as it hasn't had since NPHT was created (the Bayesians have been fighting against it, or rather fighting for alternatives, since the 60's).

Hidden in the middle of this storm, or a small tornado on the side, is the issue of data fishing or p-hacking. Since only a p-value of less than .05 is considered 'statistically significant, only such values are considered publishable leading to the problems of: selective publishing (ignoring significant non-results), and p-hacking (if one p-value isn't good enough, change you hypothesis and testing little by little until you get a p-value below the threshold). The problematic trope is stated in roughly the following manner:

At p-level threshold set at 5% (or .05 or 1/20), all you need is 20 studied hypotheses to get one hypothesis that is significant by chance.

It's so obvious!  With 1/20th probability, you need 20 tries to guarantee a hit! The intention is that you shouldn't make many hypothesis tests at a time, otherwise you'll get some false hypothesis stated as true.

But you may notice with that wording that it is a classic gambler's misinterpretation, 'the run has to end!'. Each hypothesis test is independent, and so the probability of the next test will not change if all the previous tests are all hits or all not or whatever.

Whatever you think of p-values, and whatever you think they mean, they are probabilities. Probabilities of what is complicated and nuanced and misleadingly stated and problematic and the firm basis for statistical inference for the past hundred years. But still, they are probabilities of something and a strict threshold of 5% of accepting a hypothesis over rej... forget that verbiage. it's a 1/20 probability event.


So now we're in the realm of basic probability (and its own difficulties) but they should be shared by Bayesians, Frequentists, Kolmogorov..ans, Keynesians (he had his own!). So any hypothesis has a probability of .05 of being positive, a hit. What's the probability of a hit in 1 trial? 5%. What's the probability of a hit (at least one hit) in 2 trials? 3? n trials? Those are harder but only a tiny bit, basic probability/combinatorics. What's the probability of at least one hit? One minus the probability of no hits at all. What's the probability of no hits in n trials? (probability of no hit in one trial)^n. They are independent events so you multiply. Final answer:


\[
P({\rm hit\ in\ n\ trials}) = 1-P({\rm no\ hit})^n = 1- 0.95^n
\]

Well, not final exactly, it's just a formula. We don't yet have a good picture of what it means in realation to our intuition about 'it'll take 20 to make sure we have a hit'.

So then a picture:



It starts at 0, rises in exponential decay asymptotically to 1, but is a little slower than you'd expect for an exponential because the base is so close to 1.


The usual way to present such probabilities is, like the birthday paradox, to say how many trials it takes to get %50 chance. Multiplying .95 a few times we see that it takes 14 trials for there to be more than 50/50 chance that at least one item is a hit. To get to 95% chance, it takes 59 trials. There's no guarantee that there'll be a hit, just the probability gets smaller and smaller.


I make this point to... well, to pick a nit. You do more experiments, the more likely there will be one that is 'statistically significant' totally by chance. Intuition and logic lead to that immediately. But the logic is never done and one step of it doesn't lead to correct two steps. I do realize it is a bit of a mouthful and difficult to digest to say 'after 14 trials there will be a fifty-fifty for a false positive'? 50/50? It's not obvious how that relates to 5%, but the erroneous 20*.05 = 100% does obviously relate.


In the end, to say '5% means 20 experiments', which seems so directly and intuitively obvious, is wrong. In the right direction, but wrong.








Sunday, March 27, 2016

A serious proposal about Daylight Saving Time

People everywhere hate the spring change in clocks for Daylight Saving Time.

First, it's annoyingly tiring having to get up one hour earlier. And feeling like you can stay up a whole hour later. And this leading to wanting to stay in bed that much longer even more. Annoying.

Second, the average effect of this annoyance on the population. This leads to, statistically, 'studies have been done', of reduced physical performance leading to accidents (lower light levels from sun than in the previous week), and an increase in heart attacks and strokes (not exactly excuses to lie in bed a bit longer).

DST changes things twice a year. In a preindustrial society, without fast travel and communication, time zones are not useful. DST is useful for... for what I'm not sure. More daylight in the evening? That's not for work, that's for ... I don't know, kids to stay outside and play longer?

But if that is desired, but also we'd like the feature of no changing of wakeup time, I have a radical new idea, which though it sounds strange, is not entirely a crackpot idea. My proposal:

Have clock time pegged to sunrise.

And call that time 6am. No matter where you are on Earth.

First there is the practical consideration, then there is the implementation. Practically, it fulfills the feature of never changing. You always get to wakeup at the same time. Nowadays we all choose to go to bed pretty much irrelevant to sunset anyways. And the farmers don't have their delivery schedules screwed up when the milking cows don't care about these crazy human practices.

Implementation may be the most difficult, and here is where some of the subtleties come out. However, this is the 21st century. The point is that sunrise is equivalent to longitude, and this is determinable now by any variety of methods but also by pure calculation (whose coefficients were determined by those methods or in the end skywatching).

We all have at our disposal computation devices that are minute in bulk (phones, watches, RFID chips) with GPS positioning (via satellites), that can determine longitude within a few meters (2?). So sunrise will be a few seconds off from someone 10 miles west (15 degrees = 1 hour sunrise difference, 1 degree ~= 60 miles ... 10 miles ~= 24 seconds). This continuous difference sounds like it is insurmountable... except by computation, which is currently trivially executable. You're having a phone meeting with someone at 9:30am, you're in downtown Chicago and they're in Kansas City, MO (~415 miles west). That's around 7 degrees or 28 minutes. If the appointment is at 9:30AM Chicago Loop time, in Kansas City it is ~9AM. The appointment making software would calculate the time for the recipient based on their address. For coordination there can still be a UTC which is the time at one particular longitude.

There are some very minor difficulties that are barely in the realm of practical considerations. Depending on your latitude, noon (or six hours after sunrise) may not correspond to the highest point of the sun in the sky, but that's the case now anyway (just not as much). Also for higher latitudes, sunset will change drastically in the spring in fall (sunset coming later one day to the next by up to 10 minutes, making the feel of the evening change within the week.

There's cultural precedence for this. Many ancient calendars count the change of day (but not when to rise or go to bed) to sunset, so why not sunrise. Our mental perception of the 'next day starting', despite the European style (derived from Roman) changeover at the sun point opposite to sun's zenith (yes, I'm avoiding 'midday' and 'noon' for the moment),  is truly when we wake up around dawn or more naturally at dawn.

So that's it. Free idea. Go forth and implement. Just give me credit. If there are problems, I'd like to blame the future implementation.

Of course, this doesn't fix the problems of Jews and Muslims north of the Arctic Circle with practices on cooking or fasting concerning sundown.

Tuesday, March 15, 2016

Review: Ready Player One

"Ready Player One" by Ernest Cline, is a sci-fi novel about a kid, living in an impoverished near future, escapes like most of the population into a super-powered virtual world invented by a Steve Jobs style nerd. The plot of the novel is following this kid trying to find an 'easter egg' planted by this eccentric inventor bajillionaire, all by playing different video games or by using pop cultural references from nerd culture () most of these from the 70's and 80's (the dawn of blockbuster movies, and PCs and video games (arcade games, PC games, and pre-Mario Bros game controllers).

I don't know. The nerd-pop-cultural references were all there. Star Wars, Star Trek, Blade Runner, Monty Python, War Games. Joust, Defender, Dig Dug, D&D, all well referenced, lines of dialog. All well suited to my personal memory of teen years. Lots of excitement and virtual explosions and real explosions and plot twists and I'm sure hidden easter eggs in the novel itself.

But I'm an old dude. I don't see how kids these days (2011) would appreciate the references.

The tone is very much 'young adult novel'. It should be advertised as such, because it's not Philip K Dick if you're expecting that. The concerns are all of a cliché 18 year old (does that girl like me? I will obviously win this game having played it literally hundreds of times or literally never! I hate my welfare aunt who I live with since my parents died.)

The writing is like Asimov. I don't mean that as a compliment. Despite its obvious attempts at being modern and PC, they all come across as clunky white-male-teenager-privileged like a nice mormon dad coming to terms with a child caught drinking coffee. The main character's best friend on-line, a dude like himself but cooler, turns out in real life to, wait for it, it's shocking, are you sure you won't be shocked, a shy black overweight lesbian. Oh and the girl I like in the virtual world whose avatar is totally hot, in real life, she's also hot except she has a port wine stain on her face. The horror. The shame. The forced lesson of understanding of others without being other.

And the plot should be considered simply a string of the most unmotivated deus ex machina's and Mary Sue. And this is two levels deep, both in the engineered MMORG world and in the simulated games within that world. Oh did I mention this one hidden rule that you get 40 lives if you die the right may in this video game? Convenient that the best friend of the inventor comes to help out the gang at the last minute. Oh no, the opponents destroyed absolutely everything (virtual) with a  pixel bomb (sorry, that's not in the book but it's the shortest way to give the same idea), but the one tiny thing needed to complete this one magic task left happens to have survived? I'll win a bajillion dollars, I think I'll share it with my friends even though I won the game at the end, I'm a great guy!

The book is a lot of fun (moreso if you recognize the references). But it's mostly processed sugar and wish-fulfillment. Should be a fun movie, which should be able to avoid most of these clunkers.

Monday, March 7, 2016

An annoying argument/information trope

You know what really bugs me? When you're reading an article and it says:

"Last year, 350 bajillion Americans stabbed themselves in the face with a fork"

Yes, that's tragic. But the additional thing that annoys me is that nowhere in the article does it say how many people are using forks when they don't stab themselves in the face or at all, how often did people do it last year or in the world, or what the data source is, or if this was a meal accident vs kitchen clean up accidents, or what about people stabbing others like that, or if this is counting repeat offenders or stabby motions or what.

But the biggest annoyance of all is the bigness of the number. How big is that number really? Why that number (which usually sounds like it is quite a bit more than the number of forks or people with the capacity to do such a thing). It's like magical thinking, saying a big exact number invokes some mystical 'wow'. It's a very specific thing for a very fuzzy idea. Self stabbing bad. A lot bad. Wow, lots of bad things is bad. Wow.

Oh, so you want me to be specific? OK, specifically any article anywhere that mentions a number. Mostly science popularization articles.

Sure, the main problem is the base rate fallacy, but it's so much further down the scale than that. It's the 'Wow, a number' fallacy.

Monday, February 1, 2016

Chomsky vs Google, Structuralism vs Behaviorism

It's 2015. Google has been in business almost 20 years (since 1998). Their primary tech model for web page search is (still) to split all findable web pages by words, make a word/concept vector for each page using the entire corpus of words and rank a word search using links to other web pages (the PageRank algorithm). It is primarily a data driven process rather than an expert driven process. Of course there are lots of add ons, hand-tweaking, special cases, but at the center is an automated process of using what words people have used and what they've linked to.

It's 2015. Almost 60 years ago, Chomsky revolutionized linguistics (assuredly pursuing already existing trends in structuralism). Though the revolution started in substance with 'Syntactic Structures', he also put in a a mortal wound to (already dying) behaviorism with his critique of Skinner's Verbal Behavior. Skinner's thesis was that people learned language by example, by mimicking the thousands of utterances, assimilating the patterns heard. Chomsky's critique was that, in so many ways, this was wrong. The variety of known languages had a narrow set of commonalities not explainable by the broad possibilities allowed by reaction conditioning (fixed action patterns, operant conditioning, stimulus response). People seemed to learn from negative information. Chomsky's criticisms were so incisive and convincing

So behaviorism went out of favor, and Chomsky's linguistics is still like Newtonian mechanics; even if there's a linguistic Einstein, it will only be a slight correction to the very accurate approximation that is Chomskian language theory.

But Google. It is so obviously successful (sure as a company, but their search capability). More relevantly, Google Translate is essentially an ngram analysis. It is really good! Surely it makes many groaning mistakes on uncommon languages, but as time goes by it is more and more successful. It is a poster child for behaviorism. It is behaviorism embodied in a machine. And it works so well.

The point is that Chomsky won 60 years ago with, not exactly rationalism but more like introspection (which had lost earlier). And now Google is winning with pure un-introspected descriptivism.

But why? Why is statistical NLP so (currently) successful (when based on single words or short sequences of words, not constituent phrases) and syntactical NLP (POS, parse trees) is, well, not exactly wrong, just not terribly useful?

Maybe it's machine performance and having lots of data.
Maybe the words themselves have a lot of meaning and just being in the same sentence is enough.
Maybe understanding parse trees is very informative but only under toy conditions; the bag of words has the bulk of the information.

Google is killing Chomksy. It's like saying that Popper killed Logical Positivism. He did, sort of, but along with others, and anyway it's not really dead, just a basis for what comes afterwards.

Friday, January 22, 2016

Metaphor is a Metaphor. A Leaky Abstraction is a Leaky Abstraction

Bear with me.

Metaphor is a metaphor. A metaphor is a transfer is a carrying over (meta- (GR) = trans- (L) = over, across, -phor (G) = -fer (L) = carry, bear). A metaphor is a transfer of meaning from one literal sense to a figurative one where, hopefully, the transfer to the new domain keeps things one-to-one.

Imperfect metaphors are called leaky abstractions (at least in software). And 'leaky abstraction' is a leaky abstraction. They leak because the literal source of the abstraction leaks through to the abstraction layer. An abstraction or a metaphor is metaphorically used as a container of fluid some of which leaks out (but onto what? which is the source and which is the target (in another metaphor of metaphor)). You can't take the abstraction layer literally, you have to know something about the underlying source.

Also, 'figurative' is figurative or metaphorical; I often take metaphor as a synecdoche (or metonymy; synecdoche is a metonymy of metonymy) of figure of speech, in that it, figure of speech, is not literal. A figure is a picture, which is a good metaphor for metaphors... or figures of speech. 'Literal', on the other hand, happens to be somewhat non-literal because it is about writing, which is a metaphor for verbatim, that is the primary surface definition. 'Literally' has been used non-literally (that is as a general intensifier) literally for ages, but universally (that's hyperbole which is a figure of speech which is just a lie that we all agree to and not metonymy) recognized as wrong.

Borges said (in This Craft of Verse) that Lugones said (in Lunario sentimental) that "all words are dead metaphors", which is a dead metaphor because nothing has really died. Or rather the original meaning died or faded away very slowly, but you could resurrect it a little if you tried. So it's a little leaky.

We're swimming in metaphors!

Also 'Metaphors We Live By', by Lakoff and Johnson

Thursday, January 21, 2016

Technical Debt is a Leaky Abstraction, but so what

Technical debt is a recent term in software engineering used to describe potential later problems that may be caused by known decisions now. For example, suppose a customer asks for a feature that can be implemented cheaply and quickly, but it will introduce security holes or be very difficult to generalize if done quickly, or will prevent another rare feature from working without lots of rework. You're making things work now for some pain later. You would have much less total pain overall if you 'do it right' today, but cheaply and quickly are more important.


(from collab.net)

Saying 'we have a lot of technical debt in our software' is very loose talk for 'we have either a lot of bugs that no one is complaining about' or 'we don't do code review so I bet there's a lot of crap that gets deployed'.

A leaky abstraction is another metaphor in software engineering. When one creates a new layer intended to hide all the details of a lower layer, For example, floating point numbers are leaky because they try to hide all the gross details of finite bits per number approximating perfect precision of the desired  numbers, but sometimes you end up having to be aware of the underlying implementation because they leak through the abstraction (eg a+b usually equals b+a except when b is much smaller than a, and to deal with that you need to know some details of how the floating point ops really work under the hood).

Technical debt is a leaky abstraction in that the metaphor cannot be taken too literally; if you try to follow implications of the words, it breaks. Debt is a balance in a ledger, you can owe or be owed a value. Technical debt, being very qualitative, is hard to put into numbers, you just have a vague sense of 'this is bad' vs 'this is really bad' vs 'this is tolerable'. The debt isn't really about the features themselves but about the time and mental effort needed to implement things.

The first picture is a terrible explanation, the following is much better:


(from commadot.com)

All I'm saying is that technical debt is a leaky abstraction, a faulty metaphor. There's no paying back, it's just cleanup.

Using the term is great because it is more politic than saying 'The code is a mess and needs some cleanup. Features are really hard to add'.

Tuesday, January 19, 2016

Names make theories make names

Richard Feynman, the Nobel Prize winning physicist and Challenger disaster O-ring explainer, has a couple of anecdotes about naming. 

One is about the 'map of the cat', where Feynman had to give a talk to his graduate zoology class. In preparation, he went to the library and asked for a map of the cat, to which the librarian responded "You mean a zoological chart!"  (I'm guessing 'what a funny thing to say') Then later when he presents this, his fellow bio students say "We know all that!". Feynman says that they "had wasted all their time memorizing stuff like that, when it could be looked up in fifteen minutes."

The other story is much earlier in his life. All the (nerdy) kids on the playground try to one up each other on what their dad's taught them. "What do you call that bird? What about that bird?". But Feynman's dad said  something like it's called X in language Y, Z in language W.


You can know the name of that bird in all the languages of the world, but when you’re finished, you’ll know absolutely nothing whatever about the bird. You’ll only know about humans in different places, and what they call the bird. So let’s look at the bird and see what it’s doing—that’s what counts.” (I learned very early the difference between knowing the name of something and knowing something.)

What is the point about these parables? He goes on to explain that names don't explain anything. He explains that these technical terms for anatomy or for bird varieties or for whatever science are just terms, they're not explanation themselves. There is a tendency even for these technical terms to be magical invocations, the totems of a closed guild, supplications to the gods, when these terms are barely paratactic gestures of pointing, a superficial 'behold', an inarticulate label with no explanation of the depth of the experience. "What kind of bird is that?" "It's a sparrow" "Why does it do that?" "It's a sparrow." as though emphasis is explanation enough.

Except... what do you expect? Is naming so terrible? Do you want to do away with naming and simple move on to the much more interesting explanation? Is naming so simpleminded?

Where did the names come from? Without this being an explanation of historical linguistics, presuming terms like gastrocnemus, sparrow, and inertia are somewhat random and distinct, these labels are hooks for the concepts. A lot of thought, using other labels or conceptual manipulation, led one to label this object or concept one thing, that another. If our thoughts are not necessarily tied to language, communicating them certainly is (though a gesture or picture can go a long way, like a shrinking ring in ice water). 

Names and words are little theories in themselves. We learn most of them superficially, but eventually we acquire their nuances. A new word, like inertia, is opaque to the newcomer, but at some point in time, the scientist or experimenter was playing with a number of concepts, and eventually some concepts coalesced out of that thinking and one of them was given the label 'inertia'.

Yes, knowing a name doesn't explain anything. Or rather, it explains very little. Knowing how to use a name is nontrivial, but has little explanation to it. Answering why requires being able to manipulate a number of names, but having those names is necessary. 'Black hole' is the culmination of a lot of thought. That it is a thing is the consequence of lots of thinking. And the start of a lot of thinking. Once you get to that concept, a lot of thought has gone on and not having that term would be a great loss, leaving us to swim around a number of concepts of relativity but don't quite say what they really mean. Feynman is saying that it is dumb to stop at the names of things. Sure, don't stop. But don't let that stop you from naming things because it'll be that much harder to continue without the name.

A name is itself the end product of a theory, and it makes further theories possible. 

Wednesday, January 13, 2016

Prescriptivism vs Descriptivism: which is worse?

Prescriptivism vs Descriptivism: which is worse?

These two words are used to describe one's attitude towards language usage; at its very simplest do you prescribe or describe how you speak, what are people supposed to do vs what people actually do, 'should' vs 'is'. When these terms are thrown around (and I do mean thrown, like mud pies) it's almost always meant to sting.

Objectively, prescriptivism is usually understood to mean keeping strictly to the formal rules of a language and descriptivism is more about discovering and recording the rules of language however people say things (whether they match the formal version or not). Newspaper editors and school teachers are often the supposed standards of prescriptivism and linguists as that of descriptivism. Your secondary teachers are teaching you the rules of good grammar, and the linguists are being scientific about what the rules actually are.

M-W's third edition dictionary (1961) is often cited as a classic of the descriptivist abyss, putting in words of dubious provenance, all the good profanity (which at least got some kids to crack it open at least once). The dictionary was decried by many as the nadir of pandering to idiocy, the last gasps in the decline of western civilization.

Informally, from the prescriptive point of view, descriptivists are 'anything goes'; whatever people say is what is allowed, nothing and no one is wrong, there are no mistakes, 'ain't' and 'between you and I' are now OK and that is just wrong and descriptivists are avatars of the decline nay the destruction of western civilization, regression towards the mean, the twilight of the idols, the idiotocracy, the worst are full of passionate intensity.

From the descriptive point of view, prescriptivists are stuck up old school marms, who make up arbitrary style rules, say 'you can't do this-you can't do that', split infinitives, prepositions at the end of a word, singular they, when people have been saying it that way forever and you just made that rule up because you are warped, frustrated old man. P's try to enforce their made up rules, when it's just their own repetition of some one else's peevish peeving on personal style. Descriptivists say that prescriptivists single way of speaking is reprehensible elitism and that they think they're are morally superior to others, and any other patterns are slack jawed, uneducated, lower class.

But that's just the tendentious version.

From the prescriptive point of view, there really are mistakes that people make, infer for imply is just wrong, 'literally' for not literal things is just wrong and native speakers just do not say it that way. A common error is not necessarily a common alternative.

From the descriptive point of view, there are many patterns out there for the same thing. Different contexts have different rules. People will say different things in different situations, one way speaking at the press conference and another in the bar, and neither is wrong (or the differences show up in different contexts).

People may very well avoid a split infinitive and prepositions at the end of a sentence in writing for stylistic or esthetic reasons, but in speech there's hardly getting around it what with all the phrasal verbs in English. The double negative ain't no (= "isn't a") problem as a perfectly everyday way of speaking for some varieties of English (that are, as the linguists say, not highly socially respectable (= redneck or AAVE)) or grammatically and logically appropriate either, two negatives making a positive or a form of understatement (that was not an uncomplicated sentence).

And this is where it comes down to the real difference.

Descriptivists are really prescriptivists at heart. Descriptivists just recognize many more varieties than prescriptivists, and those varieties tend to be much more informal or used by socially non-pinnacle subpopulations. People who are called by others prescriptivists do seem a little judgy, and people who are called descriptivists do seem a little too accepting of things that are  (i.e. errors). But if you just label large groups of patterns as varieties, Prescriptivists are just talking about (mostly) a single variety, the newspaper/college paper variety, and descriptivists allow for a wider range of varieties, informal or regional or inarticulate (well, maybe not the last one). There are still mistakes. It just depends on the context or variety you're in.