...by Daniel Szego
quote
"On a long enough timeline we will all become Satoshi Nakamoto.."
Daniel Szego

Wednesday, August 15, 2018

Enterprise blockchains and identity management


Enterprise blockhcain solutions has usually totally different directions regarding identity management and user authentication. They usually integrate some kind of cloud or more decentralized identity management system, efficiently authenticating and personalizing every transaction and access to the system. Examples are Azure Blockchain as a Service with the help of Active Directory Services or Hyperledger Fabric Composer with the help of Membership Service. It is totally different from the open public world however, where you generate a public private key pair and practically the private key represents the identity. However even in enterprise scenarios, the two approaches should be combined. One part of the service might be authenticated with the help of classical centralized identity services, however part of the application could follow classical antonym authentication methods with the help of private - public keys. The authentication situation might be an analogy to the a classical portal system, where you can reach some of the content publicly on the internet as other content and services are protected with classical identity management systems. 

blockid in a multi-hash blockchain system


In a multihash blockchain system, individual blocks do not have one hash identity but several ones, at least one pair for each hash link. This guarantees the consistency of the blocks. Despite if one wants to  download or refer to a block, it is much more practical if we have one common id, that can be refereed either at the communication protocol or from the blockchain explorers. Such an id can be for example the hash of all of the exiting hash links and hash pointers. Certainly, it is an open question if such a construct open a new attack vector or vulnerability, because the block hash directly is not contained in the hash list of the blockchain. 

Tuesday, August 14, 2018

Blockchain and cryptographical rolling hash


Blockchain solutions has one of the key advantage that they are pretty much hacking resistant, because of the chain of blocks and hashes that going back up to the genesis block. However, most of the applications do not necessarily require such a huge security guarantee, it is probably enough if data can not be hacked and consistent considering the last couple of months or years. On the other hand, there can be the requirement to delete from a blockchain solution or has the possibility to forget or modify data. On solution might be to design a blockshain platform where data is not hashed back to the genesis block, but only the last couple of like thousands of blocks are considered. Certainly, it is a question how such algorithm can be realized in a hacking resistant way. One solution might be to try to use instead of cryptographic hashes, a kind of cryptographical rolling hashes, where only a certain number of past values are considered as input.  

Sunday, August 12, 2018

How to implement a Blockchain from scratch - smart contract simplified


In a simple account/ balance/ state based blokchain system implementing smart contracts is pretty straightforward. Accounts represent for the first run not necessarily just balance but a kind of a general data as well that can be modified by the smart contracts. In order to create smart contracts, you should define the language or smart contract programming environment itself and the effect that a smart contract can result in the state. Certainly one way of doing it is to define a virtual machine which guarantees that the smart contract is executed exactly the same way on every peer. However we might as well consider an exiting virtual machine as well, like the java virtual machine and limit somehow the effects of the program. As an example a simple smart contract could: 
- read some of the state information of the blockchain manifested by accounts data and balances. This state information is the previous block on which we want to mine our contract. 
- having some computation on top.
- changing the data value of a certain account. 
- storing the smart contract code somehow as data or string, like with the help of serialization
- creating a special transaction containing the smart contract as data with the sign of the account that you want to modify, indicated indirectly the owner of the account allows the data to be changed. 

At mining process:
- The signature of the smart contract transaction has to be checked. 
- It has to be made sure that only the effected accounts are modified.
- The code has to be run and the new data value must be calculated. 
- It has to be made sure that the smart contract does not cause infinite loop, one way of doing it is to avoid general loops, or to terminate the contract after a certain number of iteration resulting the transaction as invalid. Certainly another way can be built in as well, like with the help of a cryptoeconomical mechanism longer smart contract runtime can disincentivized, just like as with Ethereum.
-  The new value or values have to be applied.
- The block must contain the valid transaction and the new valid state, which is the new value of the computed accounts. 

At validation process, practically the same steps have to be repeated, without the last one, which is putting the transaction and state to a new block and doing proof of work calculation or voting of a byzantian consensus mechanism:
- The signature of the smart contract transaction has to be checked. 
- It has to be made sure that only the effected accounts are modified.
- The code has to be run and the new data value must be calculated.
- The calculation must have finite time. 
- It has to be checked if the new values of the state of the given are the calculated values based on the values of the previous block.
  
The wallet functionality has to be extended:
- to have the possibility to write or integrate smart contracts.
- to transform the programs into data, like with the help of serialization.
- to create transactions based on the smart contract.
- to sign them.
- to broadcast them on the network. 

How to implement a Blockchain from scratch - extended wrapper classes


At designing at implementing a blockchain system from scratch, there might be some contradictory design perspectives: on the one hand elements of the blockchain that are stored or transported on the network must require as less storage as possible, resulting in lower bandwidth or data storage requirements. On the other hand for efficient processing, some further information is usually required. Examples are: 
- Blocks in a basic scenario should store the difficulty, the nonce and the a hash value of the these values values together with the merkle roots of the transactions and state and the previous block hash as well. 
- Blocks in an extended scenario might contain explicitly a link for the previous block, for the further, some information if it is an orphans block or on the block height.
- Accounts in a basic scenario should contain an address of an account, which is usually a public key, and some kind of a change request, like transferring money, or changing a value. On top, certainly a nonce value.
- Accounts in an advanced scenario are related rather to the accounts of a certain wallet, so they might contain explicitly the private key and meta-information if the account is synchronized with the blockchain, or still not available in the blockchain.

There might be a similar consideration for other elements of the blockchain as well, like Block Headers, Transactions or Peers. Practically every object that is moved on the network can be considered as implemented as a basic version containing only the relevant data, and as an extended version containing all the computational relevant data.

The new divide of the information society

There will be a new divine in the information society between people producing information and people consuming information. As information can be simpler and more automatically produced every day, like with the help of massive online courses, chat bots, and different online platforms, it will be more and more information produced by less and less people.

Services will become more and more automated, meaning that more customer service will be carried out by bots or robots, resulting a cost factor saving in the employees and enabling fully automated decentralized companies. As an average customer it will be more and more complicated to fight through the bureaucracy of  bots and robots and get to a "living" person.

Notes on fork resolving policy

In bitcoin network, there is a general rule, always the longest chain should be accepted as a valid on. However this is just one implementation of a so called fork resolving strategy which might be extended and generalized on a long run. Such extensions might include: 
- instead of just chains, considering the whole weighted structure of the possible forks, like in Ethereum taking also uncle or omar blocks into account.
- in a permissioned chain, blocks might be prioritized and weighted based on the fact who signed them.
- in an attempt to avoid selfish mining, the blocks might be weighted with the creation time.

Further implementation and algorithms might be also possible, it is only important to note that fork resolving strategy as taking always the longest chain is only one possible implementation.  

Saturday, August 11, 2018

How to implement a Blockchain from scratch - network protocols


At creating a blockchain solution from scratch, one of the most important part is to design a set of network protocols. It is important to note that these protocols are on the application level, under the hood they might be realized by simple socket sending on a TCP, something more abstract RPC or JSON-RPC, or even with onion routing. At any case, the following protocols should be considered:

0. Get client version: connecting to a known peer and getting the version of that peer. It prevents the usage of incompatible peers, that might be highly important especially if there are updates of the code regularly.

1. Update peer information: A brand new peers is usually started with a couple of preconfigured nodes further peer information. These peer information are queried for the further known peers until the number of known peers reach a certain limit (like 15 in Bitcoin by default). The peers may go offline and come back online again, therefore the active peers have to be checked regularly if they are still alive, if not different strategies can be implemented, like:
- deleting the inactive peer from the cached peer list.
- waiting for a certain time if the peer will be online again.
- querying all the still active node to get more peers that are hopefully active.
- or a combination of the previous strategies.

2. Syncronizing the blocks and states: As a first step of using the blockchin, the blockchain must be more or less up to date with the rest of the world. As a first step, the node can connect with all of the peers and ask the size of the blockchain. Based on that information, it knows exactly if the blockchain needs to be synchronized or not. As a second step a

3. Propagating transactions and propagating blocks: if new transactions or blocks arrive, they must be registered locally and further communicated possibly as fast as possible to all of the connecting peers in the form of a gossip protocol. If it is a transaction, it must be added to the local transaction pool, if it is a block it must be added to the possible top headers of the blockchain.  It is especially critical with the new blocks as the winner of the mining competition depends on the speed of the propagation. This phase should be available just after the blockhcain has already reached at least a quasi synchronized status. It is however an interesting general question how the propagating mechanism might stop, without effectively over flooding the network. One way might be to implement something as a finite hop system, in which a given information is propagated only at a given time. Another idea might be to implement regularly handshakes where peers exchange information on relevant transactions and blocks first before transporting them. 

In all of the communication protocols, it is always questionable, if we consider a kind of a push or pull protocol. In both protocols, the design direction should be to transfer as less amount of data as possible, and only if it is necessary. As a consequence, most protocols should transfer the inventory first, meaning the hash values of blocks and transactions. After that the nodes should be able to download the content of the hash values on-demand. 




Friday, August 10, 2018

Double spending in account - balanced based systems


Double spending in an UTXO based system is relative simple. Each transaction output can be considered as a "coin" that can be spent only as a whole and only by one transaction. It practically means that applying two transactions in a round for an unspent transaction means a kind of a double spending. The situation gets a little bit complicated if we consider an account / balance based system. One implementation of such a system is that there is a nonce both at the account and at the transactions as well: if the transaction nonce is by one bigger than the account nonce, implies that the transaction can be applied to the account. Otherwise we might as well have a replay attack. Despite of the logic it remains questionable how such a logic can be implemented. The versions are the following:
- After each transaction, the nonce is incremented by right away one. The problem is there are two transactions to the same account, than one must be considered as double spending and on the one hand must not be added to the block and it must be deleted from the transaction pool as well (as the nonce of the account already increased, there is no way that the transaction with the original nonce can be valid). As this is pretty much compatible with the classical UTXO based systes, it might not be the most performant solution and it might cause problems at use-cases where the only use-case is not just money spending, and there might be several actors working with the same account.
- Even with a simple money based blockchain system, it might be the use-case that I create to transactions in the same time for the same account in a way that they spend less money than the total amount of the account, and I certainly expect that both transactions are executed. If they are executed in the same block, the situation is actually solvable, as the miner can check and validate the transactions, and increase the nonce accordingly. However it might be a little bit tricky if the two transactions are executed in separate rounds, because the miner gets them in separate rounds from the transaction pool. 
- From a conceptual point of view however, the simplest consideration is to assume one account one transaction. In such situation a standard use would send one transaction at a time. If she wants to send several transaction as a batch of transactions, the wallet would simply create a set of transactions with an always increasing nonce. If coincidentally one nonce gets refused for some reason, the wallet can again create and sign the transaction with a bigger nonce. 

Trade of for showing information in a blockchain wallet or explorer


A blockchain wallet, usually stores only the local keys in the wallet, or from efficiency reasons some of the cached information. However everything that is relevant, like the balance of an account, the data of an account, or the existence of a transaction is stored on the blockchain. This implies an interesting dilemma for a wallet or even blockchain explorer implementation: which data should be shown to the end user. The two major versions might be  the followings:
- showing the most up-to-date data, in extreme cases even transactions that are already in the transaction pool but are still not confirmed. This solution might have the problem, that in case the blockchain forks, that might happen often on the top of the chan, the data will be changed or even disappear. 
- showing data only after a certain number of confirmations, like 3-10 have the advantage that the data will probably not change due to standard changes. However it has the disadvantage that the data might be simply old. 

A good solution might be to show every piece of data in a blockchain explorer or in the wallet with a kind of a reliability information as well, how much confirmation or reliability does the piece of data has. If it is a low-reliability information, it might be changed on a long run. In certain situations, there might be more than one different piece of data with different reliability as well, meaning that the data has been changed but the change itself is still not fully confirmed.